EDBT 2026 Demo / reviewers in the wild / expert
Feiyue Huang
dblp:53/520
· DBLP profile ↗
123ranked-venue papers
1as first author
54since 2021 · last 2026
0009-0006-8834-9646ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 99 · 1 first-author · 43 since 2021Artificial intelligence and machine learning · 80 · 1 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task-oriented medical image super-resolution via target prior guidance
Yuxiang Meng, Dengwen Zhou, Shaoxin Li 0004, Feiyue Huang, Xingkun Xu, Lifeng Zhu, Yuchen Xu 0008, Fan Tang |
Knowl. Based Syst. | 4 |
| 2026 | Semi-supervised crowd counting from unlabeled data
Haoran Duan 0001, Yawen Huang, Yang Long 0001, Xian Wu 0001, Feiyue Huang, Shaoxin Li 0001 |
Pattern Recognit. | 5 |
| 2026 | Boosting cross-domain semi-supervised medical image segmentation with internal and external regularizationsabstractCross-domain medical image segmentation has been challenging due to the extraordinary cost of collecting sufficient data in various imaging conditions. Mainstream methodologies enhance generalizability by directly adapting large-scale vision foundation models for medical image segmentation tasks. However, these approaches frequently incur high memory costs and necessitate additional prompts for deployment. In this study, we utilize dark knowledge in pretrained segmentation models as external regularization to improve the model’s generalizability. Furthermore, an activation-restricted regularization item is proposed to eliminate the noise/errors within the generated pseudo labels. Experiments on medical cross-domain datasets demonstrate the SOTA performance (above 2.66% improvement) of the proposed method with no additional cost during inference. Rui Wang 0177, Fan Tang, Feiyue Huang, Shaoxin Li 0004, Xinkun Xu, Yuchen Xu 0008, Lifeng Zhu, Weiming Dong |
Pattern Recognit. | 3 |
| 2026 | Harmonized medical federated learning via redundancy-aware client consistency
Jingjun Yi, Yuexiang Li, Qi Bi, Wei Ji 0011, Huimin Huang 0002, Yawen Huang, Yefeng Zheng 0001, Feiyue Huang |
Pattern Recognit. | 9 |
| 2025 | D-CAM: Learning Generalizable Weakly-Supervised Medical Image Segmentation from Domain-Invariant CAM
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Huimin Huang 0002, Yuexiang Li, Shaoxin Li 0001, Xian Wu 0001, Yefeng Zheng 0001, Feiyue Huang |
MICCAI (5) | 11 |
| 2025 | Towards normalized clinical information extraction in Chinese radiology report with large language models
Qinwei Xu, Xingkun Xu, Chenyi Zhou, Zuozhu Liu, Feiyue Huang, Shaoxin Li 0004, Lifeng Zhu, Zhian Bai, Yuchen Xu 0008, Weiguo Hu |
Expert Syst. Appl. | 5 |
| 2025 | Verification is All You Need: Prompting Large Language Models for Zero-Shot Clinical CodingabstractClinical coding translates medical information from Electronic Health Records (EHRs) into structured codes such as ICD-10, which are essential for healthcare applications. Advances in deep learning and natural language processing have enabled automatic ICD coding models to achieve notable accuracy metrics on in-domain datasets when adequately trained. However, the scarcity of clinical medical texts and the variability across different datasets pose significant challenges, making it difficult for current state-of-the-art models to ensure robust generalization performance across diverse data distributions. Recent advances in Large Language Models (LLMs), such as GPT-4o, have shown great generalization capabilities across general domains and potential in medical information processing tasks. However, their performance in generating clinical codes remains suboptimal. In this study, we propose a novel ICD coding paradigm based on code verification to leverage the capabilities of LLMs. Instead of directly generating accurate codes from a vast code space, we simplify the task by verifying the code assignment from a given candidate set. Through extensive experiments, we demonstrate that LLMs function more effectively as code verifiers rather than code generators, with GPT-4o achieving the best performance on the CodiEsp dataset under zero-shot settings. Furthermore, our results indicate that LLM-based systems can perform on par with state-of-the-art clinical coding systems while offering superior generalizability across institutions, languages, and ICD versions. Shaoxin Li 0004, Jiaxiang Wu 0002, Qinwei Xu, Xingkun Xu, Yingkai Sun, Zhian Bai, Yuchen Xu 0008, Lifeng Zhu, Weiguo Hu, Feiyue Huang |
IEEE J. Biomed. Health Informatics | 12 |
| 2024 | Hallucinated Style Distillation for Single Domain Generalization in Medical Image Segmentation
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Shaoxin Li 0001, Yuexiang Li, Yefeng Zheng 0001, Feiyue Huang |
MICCAI (10) | 10 |
| 2023 | Knowing What it is: Semantic-Enhanced Dual Attention TransformerabstractAttention has become an indispensable component of the models of various multimedia tasks likeImage Captioning(IC) andVisual Question Answering(VQA). However, most existing attention modules are designed for capturing the spatial dependency, and are still insufficient in semantic understanding,e.g., the categories of objects and their attributes, which is also critical for image captioning. To compensate for this defect, we propose a novel attention module termedChannel-wise Attention Block(CAB) to model channel-wise dependency for both visual modality and linguistic modality, thereby improving semantic learning and multi-modal reasoning simultaneously. Specifically, CAB has two novel designs to tackle with the high overhead of channel-wise attention, which are thereduction-reconstruction block structureand thegating-based attention prediction. Based on CAB, we further propose a novelSemantic-enhanced Dual Attention Transformer(termed SDATR), which combines the merits of spatial and channel-wise attentions. To validate SDATR, we conduct extensive experiments on the MS COCO dataset and yield new state-of-the-art performance of 134.5 CIDEr score on COCO Karpathy test split and 136.0 CIDEr score on the official online testing server. To examine the generalization of SDATR, we also apply it to the task of visual question answering, where superior performance gains are also witnessed. The code and models are publicly available athttps://github.com/xmu-xiaoma666/SDATR. Jiayi Ji, Xiaoshuai Sun, Yiyi Zhou, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
IEEE Trans. Multim. | 6 |
| 2023 | Fast Monocular Depth Estimation via Side Prediction Aggregation with Continuous Spatial RefinementabstractRecent works have validated the benefit of integrating spatial information into deep networks to improve pixel-level prediction tasks such as monocular depth estimation. However, how to efficiently and robustly integrate spatial cues retains as an open problem. In this paper, we introduce the Side Prediction Aggregation (termed SPA) method to enhance the embedding of scene structural information from low-level to high-level layers. To improve the estimation accuracy, the proposed method is further equipped with continuous Spatial Refinement Loss (termed SRL) at multiple resolutions with negligible extra computation. Besides, the proposed sequential network can further perform adversarial learning at multiple resolutions. Such an adversarial refinement strategy greatly improves the accuracy of estimated depth with a little extra computation. Without using any pre-trained models, our network achieves the the-state-of-art accuracy on KITTI, NYUD V2, and Cityscapes datasets, which has achieved real-time depth estimation online. Jipeng Wu, Rongrong Ji, Qiang Wang 0023, Shengchuan Zhang, Xiaoshuai Sun, Yan Wang 0059, Mingliang Xu 0001, Feiyue Huang |
IEEE Trans. Multim. | 8 |
| 2023 | Balance-Aware Grid Collage for Small Image CollectionsabstractGrid collages (GClg) of small image collections are popular and useful in many applications, such as personal album management, online photo posting, and graphic design. In this article, we focus on how visual effects influence individual preferences through various arrangements of multiple images under such scenarios. A novel balance-aware metric is proposed to bridge the gap between multi-image joint presentation and visual pleasure. The metric merges psychological achievements into the field of grid collage. To capture user preference, a bonus mechanism related to a user-specified special location in the grid and uniqueness values of the subimages is integrated into the metric. An end-to-end reinforcement learning mechanism empowers the model without tedious manual annotations. Experiments demonstrate that our metric can evaluate the GClg visual balance in line with human subjective perception, and the model can generate visually pleasant GClg results, which is comparable to manual designs. Fan Tang, Weiming Dong, Feiyue Huang, Tong-Yee Lee, Changsheng Xu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Lifelong Person Re-identification by Pseudo Task Knowledge PreservationabstractIn real world, training data for person re-identification (Re-ID) is collected discretely with spatial and temporal variations, which requires a model to incrementally learn new knowledge without forgetting old knowledge. This problem is called lifelong person re-identification (LReID). Variations of illumination and background for images of each task exhibit task-specific image style and lead to task-wise domain gap. In addition to missing data from the old tasks, task-wise domain gap is a key factor for catastrophic forgetting in LReID, which is ignored in existing approaches for LReID. The model tends to learn task-specific knowledge with task-wise domain gap, which results in stability and plasticity dilemma. To overcome this problem, we cast LReID as a domain adaptation problem and propose a pseudo task knowledge preservation framework to alleviate the domain gap. Our framework is based on a pseudo task transformation module which maps the features of the new task into the feature space of the old tasks to complement the limited saved exemplars of the old tasks. With extra transformed features in the task-specific feature space, we propose a task-specific domain consistency loss to implicitly alleviate the task-wise domain gap for learning task-shared knowledge instead of task-specific one. Furthermore, to guide knowledge preservation with the feature distributions of the old tasks, we propose to preserve knowledge on extra pseudo tasks which jointly distills knowledge and discriminates identity, in order to achieve a better trade-off between stability and plasticity for lifelong learning with task-wise domain gap. Extensive experiments demonstrate the superiority of our method as compared with the state-of-the-art lifelong learning and LReID methods. Wenhang Ge, Junlong Du, Ancong Wu, Yuqiao Xian, Feiyue Huang, Wei-Shi Zheng 0001 |
AAAI | 6 |
| 2022 | Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack
Yixu Wang, Jie Li 0052, Hong Liu 0009, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
ECCV (5) | 6 |
| 2022 | Deepwalk-aware graph convolutional networks
Taisong Jin, Huaqiang Dai, Liujuan Cao, Baochang Zhang 0001, Feiyue Huang, Yue Gao 0002, Rongrong Ji |
Sci. China Inf. Sci. | 5 |
| 2022 | Transformers in computational visual media: A surveyabstractTransformers, the dominant architecture for natural language processing, have also recently attracted much attention from computational visual media researchers due to their capacity for long-range representation and high performance. Transformers are sequence-to-sequence models, which use a self-attention mechanism rather than the RNN sequential structure. Thus, such models can be trained in parallel and can represent global information. This study comprehensively surveys recent visual transformer works. We categorize them according to task scenario: backbone design, high-level vision, low-level vision and generation, and multimodal learning. Their key ideas are also analyzed. Differing from previous surveys, we mainly focus on visual transformer methods in low-level vision and generation. The latest works on backbone design are also reviewed in detail. For ease of understanding, we precisely describe the main contributions of the latest works in the form of tables. As well as giving quantitative comparisons, we also present image results for low-level vision and generation tasks. Computational costs and source code links for various important works are also given in this survey to assist further development. Yifan Xu 0008, HuaPeng Wei, Minxuan Lin, Yingying Deng, Kekai Sheng, Mengdan Zhang, Fan Tang, Weiming Dong, Feiyue Huang, Changsheng Xu |
Comput. Vis. Media | 9 |
| 2022 | Fast Class-Wise Updating for Online HashingabstractOnline image hashing has received increasing research attention recently, which processes large-scale data in a streaming fashion to update the hash functions on-the-fly. To this end, most existing works exploit this problem under a supervised setting, i.e., using class labels to boost the hashing performance, which suffers from the defects in both adaptivity and efficiency: First, large amounts of training batches are required to learn up-to-date hash functions, which leads to poor online adaptivity. Second, the training is time-consuming, which contradicts with the core need of online learning. In this paper, a novel supervised online hashing scheme, termed Fast Class-wise Updating for Online Hashing (FCOH), is proposed to address the above two challenges by introducing a novel and efficient inner product operation. To achieve fast online adaptivity, a class-wise updating method is developed to decompose the binary code learning and alternatively renew the hash functions in a class-wise fashion, which well addresses the burden on large amounts of training batches. Quantitatively, such a decomposition further leads to at least 75 percent storage saving. To further achieve online efficiency, we propose a semi-relaxation optimization, which accelerates the online training by treating different binary constraints independently. Without additional constraints and variables, the time complexity is significantly reduced. Such a scheme is also quantitatively shown to well preserve past information during updating hashing functions. We have quantitatively demonstrated that the collective effort of class-wise updating and semi-relaxation optimization provides a superior performance comparing to various state-of-the-art methods, which is verified through extensive experiments on three widely-used datasets. Mingbao Lin, Rongrong Ji, Xiaoshuai Sun, Baochang Zhang 0001, Feiyue Huang, Yonghong Tian 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Towards Lightweight Transformer Via Group-Wise Transformation for Vision-and-Language TasksabstractDespite the exciting performance, Transformer is criticized for its excessive parameters and computation cost. However, compressing Transformer remains as an open problem due to its internal complexity of the layer designs, i.e., Multi-Head Attention (MHA) and Feed-Forward Network (FFN). To address this issue, we introduce Group-wise Transformation towards a universal yet lightweight Transformer for vision-and-language tasks, termed as LW-Transformer. LW-Transformer applies Group-wise Transformation to reduce both the parameters and computations of Transformer, while also preserving its two main properties, i.e., the efficient attention modeling on diverse subspaces of MHA, and the expanding-scaling feature transformation of FFN. We apply LW-Transformer to a set of Transformer-based networks, and quantitatively measure them on three vision-and-language tasks and six benchmark datasets. Experimental results show that while saving a large number of parameters and computations, LW-Transformer achieves very competitive performance against the original Transformer networks for vision-and-language tasks. To examine the generalization ability, we apply LW-Transformer to the task of image classification, and build its network based on a recently proposed image Transformer called Swin-Transformer, where the effectiveness can be also confirmed. Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Yan Wang 0059, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
IEEE Trans. Image Process. | 7 |
| 2022 | Knowledge-Driven Generative Adversarial Network for Text-to-Image SynthesisabstractText-to-Image (T2I) synthesis is a challenging task that aims to convert natural language descriptions to real images. It remains an open problem mainly due to the diversity of text descriptions, which poses a huge obstacle in generating vivid and relevant images. Moreover, the existing evaluation metrics in T2I synthesis are mainly used to evaluate the visual quality of the generated images, while the semantic consistency between the two modalities is often ignored. To address these issues, we present a novelKnowledge-Driven Generative Adversarial Network, termed KD-GAN, and a new evaluation system, namedPseudo Turing Test(PTT for short). Concretely, KD-GAN takes a further step in imitating the behavior of human painting,i.e., drawing an image according to reference knowledge. The introduction of reference knowledge in KD-GAN not only improves the quality of the generated images but also enhances the semantic consistency between them and the input texts. In addition, KD-GAN can also greatly avoid some flaws against common sense during image generation,e.g., skiing in the blue sky. The proposed PTT is an important supplement to the existing evaluation system of T2I synthesis. It includes a set of pseudo-experts of different multimedia tasks to evaluate the semantic consistency between the given texts and the generated images. To validate the proposed KD-GAN, we conducted extensive experiments on two benchmark datasets,i.e., Caltech-UCSD Birds (CUB), and MS-COCO (COCO). The experimental results demonstrate that KD-GAN outperforms state-of-the-art methods on IS, FID, and the proposed PTT metrics.11The codes of KD-GAN are at [Online]. Available:https://github.com/pengjunn/KD-GANand the codes and models of PTT are at [Online]. Available:https://github.com/pengjunn/PTT. Jun Peng 0007, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
IEEE Trans. Multim. | 6 |
| 2022 | Network Pruning Using Adaptive Exemplar FiltersabstractPopular network pruning algorithms reduce redundant information by optimizing hand-crafted models, and may cause suboptimal performance and long time in selecting filters. We innovatively introduce adaptive exemplar filters to simplify the algorithm design, resulting in an automatic and efficient pruning approach called EPruner. Inspired by the face recognition community, we use a message-passing algorithm Affinity Propagation on the weight matrices to obtain an adaptive number of exemplars, which then act as the preserved filters. EPruner breaks the dependence on the training data in determining the "important" filters and allows the CPU implementation in seconds, an order of magnitude faster than GPU-based SOTAs. Moreover, we show that the weights of exemplars provide a better initialization for the fine-tuning. On VGGNet-16, EPruner achieves a 76.34%-FLOPs reduction by removing 88.80% parameters, with 0.06% accuracy improvement on CIFAR-10. In ResNet-152, EPruner achieves a 65.12%-FLOPs reduction by removing 64.18% parameters, with only 0.71% top-5 accuracy loss on ILSVRC-2012. Our code is available at https://github.com/lmbxmu/EPruner. Mingbao Lin, Rongrong Ji, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Generalizable Representation Learning for Mixture Domain Face Anti-SpoofingabstractFace anti-spoofing approach based on domain generalization (DG) has drawn growing attention due to its robustness for unseen scenarios. Existing DG methods assume that the domain label is known. However, in real-world applications, the collected dataset always contains mixture domains, where the domain label is unknown. In this case, most of existing methods may not work. Further, even if we can obtain the domain label as existing methods, we think this is just a sub-optimal partition. To overcome the limitation, we propose domain dynamic adjustment meta-learning (D$^2$AM) without using domain labels, which iteratively divides mixture domains via discriminative domain representation and trains a generalizable face anti-spoofing with meta-learning. Specifically, we design a domain feature based on Instance Normalization (IN) and propose a domain representation learning module (DRLM) to extract discriminative domain features for clustering. Moreover, to reduce the side effect of outliers on clustering performance, we additionally utilize maximum mean discrepancy (MMD) to align the distribution of sample features to a prior distribution, which improves the reliability of clustering. Extensive experiments show that the proposed method outperforms conventional DG-based face anti-spoofing methods, including those utilizing domain labels. Furthermore, we enhance the interpretability through visualization. Taiping Yao, Kekai Sheng, Shouhong Ding, Ying Tai, Feiyue Huang |
AAAI | 7 |
| 2021 | Frequency Consistent Adaptation for Real World Super ResolutionabstractRecent deep-learning based Super-Resolution (SR) methods have achieved remarkable performance on images with known degradation. However, these methods always fail in real-world scene, since the Low-Resolution (LR) images after the ideal degradation (e.g., bicubic down-sampling) deviate from real source domain. The domain gap between the LR images and the real-world images can be observed clearly on frequency density, which inspires us to explicitly narrow the undesired gap caused by incorrect degradation. From this point of view, we design a novel Frequency Consistent Adaptation (FCA) that ensures the frequency domain consistency when applying existing SR methods to the real scene. We estimate degradation kernels from unsupervised images and generate the corresponding LR images. To provide useful gradient information for kernel estimation, we propose Frequency Density Comparator (FDC) by distinguishing the frequency density of images on different scales. Based on the domain-consistent LR-HR pairs, we train easy-implemented Convolutional Neural Network (CNN) SR models. Extensive experiments show that the proposed FCA improves the performance of the SR model under real-world setting achieving state-of-the-art results with high fidelity and plausible perception, thus providing a novel effective framework for real-world SR application. Xiaozhong Ji, Guangpin Tao, Yun Cao 0002, Ying Tai, Tong Lu 0002, Chengjie Wang 0001, Feiyue Huang |
AAAI | 8 |
| 2021 | Dual-level Collaborative Transformer for Image CaptioningabstractDescriptive region features extracted by object detection networks have played an important role in the recent advancements of image captioning. However, they are still criticized for the lack of contextual information and fine-grained details, which in contrast are the merits of traditional grid features. In this paper, we introduce a novel Dual-Level Collaborative Transformer (DLCT) network to realize the complementary advantages of the two features. Concretely, in DLCT, these two features are first processed by a novel Dual-way Self Attenion (DWSA) to mine their intrinsic properties, where a Comprehensive Relation Attention component is also introduced to embed the geometric information. In addition, we propose a Locality-Constrained Cross Attention module to address the semantic noises caused by the direct fusion of these two features, where a geometric alignment graph is constructed to accurately align and reinforce region and grid features. To validate our model, we conduct extensive experiments on the highly competitive MS-COCO dataset, and achieve new state-of-the-art performance on both local and online test sets, i.e., 133.8% CIDEr on Karpathy split and 135.4% CIDEr on the official split. Yunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang, Chia-Wen Lin, Rongrong Ji |
AAAI | 6 |
| 2021 | Learning Comprehensive Motion Representation for Action RecognitionabstractFor action recognition learning, 2D CNN-based methods are efficient but may yield redundant features due to applying the same 2D convolution kernel to each frame. Recent efforts attempt to capture motion information by establishing inter-frame connections while still suffering the limited temporal receptive field or high latency. Moreover, the feature enhancement is often only performed by channel or space dimension in action recognition. To address these issues, we first devise a Channel-wise Motion Enhancement (CME) module to adaptively emphasize the channels related to dynamic information with a channel-wise gate vector. The channel gates generated by CME incorporate the information from all the other frames in the video. We further propose a Spatial-wise Motion Enhancement (SME) module to focus on the regions with the critical target in motion, according to the point-to-point similarity between adjacent feature maps. The intuition is that the change of background is typically slower than the motion area. Both CME and SME have clear physical meaning in capturing action clues. By integrating the two modules into the off-the-shelf 2D network, we finally obtain a Comprehensive Motion Representation (CMR) learning method for action recognition, which achieves competitive performance on Something-Something V1 & V2 and Kinetics-400. On the temporal reasoning datasets Something-Something V1 and V2, our method outperforms the current state-of-the-art by 2.3% and 1.9% when using 16 frames as input, respectively. Mingyu Wu 0002, Boyuan Jiang, Donghao Luo 0001, Junchi Yan, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Xiaokang Yang 0001 |
AAAI | 9 |
| 2021 | Learning Dynamic Alignment via Meta-Filter for Few-Shot LearningabstractFew-shot learning (FSL), which aims to recognise new classes by adapting the learned knowledge with extremely limited few-shot (support) examples, remains an important open problem in computer vision. Most of the existing methods for feature alignment in few-shot learning only consider image-level or spatial-level alignment while omitting the channel disparity. Our insight is that these methods would lead to poor adaptation with redundant matching, and leveraging channel-wise adjustment is the key to well adapting the learned knowledge to new classes. Therefore, in this paper, we propose to learn a dynamic alignment, which can effectively highlight both query regions and channels according to different local support information. Specifically, this is achieved by first dynamically sampling the neighbourhood of the feature position conditioned on the input few shot, based on which we further predict a both position-dependent and channel-dependent Dynamic Meta-filter. The filter is used to align the query feature with position-specific and channel-specific knowledge. Moreover, we adopt Neural Ordinary Differential Equation (ODE) to enable a more accurate control of the alignment. In such a sense our model is able to better capture fine-grained semantic context of the few-shot example and thus facilitates dynamical knowledge adaptation for few-shot learning. The resulting framework establishes the new state-of-the-arts on major few-shot visual recognition benchmarks, including miniImageNet and tieredImageNet. Chengming Xu 0001, Yanwei Fu 0001, Chen Liu 0030, Chengjie Wang 0001, Feiyue Huang, Li Zhang 0040, Xiangyang Xue 0001 |
CVPR | 6 |
| 2021 | Learning To Aggregate and Personalize 3D Face From In-the-Wild Photo CollectionabstractNon-parametric face modeling aims to reconstruct 3D face only from images without shape assumptions. While plausible facial details are predicted, the models tend to over-depend on local color appearance and suffer from ambiguous noise. To address such problem, this paper presents a novel Learning to Aggregate and Personalize (LAP) framework for unsupervised robust 3D face modeling. Instead of using controlled environment, the proposed method implicitly disentangles ID-consistent and scene-specific face from unconstrained photo set. Specifically, to learn ID-consistent face, LAP adaptively aggregates intrinsic face factors of an identity based on a novel curriculum learning approach with relaxed consistency loss. To adapt the face for a personalized scene, we propose a novel attribute-refining network to modify ID-consistent face with target attribute and details. Based on the proposed method, we make unsupervised 3D face modeling benefit from meaningful image facial structure and possibly higher resolutions. Extensive experiments on benchmarks show LAP recovers superior or competitive face shape and texture, compared with state-of-the-art (SOTA) methods with or without prior and supervision. Zhenyu Zhang 0005, Yanhao Ge, Renwang Chen, Ying Tai, Yan Yan 0002, Jian Yang 0003, Chengjie Wang 0001, Feiyue Huang |
CVPR | 9 |
| 2021 | Image-to-Image Translation via Hierarchical Style DisentanglementabstractRecently, image-to-image translation has made significant progress in achieving both multi-label (i.e., translation conditioned on different labels) and multi-style (i.e., generation with diverse styles) tasks. However, due to the unexplored independence and exclusiveness in the labels, existing endeavors are defeated by involving uncontrolled manipulations to the translation results. In this paper, we propose Hierarchical Style Disentanglement (HiSD) to address this issue. Specifically, we organize the labels into a hierarchical tree structure, in which independent tags, exclusive attributes, and disentangled styles are allocated from top to bottom. Correspondingly, a new translation process is designed to adapt the above structure, in which the styles are identified for controllable translations. Both qualitative and quantitative results on the CelebA-HQ dataset verify the ability of the proposed HiSD. The code has been released at https://github.com/imlixinyang/HiSD. Shengchuan Zhang, Jie Hu 0018, Liujuan Cao, Xiaopeng Hong, Xudong Mao, Feiyue Huang, Yongjian Wu 0001, Rongrong Ji |
CVPR | 7 |
| 2021 | Learning Salient Boundary Feature for Anchor-free Temporal Action LocalizationabstractTemporal action localization is an important yet challenging task in video understanding. Typically, such a task aims at inferring both the action category and localization of the start and end frame for each action instance in a long, untrimmed video. While most current models achieve good results by using pre-defined anchors and numerous actionness, such methods could be bothered with both large number of outputs and heavy tuning of locations and sizes corresponding to different anchors. Instead, anchor-free methods is lighter, getting rid of redundant hyper-parameters, but gains few attention. In this paper, we propose the first purely anchor-free temporal localization method, which is both efficient and effective. Our model includes (i) an end-to-end trainable basic predictor, (ii) a saliency-based refinement module to gather more valuable boundary features for each proposal with a novel boundary pooling, and (iii) several consistency constraints to make sure our model can find the accurate boundary given arbitrary proposals. Extensive experiments show that our method beats all anchor-based and actionness-guided methods with a remarkable margin on THUMOS14, achieving state-of-the-art results, and comparable ones on ActivityNet v1.3. Code is available at https://github.com/TencentYoutuResearch/ActionDetection-AFSD. Chuming Lin, Chengming Xu 0001, Donghao Luo 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Yanwei Fu 0001 |
CVPR | 8 |
| 2021 | Unveiling the Potential of Structure Preserving for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) remains an open problem given the deficiency of finding object extent information using a classification network. Although prior works struggled to localize objects through various spatial regularization strategies, we argue that how to extract object structural information from the trained classification network is neglected. In this paper, we propose a two-stage approach, termed structure-preserving activation (SPA), toward fully leveraging the structure information incorporated in convolutional features for WSOL. First, a restricted activation module (RAM) is designed to alleviate the structure-missing issue caused by the classification network on the basis of the observation that the unbounded classification map and global average pooling layer drive the network to focus only on object parts. Second, we designed a post-process approach, termed self-correlation map generating (SCG) module to obtain structure-preserving localization maps on the basis of the activation maps acquired from the first stage. Specifically, we utilize the high-order self-correlation (HSC) to extract the inherent structural information retained in the learned model and then aggregate HSC of multiple points for precise object localization. Extensive experiments on two publicly available benchmarks including CUB-2002011 and ILSVRC show that the proposed SPA achieves substantial and consistent performance gains compared with baseline approaches. Code and models are available at github.com/Panxjia/SPA CVPR2021. Xingjia Pan, Yingguo Gao, Zhiwen Lin, Fan Tang, Weiming Dong, Haolei Yuan, Feiyue Huang, Changsheng Xu |
CVPR | 7 |
| 2021 | Toward Joint Thing-and-Stuff Mining for Weakly Supervised Panoptic SegmentationabstractPanoptic segmentation aims to partition an image to object instances and semantic content for thing and stuff categories, respectively. To date, learning weakly supervised panoptic segmentation (WSPS) with only image-level labels remains unexplored. In this paper, we propose an efficient jointly thing-and-stuff mining (JTSM) framework for WSPS. To this end, we design a novel mask of interest pooling (MoIPool) to extract fixed-size pixel-accurate feature maps of arbitrary-shape segmentations. MoIPool enables a panoptic mining branch to leverage multiple instance learning (MIL) to recognize things and stuff segmentation in a unified manner. We further refine segmentation masks with parallel instance and semantic segmentation branches via self-training, which collaborates the mined masks from panoptic mining with bottom-up object evidence as pseudo-ground-truth labels to improve spatial coherence and contour localization. Experimental results demonstrate the effectiveness of JTSM on PASCAL VOC and MS COCO. As a by-product, we achieve competitive results for weakly supervised object detection and instance segmentation. This work is a first step towards tackling challenge panoptic segmentation task with only image-level labels. Yunhang Shen, Liujuan Cao, Feihong Lian, Baochang Zhang 0001, Chi Su, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
CVPR | 8 |
| 2021 | Removing the Background by Adding the Background: Towards Background Robust Self-Supervised Video Representation LearningabstractSelf-supervised learning has shown great potentials in improving the video representation ability of deep neural networks by getting supervision from the data itself. However, some of the current methods tend to cheat from the background, i.e., the prediction is highly dependent on the video background instead of the motion, making the model vulnerable to background changes. To mitigate the model reliance towards the background, we propose to remove the background impact by adding the background. That is, given a video, we randomly select a static frame and add it to every other frames to construct a distracting video sample. Then we force the model to pull the feature of the distracting video and the feature of the original video closer, so that the model is explicitly restricted to resist the background influence, focusing more on the motion changes. We term our method as Background Erasing (BE). It is worth noting that the implementation of our method is so simple and neat and can be added to most of the SOTA methods without much efforts. Specifically, BE brings 16.4% and 19.1% improvements with MoCo on the severely biased datasets UCF101 and HMDB51, and 14.5% improvement on the less biased dataset Diving48. Ke Li 0015, Andy Jinhua Ma, Hao Cheng 0012, Feiyue Huang, Rongrong Ji, Xing Sun 0001 |
CVPR | 8 |
| 2021 | Delving into Data: Effectively Substitute Training for Black-box AttackabstractDeep models have shown their vulnerability when processing adversarial samples. As for the black-box attack, without access to the architecture and weights of the attacked model, training a substitute model for adversarial attacks has attracted wide attention. Previous substitute training approaches focus on stealing the knowledge of the target model based on real training data or synthetic data, without exploring what kind of data can further improve the transferability between the substitute and target models. In this paper, we propose a novel perspective substitute training that focuses on designing the distribution of data used in the knowledge stealing process. More specifically, a diverse data generation module is proposed to synthesize large-scale data with wide distribution. And adversarial substitute training strategy is introduced to focus on the data distributed near the decision boundary. The combination of these two modules can further boost the consistency of the substitute model and target model, which greatly improves the effectiveness of adversarial attack. Extensive experiments demonstrate the efficacy of our method against state-of-the-art competitors under non-target and target at-tack settings. Detailed visualization and analysis are also provided to help understand the advantage of our method. Wenxuan Wang 0003, Bangjie Yin, Taiping Yao, Li Zhang 0040, Yanwei Fu 0001, Shouhong Ding, Feiyue Huang, Xiangyang Xue 0001 |
CVPR | 8 |
| 2021 | Discover Cross-Modality Nuances for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (Re-ID) aims to match the pedestrian images of the same identity from different modalities. Existing works mainly focus on alleviating the modality discrepancy by aligning the distributions of features from different modalities. However, nuanced but discriminative information, such as glasses, shoes, and the length of clothes, has not been fully explored, especially in the infrared modality. Without discovering nuances, it is challenging to match pedestrians across modalities using modality alignment solely, which inevitably reduces feature distinctiveness. In this paper, we propose a joint Modality and Pattern Alignment Network (MPANet) to discover cross-modality nuances in different patterns for visible-infrared person Re-ID, which introduces a modality alleviation module and a pattern alignment module to jointly extract discriminative features. Specifically, we first propose a modality alleviation module to dislodge the modality information from the extracted feature maps. Then, We devise a pattern alignment module, which generates multiple pattern maps for the diverse patterns of a person, to discover nuances. Finally, we introduce a mutual mean learning fashion to alleviate the modality discrepancy and propose a center cluster loss to guide both identity learning and nuances discovering. Extensive experiments on the public SYSU-MM01 and RegDB datasets demonstrate the superiority of MPANet over state-of-the-arts. Qiong Wu 0012, Pingyang Dai, Jie Chen 0001, Chia-Wen Lin, Yongjian Wu 0001, Feiyue Huang, Bineng Zhong 0001, Rongrong Ji |
CVPR | 6 |
| 2021 | Consistent Instance False Positive Improves Fairness in Face RecognitionabstractDemographic bias is a significant challenge in practical face recognition systems. Existing methods heavily rely on accurate demographic annotations. However, such annotations are usually unavailable in real scenarios. Moreover, these methods are typically designed for a specific demographic group and are not general enough. In this paper, we propose a false positive rate penalty loss, which mitigates face recognition bias by increasing the consistency of instance False Positive Rate (FPR). Specifically, we first define the instance FPR as the ratio between the number of the non-target similarities above a unified threshold and the total number of the non-target similarities. The unified threshold is estimated for a given total FPR. Then, an additional penalty term, which is in proportion to the ratio of instance FPR overall FPR, is introduced into the denominator of the softmax-based loss. The larger the instance FPR, the larger the penalty. By such unequal penalties, the instance FPRs are supposed to be consistent. Compared with the previous debiasing methods, our method requires no demographic annotations. Thus, it can mitigate the bias among demographic groups divided by various attributes, and these attributes are not needed to be previously predefined during training. Extensive experimental results on popular benchmarks demonstrate the superiority of our method over state-of-the-art competitors. Code and pre-trained models are available at https://github.com/xkx0430/FairnessFR. Xingkun Xu, Yuge Huang, Pengcheng Shen, Shaoxin Li 0001, Feiyue Huang, Yong Li 0044, Zhen Cui 0001 |
CVPR | 6 |
| 2021 | Learning To Restore Hazy Video: A New Real-World Dataset and a New MethodabstractMost of the existing deep learning-based dehazing methods are trained and evaluated on the image dehazing datasets, where the dehazed images are generated by only exploiting the information from the corresponding hazy ones. On the other hand, video dehazing algorithms, which can acquire more satisfying dehazing results by exploiting the temporal redundancy from neighborhood hazy frames, receive less attention due to the absence of the video dehazing datasets. Therefore, we propose the first REal-world VIdeo DEhazing (REVIDE) dataset which can be used for the supervised learning of the video dehazing algorithms. By utilizing a well-designed video acquisition system, we can capture paired real-world hazy and haze-free videos that are perfectly aligned by recording the same scene (with or without haze) twice. Considering the challenge of exploiting temporal redundancy among the hazy frames, we also develop a Confidence Guided and Improved Deformable Network (CG-IDN) for video dehazing. The experiments demonstrate that the hazy scenes in the REVIDE dataset are more realistic than the synthetic datasets and the proposed algorithm also performs favorably against state-of-the-art dehazing methods. Xinyi Zhang 0005, Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Fei Wang 0008 |
CVPR | 8 |
| 2021 | RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual WordsabstractRecent progress on visual question answering has explored the merits of grid features for vision language tasks. Meanwhile, transformer-based models have shown remarkable performance in various sequence prediction problems. However, the spatial information loss of grid features caused by flattening operation, as well as the defect of the transformer model in distinguishing visual words and non visual words, are still left unexplored. In this paper, we first propose Grid-Augmented (GA) module, in which relative geometry features between grids are incorporated to enhance visual representations. Then, we build a BERT-based language model to extract language context and propose Adaptive-Attention (AA) module on top of a transformer decoder to adaptively measure the contribution of visual and language cues before making decisions for word prediction. To prove the generality of our proposals, we apply the two modules to the vanilla transformer model to build our Relationship-Sensitive Transformer (RSTNet) for image captioning task. The proposed model is tested on the MSCOCO benchmark, where it achieves new state-of-art results on both the Karpathy test split and the online test server. Source code is available at GitHub1. Xuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji, Yiyi Zhou, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
CVPR | 7 |
| 2021 | Graph-Based High-Order Relation Discovery for Fine-Grained RecognitionabstractFine-grained object recognition aims to learn effective features that can identify the subtle differences between visually similar objects. Most of the existing works tend to amplify discriminative part regions with attention mechanisms. Besides its unstable performance under complex backgrounds, the intrinsic interrelationship between different semantic features is less explored. Toward this end, we propose an effective graph-based relation discovery approach to build a contextual understanding of high-order relationships. In our approach, a high-dimensional feature bank is first formed and jointly regularized with semantic- and positional-aware high-order constraints, endowing rich attributes to feature representations. Second, to overcome the high-dimension curse, we propose a graph-based semantic grouping strategy to embed this high-order tensor bank into a low-dimensional space. Meanwhile, a group-wise learning strategy is proposed to regularize the features focusing on the cluster embedding center. With the collaborative learning of three modules, our module is able to grasp the stronger contextual details of fine-grained objects. Experimental evidence demonstrates our approach achieves new state-of-the-art on 4 widely-used fine-grained object recognition benchmarks. Yifan Zhao 0002, Ke Yan 0007, Feiyue Huang, Jia Li 0003 |
CVPR | 3 |
| 2021 | Architecture Disentanglement for Deep Neural NetworksabstractUnderstanding the inner workings of deep neural networks (DNNs) is essential to provide trustworthy artificial intelligence techniques for practical applications. Existing studies typically involve linking semantic concepts to units or layers of DNNs, but fail to explain the inference process. In this paper, we introduce neural architecture disentanglement (NAD) to fill the gap. Specifically, NAD learns to disentangle a pre-trained DNN into sub-architectures according to independent tasks, forming information flows that describe the inference processes. We investigate whether, where, and how the disentanglement occurs through experiments conducted with handcrafted and automatically-searched network architectures, on both object-based and scene-based datasets. Based on the experimental results, we present three new findings that provide fresh insights into the inner logic of DNNs. First, DNNs can be divided into sub-architectures for independent tasks. Second, deeper layers do not always correspond to higher semantics. Third, the connection type in a DNN affects how the information flows across layers, leading to different disentanglement behaviors. With NAD, we further explain why DNNs sometimes give wrong predictions. Experimental results show that misclassified images have a high probability of being assigned to task sub-architectures similar to the correct ones. Our code is available at https://github.com/hujiecpp/NAD. Jie Hu 0018, Liujuan Cao, Qixiang Ye, Shengchuan Zhang, Ke Li 0015, Feiyue Huang, Ling Shao 0001, Rongrong Ji |
ICCV | 7 |
| 2021 | Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial QueryabstractText-based image retrieval has seen considerable progress in recent years. However, the performance of existing methods suffers in real life since the user is likely to provide an incomplete description of an image, which often leads to results filled with false positives that fit the incomplete description. In this work, we introduce the partial-query problem and extensively analyze its influence on text-based image retrieval. Previous interactive methods tackle the problem by passively receiving users’ feedback to supplement the incomplete query iteratively, which is time-consuming and requires heavy user effort. Instead, we propose a novel retrieval framework that conducts the interactive process in an Ask-and-Confirm fashion, where AI actively searches for discriminative details missing in the current query, and users only need to confirm AI’s proposal. Specifically, we propose an object-based interaction to make the interactive retrieval more user-friendly and present a reinforcement-learning-based policy to search for discriminative objects. Furthermore, since fully-supervised training is often infeasible due to the difficulty of obtaining human-machine dialog data, we present a weakly-supervised training strategy that needs no human-annotated dialogs other than a text-image dataset. Experiments show that our framework significantly improves the performance of text-based image retrieval. Code is available at https://github.com/CuthbertCai/Ask-Confirm. Guanyu Cai, Jun Zhang 0018, Xinyang Jiang, Yifei Gong, Lianghua He, Fufu Yu, Feiyue Huang, Xing Sun 0001 |
ICCV | 9 |
| 2021 | Aha! Adaptive History-driven Attack for Decision-based Black-box ModelsabstractThe decision-based black-box attack means to craft adversarial examples with only the top-1 label of the victim model available. A common practice is to start from a large perturbation and then iteratively reduce it with a deterministic direction and a random one while keeping it adversarial. The limited information obtained from each query and inefficient direction sampling impede attack efficiency, making it hard to obtain a small enough perturbation within a limited number of queries. To tackle this problem, we propose a novel attack method termed Adaptive History-driven Attack (AHA) which gathers information from all historical queries as the prior for current sampling. Moreover, to balance between the deterministic direction and the random one, we dynamically adjust the coefficient according to the ratio of the actual magnitude reduction to the expected one. Such a strategy improves the success rate of queries during optimization, letting adversarial examples move swiftly along the decision boundary. Our method can also integrate with subspace optimization like dimension reduction to further improve efficiency. Extensive experiments on both ImageNet and CelebA datasets demonstrate that our method achieves at least 24.3% lower magnitude of perturbation on average with the same number of queries. Finally, we prove the practical potential of our method by evaluating it on popular defense methods and a real-world system provided by MEGVII Face++. Jie Li 0052, Rongrong Ji, Peixian Chen, Baochang Zhang 0001, Xiaopeng Hong, Shaoxin Li 0001, Feiyue Huang, Yongjian Wu 0001 |
ICCV | 9 |
| 2021 | Parallel Detection-and-Segmentation Learning for Weakly Supervised Instance SegmentationabstractWeakly supervised instance segmentation (WSIS) with only image-level labels has recently drawn much attention. To date, bottom-up WSIS methods refine discriminative cues from classifiers with sophisticated multi-stage training procedures, which also suffer from inconsistent object boundaries. And top-down WSIS methods are formulated as cascade detection-to-segmentation pipeline, in which the quality of segmentation learning heavily depends on pseudo masks generated from detectors. In this paper, we propose a unified parallel detection-and-segmentation learning (PDSL) framework to learn instance segmentation with only image-level labels, which draws inspiration from both top-down and bottom-up instance segmentation approaches. The detection module is the same as the typical design of any weakly supervised object detection, while the segmentation module leverages self-supervised learning to model class-agnostic foreground extraction, following by self-training to refine class-specific segmentation. We further design instance-activation correlation module to improve the coherence between detection and segmentation branches. Extensive experiments verify that the proposed method outperforms baselines and achieves the state-of-the-art results on PASCAL VOC and MS COCO. Yunhang Shen, Liujuan Cao, Baochang Zhang 0001, Chi Su, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
ICCV | 7 |
| 2021 | Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkabstractLocalizing individuals in crowds is more in accordance with the practical demands of subsequent high-level crowd analysis tasks than simply counting. However, existing localization based methods relying on intermediate representations (i.e., density maps or pseudo boxes) serving as learning targets are counter-intuitive and error-prone. In this paper, we propose a purely point-based framework for joint crowd counting and individual localization. For this framework, instead of merely reporting the absolute counting error at image level, we propose a new metric, called density Normalized Average Precision (nAP), to provide more comprehensive and more precise performance evaluation. Moreover, we design an intuitive solution under this framework, which is called Point to Point Network (P2PNet). P2PNet discards superfluous steps and directly predicts a set of point proposals to represent heads in an image, being consistent with the human annotation results. By thorough analysis, we reveal the key step towards implementing such a novel idea is to assign optimal learning targets for these proposals. Therefore, we propose to conduct this crucial association in an one-to-one matching manner using the Hungarian algorithm. The P2PNet not only significantly surpasses state-of-the-art methods on popular counting benchmarks, but also achieves promising localization accuracy. The codes will be available at: TencentYoutuResearch/CrowdCounting-P2PNet. Qingyu Song 0001, Changan Wang, Zhengkai Jiang 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Yang Wu 0001 |
ICCV | 8 |
| 2021 | Transformer-based Dual Relation Graph for Multi-label Image RecognitionabstractThe simultaneous recognition of multiple objects in one image remains a challenging task, spanning multiple events in the recognition field such as various object scales, inconsistent appearances, and confused inter-class relationships. Recent research efforts mainly resort to the statistic label co-occurrences and linguistic word embedding to enhance the unclear semantics. Different from these researches, in this paper, we propose a novel Transformer-based Dual Relation learning framework, constructing complementary relationships by exploring two aspects of correlation, i.e., structural relation graph and semantic relation graph. The structural relation graph aims to capture long-range correlations from object context, by developing a cross-scale transformer-based architecture. The semantic graph dynamically models the semantic meanings of image objects with explicit semantic-aware constraints. In addition, we also incorporate the learnt structural relationship into the semantic graph, constructing a joint relation graph for robust representations. With the collaborative learning of these two effective relation graphs, our approach achieves new state-of-the-art on two popular multi-label recognition benchmarks, i.e. MS-COCO and VOC 2007 dataset. Ke Yan 0007, Yifan Zhao 0002, Feiyue Huang, Jia Li 0003 |
ICCV | 5 |
| 2021 | HifiFace: 3D Shape and Semantic Prior Guided High Fidelity Face SwappingabstractIn this work, we propose a high fidelity face swapping method, called HifiFace, which can well preserve the face shape of the source face and generate photo-realistic results. Unlike other existing face swapping works that only use face recognition model to keep the identity similarity, we propose 3D shape-aware identity to control the face shape with the geometric supervision from 3DMM and 3D face reconstruction method. Meanwhile, we introduce the Semantic Facial Fusion module to optimize the combination of encoder and decoder features and make adaptive blending, which makes the results more photo-realistic. Extensive experiments on faces in the wild demonstrate that our method can preserve better identity, especially on the face shape, and can generate more photo-realistic results than previous state-of-the-art methods. Code is available at: https://johann.wang/HifiFace Yuhan Wang 0002, Xu Chen 0024, Wenqing Chu, Ying Tai, Chengjie Wang 0001, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
IJCAI | 9 |
| 2021 | Distributed Attention for Grounded Image CaptioningabstractWe study the problem of weakly supervised grounded image captioning. That is, given an image, the goal is to automatically generate a sentence describing the context of the image with each noun word grounded to the corresponding region in the image. This task is challenging due to the lack of explicit fine-grained region word alignments as supervision. Previous weakly supervised methods mainly explore various kinds of regularization schemes to improve attention accuracy. However, their performances are still far from the fully supervised ones. One main issue that has been ignored is that the attention for generating visually groundable words may only focus on the most discriminate parts and can not cover the whole object. To this end, we propose a simple yet effective method to alleviate the issue, termed as partial grounding problem in our paper. Specifically, we design a distributed attention mechanism to enforce the network to aggregate information from multiple spatially different regions with consistent semantics while generating the words. Therefore, the union of the focused region proposals should form a visual region that encloses the object of interest completely. Extensive experiments have demonstrated the superiority of our proposed method compared with the state-of-the-arts. Nenglun Chen, Xingjia Pan, Runnan Chen, Lei Yang 0048, Zhiwen Lin, Yuqiang Ren, Haolei Yuan, Feiyue Huang, Wenping Wang 0001 |
ACM Multimedia | 9 |
| 2021 | Spatiotemporal Inconsistency Learning for DeepFake Video DetectionabstractThe rapid development of facial manipulation techniques has aroused public concerns in recent years. Following the success of deep learning, existing methods always formulate DeepFake video detection as a binary classification problem and develop frame-based and video-based solutions. However, little attention has been paid to capturing the spatial-temporal inconsistency in forged videos. To address this issue, we term this task as a Spatial-Temporal Inconsistency Learning (STIL) process and instantiate it into a novel STIL block, which consists of a Spatial Inconsistency Module (SIM), a Temporal Inconsistency Module (TIM), and an Information Supplement Module (ISM). Specifically, we present a novel temporal modeling paradigm in TIM by exploiting the temporal difference over adjacent frames along with both horizontal and vertical directions. And the ISM simultaneously utilizes the spatial information from SIM and temporal information from TIM to establish a more comprehensive spatial-temporal representation. Moreover, our STIL block is flexible and could be plugged into existing 2D CNNs. Extensive experiments and visualizations are presented to demonstrate the effectiveness of our method against the state-of-the-art competitors. Taiping Yao, Shouhong Ding, Feiyue Huang, Lizhuang Ma |
ACM Multimedia | 6 |
| 2021 | LSTC: Boosting Atomic Action Detection with Long-Short-Term ContextabstractIn this paper, we place the atomic action detection problem intoa Long-Short Term Context (LSTC) to analyze how the temporalreliance among video signals affect the action detection results. Todo this, we decompose the action recognition pipeline into short-term and long-term reliance, in terms of the hypothesis that the twokinds of context are conditionally independent given the objectiveaction instance. Within our design, a local aggregation branch isutilized to gather dense and informative short-term cues, while ahigh order long-term inference branch is designed to reason theobjective action class from high-order interaction between actor andother person or person pairs. Both branches independently predictthe context-specific actions and the results are merged in the end.We demonstrate that both temporal grains are beneficial to atomicaction recognition. On the mainstream benchmarks of atomic actiondetection, our design can bring significant performance gain fromthe existing state-of-the-art pipeline. Yuxi Li 0009, Boshen Zhang, Jian Li 0062, Yabiao Wang, Weiyao Lin, Chengjie Wang 0001, Feiyue Huang |
ACM Multimedia | 8 |
| 2021 | Adaptive Normalized Representation Learning for Generalizable Face Anti-SpoofingabstractWith various face presentation attacks arising under unseen scenarios, face anti-spoofing (FAS) based on domain generalization (DG) has drawn growing attention due to its robustness. Most existing methods utilize DG frameworks to align the features to seek a compact and generalized feature space. However, little attention has been paid to the feature extraction process for the FAS task, especially the influence of normalization, which also has a great impact on the generalization of the learned representation. To address this issue, we propose a novel perspective of face anti-spoofing that focuses on the normalization selection in the feature extraction process. Concretely, an Adaptive Normalized Representation Learning (ANRL) framework is devised, which adaptively selects feature normalization methods according to the inputs, aiming to learn domain-agnostic and discriminative representation. Moreover, to facilitate the representation learning, Dual Calibration Constraints are designed, including Inter-Domain Compatible loss and Inter-Class Separable loss, which provide a better optimization direction for generalizable representation. Extensive experiments and visualizations are presented to demonstrate the effectiveness of our method against the SOTA competitors. Shubao Liu, Ke-Yue Zhang, Taiping Yao, Mingwei Bi, Shouhong Ding, Feiyue Huang, Lizhuang Ma |
ACM Multimedia | 7 |
| 2021 | Discriminator-free Generative Adversarial AttackabstractThe Deep Neural Networks are vulnerable to adversarial examples (Figure 1), making the DNNs-based systems collapsed by adding the inconspicuous perturbations to the images. Most of the existing works for adversarial attack are gradient-based and suffer from the latency efficiencies and the load on GPU memory. The generative-based adversarial attacks can get rid of this limitation, and some relative works propose the approaches based on GAN. However, suffering from the difficulty of the convergence of training a GAN, the adversarial examples have either bad attack ability or bad visual quality. In this work, we find that the discriminator could be not necessary for generative-based adversarial attack, and propose the Symmetric Saliency-based Auto-Encoder (SSAE) to generate the perturbations, which is composed of the saliency map module and the angle-norm disentanglement of the features module. The advantage of our proposed method lies in that it is not depending on discriminator, and uses the generative saliency map to pay more attention to label-relevant regions. The extensive experiments among the various tasks, datasets, and models demonstrate that the adversarial examples generated by SSAE not only make the widely-used models collapse, but also achieves good visual quality. The code is available at: https://github.com/BravoLu/SSAE. Shaohao Lu, Yuqiao Xian, Xing Sun 0001, Feiyue Huang, Wei-Shi Zheng 0001 |
ACM Multimedia | 7 |
| 2021 | Analogous to Evolutionary Algorithm: Designing a Unified Sequence ModelabstractInspired by biological evolution, we explain the rationality of Vision Transformer by analogy with the proven practical Evolutionary Algorithm (EA) and derive that both of them have consistent mathematical representation. Analogous to the dynamic local population in EA, we improve the existing transformer structure and propose a more efficient EAT model, and design task-related heads to deal with different tasks more flexibly. Moreover, we introduce the spatial-filling curve into the current vision transformer to sequence image data into a uniform sequential format. Thus we can design a unified EAT framework to address multi-modal tasks, separating the network architecture from the data format adaptation. Our approach achieves state-of-the-art results on the ImageNet classification task compared with recent vision transformer works while having smaller parameters and greater throughput. We further conduct multi-modal tasks to demonstrate the superiority of the unified EAT, \eg, Text-Based Image Retrieval, and our approach improves the rank-1 by +3.7 points over the baseline on the CSS dataset. Jiangning Zhang, Chao Xu 0023, Jian Li 0062, Wenzhou Chen, Yabiao Wang, Ying Tai, Shuo Chen 0003, Chengjie Wang 0001, Feiyue Huang, Yong Liu 0007 |
NeurIPS | 9 |
| 2021 | MIGO-NAS: Towards Fast and Generalizable Neural Architecture SearchabstractNeural architecture search (NAS) has achieved unprecedented performance in various computer vision tasks. However, most existing NAS methods are defected in search efficiency and model generalizability. In this paper, we propose a novel NAS framework, termed MIGO-NAS, with the aim to guarantee the efficiency and generalizability in arbitrary search spaces. On the one hand, we formulate the search space as a multivariate probabilistic distribution, which is then optimized by a novel multivariate information-geometric optimization (MIGO). By approximating the distribution with a sampling, training, and testing pipeline, MIGO guarantees the memory efficiency, training efficiency, and search flexibility. Besides, MIGO is the first time to decrease the estimation error of natural gradient in multivariate distribution. On the other hand, for a set of specific constraints, the neural architectures are generated by a novel dynamic programming network generation (DPNG), which significantly reduces the training cost under various hardware environments. Experiments validate the advantages of our approach over existing methods by establishing a superior accuracy and efficiency i.e., 2.39 test error on CIFAR-10 benchmark and 21.7 on ImageNet benchmark, with only 1.5 GPU hours and 96 GPU hours for searching, respectively. Besides, the searched architectures can be well generalize to computer vision tasks including object detection and semantic segmentation, i.e., 25× FLOPs compression, with 6.4 mAP gain over Pascal VOC dataset, and 29.9× FLOPs compression, with only 1.41 percent performance drop over Cityscapes dataset. The code is publicly available. Xiawu Zheng, Rongrong Ji, Qiang Wang 0060, Baochang Zhang 0001, Jie Chen 0001, Qixiang Ye, Feiyue Huang, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2021 | Exploring the Representativity of Art PaintingsabstractArt painting evaluation is sophisticated for a novice with no or limited knowledge on art criticism, and history. In this study, we propose the concept ofrepresentativityto evaluate paintings instead of using professional concepts, such as genre, media, and style, which may be confusing to non-professionals. We define the concept of representativity to evaluate quantitatively the extent to which a painting can represent the characteristics of an artists creations. We begin by proposing a novel deep representation of art paintings, which is enhanced by style information through a weighted pooling feature fusion module. In contrast to existing feature extraction approaches, the proposed framework embeds painting styles, and authorship information, and learns specific artwork characteristics in a single framework. Subsequently, we propose a graph-based learning method for representativity learning, which considers intra-category, and extra-category information. In view of the significance of historical factors in the art domain, we introduce the creation time of a painting into the learning process. User studies demonstrate our approach helps the public effectively access the creation characteristics of artists through sorting paintings by representativity from highest to lowest. Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Feiyue Huang, Oliver Deussen, Changsheng Xu |
IEEE Trans. Multim. | 5 |
| 2021 | Aggregating Global and Local Visual Representation for Vehicle Re-IDentificationabstractVehicle Re-Identification targets at searching for vehicle instances of the same identity with a given query. It has gained an increasing attention recently with a wide application prospects in video surveillance and intelligent transportation. The main challenge lies in how to distinguish the subtle differences between different vehicles, while capturing the slight similarity between the instance of the same vehicle in different viewpoints or illuminations. To this end, most existing methods focus on learning discriminative global representations, which leave the unique local details such as stickers, inspection labels and driver wearing being unexploited. In this paper, we present a novel coarse-to-fine scheme that aggregates global and local visual representations to boost the retrieval accuracy. Specifically, a multi-task learning framework combining an Attribute Learning branch and a Deep Ranking branch (termed ALDR) is first adopted to learn robust global features, which produces an initial ranking list. Then the local similarities between image patches in the initial ranking list and in the query is computed via a Multi-Channel and Multi-Scale Siamese network (termed MCMS-Siam). Finally, the retrieval result is returned after re-ranking the initial list according to such a combination of global and local similarities. Experimental results on the widely-used VehicleID dataset and VECH-WILD dataset demonstrate the merits of the proposed method over the state-of-the-art methods. Xianming Lin, Run Li, Xiawu Zheng, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
IEEE Trans. Multim. | 6 |
| 2021 | Part-aware Progressive Unsupervised Domain Adaptation for Person Re-IdentificationabstractUnsupervised domain adaptation (UDA) aims to mitigate the domain shift that occurs when transferring knowledge from a labeled source domain to an unlabeled target domain. While it has been studied for application in unsupervised person re-identification (ReID), the relations of feature distribution across the source and target domains remain underexplored, as they either ignore the local relations or omit the in-depth consideration of negative transfer when two domains do not share identical label spaces. In light of the above, this paper presents an innovative part-aware progressive adaptation network (PPAN) that exploits global and local relations for UDA-based ReID across domains. A multi-branch network is developed that explicitly learns discriminative feature representation from both whole-body images and body-part images under the supervision of a labeled source domain. Within each network branch, an independent UDA constraint is designed that aligns the global and local feature distributions from a labeled source domain with those of an unlabeled target domain. In addition, a novel progressive adaptation strategy (PAS) is designed that effectively alleviates the negative influence of outlier source identities. The proposed unsupervised ReID model is evaluated on five widely used datasets (Market-1501, DukeMTMC-reID, CUHK03, VIPeR and PRID), and experimental results demonstrate its superior robustness and effectiveness relative to state-of-the-art approaches. Fan Yang 0053, Shijian Lu, Huizhu Jia, Don Xie, Zongqiao Yu, Feiyue Huang, Wen Gao 0001 |
IEEE Trans. Multim. | 8 |
| 2021 | Content-Based Visual Summarization for Image CollectionsabstractWith the surge of images in the information era, people demand an effective and accurate way to access meaningful visual information. Accordingly, effective and accurate communication of information has become indispensable. In this article, we propose a content-based approach that automatically generates a clear and informative visual summarization based on design principles and cognitive psychology to represent image collections. We first introduce a novel method to make representative and nonredundant summarizations of image collections, thereby ensuring data cleanliness and emphasizing important information. Then, we propose a tree-based algorithm with a two-step optimization strategy to generate the final layout that operates as follows: (1) an initial layout is created by constructing a tree randomly based on the grouping results of the input image set; (2) the layout is refined through a coarse adjustment in a greedy manner, followed by gradient back propagation drawing on the training procedure of neural networks. We demonstrate the usefulness and effectiveness of our method via extensive experimental results and user studies. Our visual summarization algorithm can precisely and efficiently capture the main content of image collections better than alternative methods or commercial tools. Xingjia Pan, Fan Tang, Weiming Dong, Chongyang Ma, Yiping Meng, Feiyue Huang, Tong-Yee Lee, Changsheng Xu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | Rethinking Temporal Fusion for Video-Based Person Re-Identification on Semantic and Time AspectabstractRecently, the research interest of person re-identification (ReID) has gradually turned to video-based methods, which acquire a person representation by aggregating frame features of an entire video. However, existing video-based ReID methods do not consider the semantic difference brought by the outputs of different network stages, which potentially compromises the information richness of the person features. Furthermore, traditional methods ignore important relationship among frames, which causes information redundancy in fusion along the time axis. To address these issues, we propose a novel general temporal fusion framework to aggregate frame features on both semantic aspect and time aspect. As for the semantic aspect, a multi-stage fusion network is explored to fuse richer frame features at multiple semantic levels, which can effectively reduce the information loss caused by the traditional single-stage fusion. While, for the time axis, the existing intra-frame attention method is improved by adding a novel inter-frame attention module, which effectively reduces the information redundancy in temporal fusion by taking the relationship among frames into consideration. The experimental results show that our approach can effectively improve the video-based re-identification accuracy, achieving the state-of-the-art performance. Xinyang Jiang, Yifei Gong, Qize Yang, Feiyue Huang, Wei-Shi Zheng 0001, Feng Zheng 0001, Xing Sun 0001 |
AAAI | 5 |
| 2020 | Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorabstractGenerating temporal action proposals remains a very challenging problem, where the main issue lies in predicting precise temporal proposal boundaries and reliable action confidence in long and untrimmed real-world videos. In this paper, we propose an efficient and unified framework to generate temporal action proposals named Dense Boundary Generator (DBG), which draws inspiration from boundary-sensitive methods and implements boundary classification and action completeness regression for densely distributed proposals. In particular, the DBG consists of two modules: Temporal boundary classification (TBC) and Action-aware completeness regression (ACR). The TBC aims to provide two temporal boundary confidence maps by low-level two-stream features, while the ACR is designed to generate an action completeness score map by high-level action-aware features. Moreover, we introduce a dual stream BaseNet (DSB) to encode RGB and optical flow information, which helps to capture discriminative boundary and actionness features. Extensive experiments on popular benchmarks ActivityNet-1.3 and THUMOS14 demonstrate the superiority of DBG over the state-of-the-art proposal generator (e.g., MGG and BMN). Chuming Lin, Jian Li 0062, Yabiao Wang, Ying Tai, Donghao Luo 0001, Zhipeng Cui, Chengjie Wang 0001, Feiyue Huang, Rongrong Ji |
AAAI | 9 |
| 2020 | TEINet: Towards an Efficient Architecture for Video RecognitionabstractEfficiency is an important issue in designing video architectures for action recognition. 3D CNNs have witnessed remarkable progress in action recognition from videos. However, compared with their 2D counterparts, 3D convolutions often introduce a large amount of parameters and cause high computational cost. To relieve this problem, we propose an efficient temporal module, termed as Temporal Enhancement-and-Interaction (TEI Module), which could be plugged into the existing 2D CNNs (denoted by TEINet). The TEI module presents a different paradigm to learn temporal features by decoupling the modeling of channel correlation and temporal interaction. First, it contains a Motion Enhanced Module (MEM) which is to enhance the motion-related features while suppress irrelevant information (e.g., background). Then, it introduces a Temporal Interaction Module (TIM) which supplements the temporal contextual information in a channel-wise manner. This two-stage modeling scheme is not only able to capture temporal structure flexibly and effectively, but also efficient for model inference. We conduct extensive experiments to verify the effectiveness of TEINet on several benchmarks (e.g., Something-Something V1&V2, Kinetics, UCF101 and HMDB51). Our proposed TEINet can achieve a good recognition accuracy on these datasets but still preserve a high efficiency. Zhaoyang Liu 0001, Donghao Luo 0001, Yabiao Wang, Limin Wang 0002, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Tong Lu 0002 |
AAAI | 8 |
| 2020 | Revisiting Image Aesthetic Assessment via Self-Supervised Feature LearningabstractVisual aesthetic assessment has been an active research field for decades. Although latest methods have achieved promising performance on benchmark datasets, they typically rely on a large number of manual annotations including both aesthetic labels and related image attributes. In this paper, we revisit the problem of image aesthetic assessment from the self-supervised feature learning perspective. Our motivation is that a suitable feature representation for image aesthetic assessment should be able to distinguish different expert-designed image manipulations, which have close relationships with negative aesthetic effects. To this end, we design two novel pretext tasks to identify the types and parameters of editing operations applied to synthetic instances. The features from our pretext tasks are then adapted for a one-layer linear classifier to evaluate the performance in terms of binary aesthetic classification. We conduct extensive quantitative experiments on three benchmark datasets and demonstrate that our approach can faithfully extract aesthetics-aware features and outperform alternative pretext schemes. Moreover, we achieve comparable results to state-of-the-art supervised methods that use 10 million labels from ImageNet. Kekai Sheng, Weiming Dong, Menglei Chai, Feiyue Huang, Bao-Gang Hu, Rongrong Ji, Chongyang Ma |
AAAI | 6 |
| 2020 | Asymmetric Co-Teaching for Unsupervised Cross-Domain Person Re-IdentificationabstractPerson re-identification (re-ID), is a challenging task due to the high variance within identity samples and imaging conditions. Although recent advances in deep learning have achieved remarkable accuracy in settled scenes, i.e., source domain, few works can generalize well on the unseen target domain. One popular solution is assigning unlabeled target images with pseudo labels by clustering, and then retraining the model. However, clustering methods tend to introduce noisy labels and discard low confidence samples as outliers, which may hinder the retraining process and thus limit the generalization ability. In this study, we argue that by explicitly adding a sample filtering procedure after the clustering, the mined examples can be much more efficiently used. To this end, we design an asymmetric co-teaching framework, which resists noisy labels by cooperating two models to select data with possibly clean labels for each other. Meanwhile, one of the models receives samples as pure as possible, while the other takes in samples as diverse as possible. This procedure encourages that the selected training samples can be both clean and miscellaneous, and that the two models can promote each other iteratively. Extensive experiments show that the proposed framework can consistently benefit most clustering based methods, and boost the state-of-the-art adaptation accuracy. Our code is available at https://github.com/FlyingRoastDuck/ACT_AAAI20. Fengxiang Yang, Ke Li 0015, Zhun Zhong, Zhiming Luo, Xing Sun 0001, Hao Cheng 0012, Feiyue Huang, Rongrong Ji, Shaozi Li |
AAAI | 8 |
| 2020 | Viewpoint-Aware Loss with Angular Regularization for Person Re-IdentificationabstractAlthough great progress in supervised person re-identification (Re-ID) has been made recently, due to the viewpoint variation of a person, Re-ID remains a massive visual challenge. Most existing viewpoint-based person Re-ID methods project images from each viewpoint into separated and unrelated sub-feature spaces. They only model the identity-level distribution inside an individual viewpoint but ignore the underlying relationship between different viewpoints. To address this problem, we propose a novel approach, called Viewpoint-Aware Loss with Angular Regularization (VA-reID). Instead of one subspace for each viewpoint, our method projects the feature from different viewpoints into a unified hypersphere and effectively models the feature distribution on both the identity-level and the viewpoint-level. In addition, rather than modeling different viewpoints as hard labels used for conventional viewpoint classification, we introduce viewpoint-aware adaptive label smoothing regularization (VALSR) that assigns the adaptive soft label to feature representation. VALSR can effectively solve the ambiguity of the viewpoint cluster label assignment. Extensive experiments on the Market1501 and DukeMTMC-reID datasets demonstrated that our method outperforms the state-of-the-art supervised Re-ID methods. Zhihui Zhu, Xinyang Jiang, Feng Zheng 0001, Feiyue Huang, Xing Sun 0001, Wei-Shi Zheng 0001 |
AAAI | 5 |
| 2020 | Adversarial Refinement Network for Human Motion Prediction
Xianjin Chao, Yanrui Bin, Wenqing Chu, Xuan Cao, Yanhao Ge, Chengjie Wang 0001, Feiyue Huang, Howard Leung |
ACCV (2) | 8 |
| 2020 | CurricularFace: Adaptive Curriculum Learning Loss for Deep Face RecognitionabstractAs an emerging topic in face recognition, designing margin-based loss functions can increase the feature margin between different classes for enhanced discriminability. More recently, the idea of mining-based strategies is adopted to emphasize the misclassified samples, achieving promising results. However, during the entire training process, the prior methods either do not explicitly emphasize the sample based on its importance that renders the hard samples not fully exploited; or explicitly emphasize the effects of semi-hard/hard samples even at the early training stage that may lead to convergence issue. In this work, we propose a novel Adaptive Curriculum Learning loss (CurricularFace) that embeds the idea of curriculum learning into the loss function to achieve a novel training strategy for deep face recognition, which mainly addresses easy samples in the early training stage and hard ones in the later stage. Specifically, our CurricularFace adaptively adjusts the relative importance of easy and hard samples during different training stages. In each stage, different samples are assigned with different importance according to their corresponding difficultness. Extensive experimental results on popular benchmarks demonstrate the superiority of our CurricularFace over the state-of-the-art competitors. Yuge Huang, Yuhan Wang 0002, Ying Tai, Xiaoming Liu 0002, Pengcheng Shen, Shaoxin Li 0001, Feiyue Huang |
CVPR | 8 |
| 2020 | Attention Convolutional Binary Neural Tree for Fine-Grained Visual CategorizationabstractFine-grained visual categorization (FGVC) is an important but challenging task due to high intra-class variances and low inter-class variances caused by deformation, occlusion, illumination, etc. An attention convolutional binary neural tree architecture is presented to address those problems for weakly supervised FGVC. Specifically, we incorporate convolutional operations along edges of the tree structure, and use the routing functions in each node to determine the root-to-leaf computational paths within the tree. The final decision is computed as the summation of the predictions from leaf nodes. The deep convolutional operations learn to capture the representations of objects, and the tree structure characterizes the coarse-to-fine hierarchical feature learning process. In addition, we use the attention transformer module to enforce the network to capture discriminative features. The negative log-likelihood loss is used to train the entire network in an end-to-end fashion by SGD with back-propagation. Several experiments on the CUB-200-2011, Stanford Cars and Aircraft datasets demonstrate that the proposed method performs favorably against the state-of-the-arts. Ruyi Ji, Longyin Wen, Libo Zhang 0001, Dawei Du, Chen Zhao 0024, Xianglong Liu 0001, Feiyue Huang |
CVPR | 8 |
| 2020 | Learning by Analogy: Reliable Supervision From Transformations for Unsupervised Optical Flow EstimationabstractUnsupervised learning of optical flow, which leverages the supervision from view synthesis, has emerged as a promising alternative to supervised methods. However, the objective of unsupervised learning is likely to be unreliable in challenging scenes. In this work, we present a framework to use more reliable supervision from transformations. It simply twists the general unsupervised learning pipeline by running another forward pass with transformed data from augmentation, along with using transformed predictions of original data as the self-supervision signal. Besides, we further introduce a lightweight network with multiple frames by a highly-shared flow decoder. Our method consistently gets a leap of performance on several benchmarks with the best accuracy among deep unsupervised methods. Also, our method achieves competitive results to recent fully supervised methods while with much fewer parameters. Liang Liu 0007, Jiangning Zhang, Ruifei He, Yong Liu 0007, Yabiao Wang, Ying Tai, Donghao Luo 0001, Chengjie Wang 0001, Feiyue Huang |
CVPR | 10 |
| 2020 | Adversarial Semantic Data Augmentation for Human Pose Estimation
Yanrui Bin, Xuan Cao, Xinya Chen, Yanhao Ge, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Changxin Gao, Nong Sang |
ECCV (19) | 8 |
| 2020 | SSCGAN: Facial Attribute Editing via Style Skip Connections
Wenqing Chu, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Rongrong Ji |
ECCV (15) | 5 |
| 2020 | Improving Face Recognition from Hard Samples via Distribution Distillation Loss
Yuge Huang, Pengcheng Shen, Ying Tai, Shaoxin Li 0001, Xiaoming Liu 0002, Feiyue Huang, Rongrong Ji |
ECCV (30) | 7 |
| 2020 | Learning Semantic Neural Tree for Human Parsing
Ruyi Ji, Dawei Du, Libo Zhang 0001, Longyin Wen, Chen Zhao 0024, Feiyue Huang, Siwei Lyu |
ECCV (13) | 7 |
| 2020 | Interpretable Neural Network Decoupling
Rongrong Ji, Shaohui Lin, Baochang Zhang 0001, Chenqian Yan, Yongjian Wu 0001, Feiyue Huang, Ling Shao 0001 |
ECCV (15) | 7 |
| 2020 | Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking
Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Yanwei Fu 0001 |
ECCV (4) | 9 |
| 2020 | Enabling Deep Residual Networks for Weakly Supervised Object Detection
Yunhang Shen, Rongrong Ji, Yan Wang 0059, Feng Zheng 0001, Feiyue Huang, Yunsheng Wu |
ECCV (8) | 6 |
| 2020 | Temporal Distinct Representation Learning for Action Recognition
Junwu Weng, Donghao Luo 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Xudong Jiang 0001, Junsong Yuan 0001 |
ECCV (7) | 7 |
| 2020 | Face Anti-Spoofing via Disentangled Representation Learning
Ke-Yue Zhang, Taiping Yao, Jian Zhang 0079, Ying Tai, Shouhong Ding, Feiyue Huang, Lizhuang Ma |
ECCV (19) | 7 |
| 2020 | Arbitrary Style Transfer via Multi-Adaptation NetworkabstractArbitrary style transfer is a significant topic with research value and application prospect. A desired style transfer, given a content image and referenced style painting, would render the content image with the color tone and vivid stroke patterns of the style painting while synchronously maintaining the detailed content structure information. Style transfer approaches would initially learn content and style representations of the content and style references and then generate the stylized images guided by these representations. In this paper, we propose the multi-adaptation network which involves two self-adaptation (SA) modules and one co-adaptation (CA) module:the SA modules adaptively disentangle the content and style representations, i.e., content SA module uses position-wise self-attention to enhance content representation and style SA module uses channel-wise self-attention to enhance style representation; the CA module rearranges the distribution of style representation based on content representation distribution by calculating the local similarity between the disentangled content and style features in a non-local fashion. Moreover, a new disentanglement loss function enables our network to extract main style patterns and exact content structures to adapt to various input images, respectively. Various qualitative and quantitative experiments demonstrate that the proposed multi-adaptation network leads to better results than the state-of-the-art style transfer methods. Yingying Deng, Fan Tang, Weiming Dong, Feiyue Huang, Changsheng Xu |
ACM Multimedia | 5 |
| 2020 | Dense Scene Multiple Object Tracking with Box-Plane MatchingabstractMultiple Object Tracking (MOT) is an important task in computer vision. MOT is still challenging due to the occlusion problem, especially in dense scenes. Following the tracking-by-detection framework, we propose the Box-Plane Matching (BPM) method to improve the MOT performacne in dense scenes. First, we design the Layer-wise Aggregation Discriminative Model (LADM) to filter the noisy detections. Then, to associate remaining detections correctly, we introduce the Global Attention Feature Model (GAFM) to extract appearance feature and use it to calculate the appearance similarity between history tracklets and current detections. Finally, we propose the Box-Plane Matching strategy to achieve data association according to the motion similarity and appearance similarity between tracklets and detections. With the effectiveness of the three modules, our team achieves the 1st place on the Track-1 leaderboard in the ACM MM Grand Challenge HiEve 2020. Jinlong Peng, Yueyang Gu, Yabiao Wang, Chengjie Wang 0001, Feiyue Huang |
ACM Multimedia | 6 |
| 2020 | Exploring Language Prior for Mode-Sensitive Visual Attention ModelingabstractModeling human visual attention mechanism is a fundamental problem for the understanding of human vision, which has also been demonstrated as an important module for various multimedia applications such as image captioning and visual question answering. In this paper, we propose a new probabilistic framework for attention, and introduce the concept ofmode to model the flexibility and adaptability of attention modulation in complex environments. We characterize the correlations between the visual input, the activated mode, the saliency and the spatial allocation of attention via a graphical model representation, based on which we explore the lingual guidance from captioning data for the implementation of a mode-sensitive attention (MSA) model. The proposed framework explicitly justifies the usage of center bias for fixation prediction and can convert an arbitrary learning-based backbone attention model to a more robust multi-mode version. Experimental results on the York120, MIT1003 and PASCAL datasets demonstrate the effectiveness of the proposed method. Xiaoshuai Sun, Xuying Zhang, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
ACM Multimedia | 5 |
| 2020 | NOH-NMS: Improving Pedestrian Detection by Nearby Objects HallucinationabstractGreedy-NMS inherently raises a dilemma, where a lower NMS threshold will potentially lead to a lower recall rate and a higher threshold introduces more false positives. This problem is more severe in pedestrian detection because the instance density varies more intensively. However, previous works on NMS don't consider or vaguely consider the factor of the existent of nearby pedestrians. Thus, we propose \heatmapname (\heatmapnameshort ), which pinpoints the objects nearby each proposal with a Gaussian distribution, together with \nmsname, which dynamically eases the suppression for the space that might contain other objects with a high likelihood. Compared to Greedy-NMS, our method, as the state-of-the-art, improves by $3.9%$ AP, $5.1%$ Recall, and $0.8%$ MR\textsuperscript-2 on CrowdHuman to $89.0%$ AP and $92.9%$ Recall, and $43.9%$ MR\textsuperscript-2 respectively. Penghao Zhou, Chong Zhou, Junlong Du, Xing Sun 0001, Feiyue Huang |
ACM Multimedia | 7 |
| 2020 | Rotated Binary Neural NetworkabstractBinary Neural Network (BNN) shows its predominance in reducing the complexity of deep neural networks. However, it suffers severe performance degradation. One of the major impediments is the large quantization error between the full-precision weight vector and its binary vector. Previous works focus on compensating for the norm gap while leaving the angular bias hardly touched. In this paper, for the first time, we explore the influence of angular bias on the quantization error and then introduce a Rotated Binary Neural Network (RBNN), which considers the angle alignment between the full-precision weight vector and its binarized version. At the beginning of each training epoch, we propose to rotate the full-precision weight vector to its binary vector to reduce the angular bias. To avoid the high complexity of learning a large rotation matrix, we further introduce a bi-rotation formulation that learns two smaller rotation matrices. In the training stage, we devise an adjustable rotated weight vector for binarization to escape the potential local optimum. Our rotation leads to around 50% weight flips which maximize the information gain. Finally, we propose a training-aware approximation of the sign function for the gradient backward. Experiments on CIFAR-10 and ImageNet demonstrate the superiorities of RBNN over many state-of-the-arts. Our source code, experimental settings, training logs and binary models are available at https://github.com/lmbxmu/RBNN. Mingbao Lin, Rongrong Ji, Baochang Zhang 0001, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Chia-Wen Lin |
NeurIPS | 7 |
| 2020 | UWSOD: Toward Fully-Supervised-Level Capacity Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) has attracted extensive research attention due to its great flexibility of exploiting large-scale dataset with only image-level annotations for detector training. Despite its great advance in recent years, WSOD still suffers limited performance, which is far below that of fully supervised object detection (FSOD). As most WSOD methods depend on object proposal algorithms to generate candidate regions and are also confronted with challenges like low-quality predicted bounding boxes and large scale variation. In this paper, we propose a unified WSOD framework, termed UWSOD, to develop a high-capacity general detection model with only image-level labels, which is self-contained and does not require external modules or additional supervision. To this end, we exploit three important components, i.e., object proposal generation, bounding-box fine-tuning and scale-invariant features. First, we propose an anchor-based self-supervised proposal generator to hypothesize object locations, which is trained end-to-end with supervision created by UWSOD for both objectness classification and regression. Second, we develop a step-wise bounding-box fine-tuning to refine both detection scores and coordinates by progressively select high-confidence object proposals as positive samples, which bootstraps the quality of predicted bounding boxes. Third, we construct a multi-rate resampling pyramid to aggregate multi-scale contextual information, which is the first in-network feature hierarchy to handle scale variation in WSOD. Extensive experiments on PASCAL VOC and MS COCO show that the proposed UWSOD achieves competitive results with the state-of-the-art WSOD methods while not requiring external modules or additional supervision. Moreover, the upper-bound performance of UWSOD with class-agnostic ground-truth bounding boxes approaches Faster R-CNN, which demonstrates UWSOD has fully-supervised-level capacity. Yunhang Shen, Rongrong Ji, Yongjian Wu 0001, Feiyue Huang |
NeurIPS | 5 |
| 2020 | Semi-Supervised Adversarial Monocular Depth EstimationabstractIn this paper, we address the problem of monocular depth estimation when only a limited number of training image-depth pairs are available. To achieve a high regression accuracy, the state-of-the-art estimation methods rely on CNNs trained with a large number of image-depth pairs, which are prohibitively costly or even infeasible to acquire. Aiming to break the curse of such expensive data collections, we propose a semi-supervised adversarial learning framework that only utilizes a small number of image-depth pairs in conjunction with a large number of easily-available monocular images to achieve high performance. In particular, we use one generator to regress the depth and two discriminators to evaluate the predicted depth, i.e., one inspects the image-depth pair while the other inspects the depth channel alone. These two discriminators provide their feedbacks to the generator as the loss to generate more realistic and accurate depth predictions. Experiments show that the proposed approach can (1) improve most state-of-the-art models on the NYUD v2 dataset by effectively leveraging additional unlabeled data sources; (2) reach state-of-the-art accuracy when the training set is small, e.g., on the Make3D dataset; (3) adapt well to an unseen new dataset (Make3D in our case) after training on an annotated dataset (KITTI in our case). Rongrong Ji, Ke Li 0015, Yan Wang 0059, Xiaoshuai Sun, Feng Guo 0005, Yongjian Wu 0001, Feiyue Huang, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2020 | Rio: a personal storage system in multi-device and cloud
Feiyue Huang |
J. Supercomput. | 3 |
| 2020 | Correction to: Rio: a personal storage system in multi-device and cloud
Feiyue Huang |
J. Supercomput. | 3 |
| 2019 | Scale Invariant Fully Convolutional Network: Detecting Hands EfficientlyabstractExisting hand detection methods usually follow the pipeline of multiple stages with high computation cost, i.e., feature extraction, region proposal, bounding box regression, and additional layers for rotated region detection. In this paper, we propose a new Scale Invariant Fully Convolutional Network (SIFCN) trained in an end-to-end fashion to detect hands efficiently. Specifically, we merge the feature maps from high to low layers in an iterative way, which handles different scales of hands better with less time overhead comparing to concatenating them simply. Moreover, we develop the Complementary Weighted Fusion (CWF) block to make full use of the distinctive features among multiple layers to achieve scale invariance. To deal with rotated hand detection, we present the rotation map to get rid of complex rotation and derotation layers. Besides, we design the multi-scale loss scheme to accelerate the training process significantly by adding supervision to the intermediate layers of the network. Compared with the state-of-the-art methods, our algorithm shows comparable accuracy and runs a 4.23 times faster speed on the VIVA dataset and achieves better average precision on Oxford hand detection dataset at a speed of 62.5 fps. Dawei Du, Libo Zhang 0001, Tiejian Luo, Feiyue Huang, Siwei Lyu |
AAAI | 6 |
| 2019 | Towards Highly Accurate and Stable Face Alignment for High-Resolution VideosabstractIn recent years, heatmap regression based models have shown their effectiveness in face alignment and pose estimation. However, Conventional Heatmap Regression (CHR) is not accurate nor stable when dealing with high-resolution facial videos, since it finds the maximum activated location in heatmaps which are generated from rounding coordinates, and thus leads to quantization errors when scaling back to the original high-resolution space. In this paper, we propose a Fractional Heatmap Regression (FHR) for high-resolution video-based face alignment. The proposed FHR can accurately estimate the fractional part according to the 2D Gaussian function by sampling three points in heatmaps. To further stabilize the landmarks among continuous video frames while maintaining the precise at the same time, we propose a novel stabilization loss that contains two terms to address time delay and non-smooth issues, respectively. Experiments on 300W, 300VW and Talking Face datasets clearly demonstrate that the proposed method is more accurate and stable than the state-ofthe-art models. Ying Tai, Yicong Liang, Xiaoming Liu 0002, Lei Duan, Chengjie Wang 0001, Feiyue Huang, Yu Chen 0037 |
AAAI | 7 |
| 2019 | Towards Optimal Fine Grained Retrieval via Decorrelated Centralized Loss with Normalize-Scale LayerabstractRecent advances on fine-grained image retrieval prefer learning convolutional neural network (CNN) with specific fullyconnect layer designed loss function for discriminative feature representation. Essentially, such loss should establish a robust metric to efficiently distinguish high-dimensional features within and outside fine-grained categories. To this end, the existing loss functions are defected in two aspects: (a) The feature relationship is encoded inside the training batch. Such a local scope leads to low accuracy. (b) The error is established by the mean square, which needs pairwise distance computation in training set and results in low efficiency. In this paper, we propose a novel metric learning scheme, termed Normalize-Scale Layer and Decorrelated Global Centralized Ranking Loss, which achieves extremely efficient and discriminative learning, i.e., 5× speedup over triplet loss and 12% recall boost on CARS196. Our method originates from the classic softmax loss, which has a global structure but does not directly optimize the distance metric as well as the inter/intra class distance. We tackle this issue through a hypersphere layer and a global centralized ranking loss with a pairwise decorrelated learning. In particular, we first propose a Normalize-Scale Layer to eliminate the gap between metric distance (for measuring distance in retrieval) and dot product (for dimension reduction in classification). Second, the relationship between features is encoded under a global centralized ranking loss, which targets at optimizing metric distance globally and accelerating learning procedure. Finally, the centers are further decorrelated by Gram-Schmidt process, leading to extreme efficiency (with 20 epochs in training procedure) and discriminability in feature learning. We have conducted quantitative evaluations on two fine-grained retrieval benchmark. The superior performance demonstrates the merits of the proposed approach over the state-of-the-arts. Xiawu Zheng, Rongrong Ji, Xiaoshuai Sun, Baochang Zhang 0001, Yongjian Wu 0001, Feiyue Huang |
AAAI | 6 |
| 2019 | Exploiting Kernel Sparsity and Entropy for Interpretable CNN CompressionabstractCompressing convolutional neural networks (CNNs) has received ever-increasing research focus. However, most existing CNN compression methods do not interpret their inherent structures to distinguish the implicit redundancy. In this paper, we investigate the problem of CNN compression from a novel interpretable perspective. The relationship between the input feature maps and 2D kernels is revealed in a theoretical framework, based on which a kernel sparsity and entropy (KSE) indicator is proposed to quantitate the feature map importance in a feature-agnostic manner to guide model compression. Kernel clustering is further conducted based on the KSE indicator to accomplish high-precision CNN compression. KSE is capable of simultaneously compressing each layer in an efficient way, which is significantly faster compared to previous data-driven feature map pruning methods. We comprehensively evaluate the compression and speedup of the proposed method on CIFAR-10, SVHN and ImageNet 2012. Our method demonstrates superior performance gains over previous ones. In particular, it achieves 4.7× FLOPs reduction and 2.9× compression on ResNet-50 with only a top-5 accuracy drop of 0.35% on ImageNet 2012, which significantly outperforms state-of-the-art methods. Shaohui Lin, Baochang Zhang 0001, Jianzhuang Liu, David S. Doermann, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
CVPR | 7 |
| 2019 | DSFD: Dual Shot Face DetectorabstractRecently, Convolutional Neural Network (CNN) has achieved great success in face detection. However, it remains a challenging problem for the current face detection methods owing to high degree of variability in scale, pose, occlusion, expression, appearance and illumination. In this Paper, we propose a novel detection network named Dual Shot face Detector(DSFD). which inherits the architecture of SSD and introduces a Feature Enhance Module (FEM) for transferring the original feature maps to extend the single shot detector to dual shot detector. Specially, progressive anchor loss (PAL) computed by using two set of anchors is adopted to effectively facilitate the features. Additionally, we propose an improved anchor matching (IAM) method by integrating novel data augmentation techniques and anchor design strategy in our DSFD to provide better initialization for the regressor. Extensive experiments on popular benchmarks: WIDER FACE (easy: 0.966, medium: 0.957, hard: 0.904) and FDDB ( discontinuous: 0.991, continuous: 0.862 ) demonstrate the superiority of DSFD over the state-of-the-art face detection methods (e.g., PyramidBox and SRN). Code will be made available upon publication. Jian Li 0062, Yabiao Wang, Changan Wang, Ying Tai, Jianjun Qian, Jian Yang 0003, Chengjie Wang 0001, Feiyue Huang |
CVPR | 9 |
| 2019 | Towards Optimal Structured CNN Pruning via Generative Adversarial LearningabstractStructured pruning of filters or neurons has received increased focus for compressing convolutional neural networks. Most existing methods rely on multi-stage optimizations in a layer-wise manner for iteratively pruning and retraining which may not be optimal and may be computation intensive. Besides, these methods are designed for pruning a specific structure, such as filter or block structures without jointly pruning heterogeneous structures. In this paper, we propose an effective structured pruning approach that jointly prunes filters as well as other structures in an end-to-end manner. To accomplish this, we first introduce a soft mask to scale the output of these structures by defining a new objective function with sparsity regularization to align the output of baseline and network with this mask. We then effectively solve the optimization problem by generative adversarial learning (GAL), which learns a sparse soft mask in a label-free and an end-to-end manner. By forcing more scale factors in the soft mask to zero, the fast iterative shrinkage-thresholding algorithm (FISTA) can be leveraged to fast and reliably remove the corresponding structures. Extensive experiments demonstrate the effectiveness of GAL on different datasets, including MNIST, CIFAR-10 and ImageNet ILSVRC 2012. For example, on ImageNet ILSVRC 2012, the pruned ResNet-50 achieves 10.88% Top-5 error and results in a factor of 3.7x speedup. This significantly outperforms state-of-the-art methods. Shaohui Lin, Rongrong Ji, Chenqian Yan, Baochang Zhang 0001, Liujuan Cao, Qixiang Ye, Feiyue Huang, David S. Doermann |
CVPR | 7 |
| 2019 | Pyramidal Person Re-IDentification via Multi-Loss Dynamic TrainingabstractMost existing Re-IDentification (Re-ID) methods are highly dependent on precise bounding boxes that enable images to be aligned with each other. However, due to the challenging practical scenarios, current detection models often produce inaccurate bounding boxes, which inevitably degenerate the performance of existing Re-ID algorithms. In this paper, we propose a novel coarse-to-fine pyramid model to relax the need of bounding boxes, which not only incorporates local and global information, but also integrates the gradual cues between them. The pyramid model is able to match at different scales and then search for the correct image of the same identity, even when the image pairs are not aligned. In addition, in order to learn discriminative identity representation, we explore a dynamic training scheme to seamlessly unify two losses and extract appropriate shared information between them. Experimental results clearly demonstrate that the proposed method achieves the state-of-the-art results on three datasets. Especially, our approach exceeds the current best method by 9.5% on the most challenging CUHK03 dataset. Feng Zheng 0001, Cheng Deng 0002, Xing Sun 0001, Xinyang Jiang, Zongqiao Yu, Feiyue Huang, Rongrong Ji |
CVPR | 7 |
| 2019 | Vocal Melody Extraction via DNN-based Pitch Estimation and Salience-based Pitch RefinementabstractData-driven methods for melody extraction from polyphonic music generally require large amounts of labeled data for model training. However, musical data with annotations of melody fundamental frequency (F0) are rare and hard to obtain. To overcome this limitation, in this paper we propose to use melody MIDI files, which are more massively available, as the sources of labels to train a deep neural network (DNN) model for melody extraction. For each testing audio, the pitch sequence estimated by DNN is comprised of note numbers quantized at semitone level, and their resolution is relatively low. Therefore, we further propose a salience-based method to refine the pitch estimate of DNN to a higher resolution of 10 cents. Experimental results on three public datasets indicate that our method outperforms four state-of-the-art melody extraction methods in most cases. Yongwei Gao, Bilei Zhu, Wei Li 0012, Ke Li 0015, Yongjian Wu 0001, Feiyue Huang |
ICASSP | 6 |
| 2019 | Universal Adversarial Perturbation via Prior Driven Uncertainty ApproximationabstractDeep learning models have shown their vulnerabilities to universal adversarial perturbations (UAP), which are quasi-imperceptible. Compared to the conventional supervised UAPs that suffer from the knowledge of training data, the data-independent unsupervised UAPs are more applicable. Existing unsupervised methods fail to take advantage of the model uncertainty to produce robust perturbations. In this paper, we propose a new unsupervised universal adversarial perturbation method, termed as Prior Driven Uncertainty Approximation (PD-UA), to generate a robust UAP by fully exploiting the model uncertainty at each network layer. Specifically, a Monte Carlo sampling method is deployed to activate more neurons to increase the model uncertainty for a better adversarial perturbation. Thereafter, a textural bias prior to revealing a statistical uncertainty is proposed, which helps to improve the attacking performance. The UAP is crafted by the stochastic gradient descent algorithm with a boosted momentum optimizer, and a Laplacian pyramid frequency model is finally used to maintain the statistical uncertainty. Extensive experiments demonstrate that our method achieves well attacking performances on the ImageNet validation set, and significantly improves the fooling rate compared with the state-of-the-art methods. Hong Liu 0009, Rongrong Ji, Jie Li 0052, Baochang Zhang 0001, Yue Gao 0002, Yongjian Wu 0001, Feiyue Huang |
ICCV | 7 |
| 2019 | LGM-Net: Learning to Generate Matching Networks for Few-Shot LearningabstractIn this work, we propose a novel meta-learning approach for few-shot classification, which learns transferable prior knowledge across tasks and directly produces network parameters for similar unseen tasks with training samples. Our approach, called LGM-Net, includes two key modules, namely, TargetNet and MetaNet. The TargetNet module is a neural network for solving a specific task and the MetaNet module aims at learning to generate functional weights for TargetNet by observing training samples. We also present an intertask normalization strategy for the training process to leverage common information shared across different tasks. The experimental results on Omniglot and miniImageNet datasets demonstrate that LGM-Net can effectively adapt to similar unseen tasks and achieve competitive performance, and the results on synthetic datasets show that transferable prior knowledge is learned by the MetaNet module via mapping training data to functional weights. LGM-Net enables fast learning and adaptation since no further tuning steps are required compared to other meta-learning approaches Huai-Yu Li, Weiming Dong, Xing Mei, Chongyang Ma, Feiyue Huang, Bao-Gang Hu |
ICML | 5 |
| 2019 | A Part Power Set Model for Scale-Free Person RetrievalabstractRecently, person re-identification (re-ID) has attracted increasing research attention, which has broad application prospects in video surveillance and beyond. To this end, most existing methods highly relied on well-aligned pedestrian images and hand-engineered part-based model on the coarsest feature map. In this paper, to lighten the restriction of such fixed and coarse input alignment, an end-to-end part power set model with multi-scale features is proposed, which captures the discriminative parts of pedestrians from global to local, and from coarse to fine, enabling part-based scale-free person re-ID. In particular, we first factorize the visual appearance by enumerating $k$-combinations for all $k$ of $n$ body parts to exploit rich global and partial information to learn discriminative feature maps. Then, a combination ranking module is introduced to guide the model training with all combinations of body parts, which alternates between ranking combinations and estimating an appearance model. To enable scale-free input, we further exploit the pyramid architecture of deep networks to construct multi-scale feature maps with a feasible amount of extra cost in term of memory and time. Extensive experiments on the mainstream evaluation datasets, including Market-1501, DukeMTMC-reID and CUHK03, validate that our method achieves the state-of-the-art performance. Yunhang Shen, Rongrong Ji, Xiaopeng Hong, Feng Zheng 0001, Yongjian Wu 0001, Feiyue Huang |
IJCAI | 7 |
| 2019 | Multi-modal Multi-layer Fusion Network with Average Binary Center Loss for Face Anti-spoofingabstractFace anti-spoofing detection is critical to guarantee the security of biometric face recognition systems. Despite extensive advances in facial anti-spoofing based on single-model image, little work has been devoted to multi-modal anti-spoofing, which is however widely encountered in real-world scenarios. Following the recent progress, this paper mainly focuses on multi-modal face anti-spoofing and aims to solve the following two challenges: (1) how to effectively fuse multi-modal information; and (2) how to effectively learn distinguishable features despite single cross-entropy loss. We propose a novel Multi-modal Multi-layer Fusion Convolutional Neural Network (mmfCNN), which targets at finding a discriminative model for recognizing the subtle differences between live and spoof faces. The mmfCNN can fully use different information provided by diverse modalities, which is based on a weight-adaptation aggregation approach. Specifically, we utilize a multi-layer fusion model to further aggregate the features from different layers, which fuses the low-, mid- and high-level information from different modalities in a unified framework. Moreover, a novel Average Binary Center (ABC) loss is proposed to maximize the dissimilarity between the features of live and spoof faces, which helps to stabilize the training to generate a robust and discriminative model. Extensive experiments conducted on the CISIA-SURF and 3DMAD datasets verify the significance and generalization capability of the proposed method for the face anti-spoofing task. Code is available at: https://github.com/SkyKuang/Face-anti-spoofing. Huafeng Kuang, Rongrong Ji, Hong Liu 0009, Shengchuan Zhang, Xiaoshuai Sun, Feiyue Huang, Baochang Zhang 0001 |
ACM Multimedia | 6 |
| 2019 | Variational Structured Semantic Inference for Diverse Image CaptioningabstractDespite the exciting progress in image captioning, generating diverse captions for a given image remains as an open problem. Existing methods typically apply generative models such as Variational Auto-Encoder to diversify the captions, which however neglect two key factors of diverse expression, i.e., the lexical diversity and the syntactic diversity. To model these two inherent diversities in image captioning, we propose a Variational Structured Semantic Inferring model (termed VSSI-cap) executed in a novel structured encoder-inferer-decoder schema. VSSI-cap mainly innovates in a novel structure, i.e., Variational Multi-modal Inferring tree (termed VarMI-tree). In particular, conditioned on the visual-textual features from the encoder, the VarMI-tree models the lexical and syntactic diversities by inferring their latent variables (with variations) in an approximate posterior inference guided by a visual semantic prior. Then, a reconstruction loss and the posterior-prior KL-divergence are jointly estimated to optimize the VSSI-cap model. Finally, diverse captions are generated upon the visual features and the latent variables from this structured encoder-inferer-decoder model. Experiments on the benchmark dataset show that the proposed VSSI-cap achieves significant improvements over the state-of-the-arts. Fuhai Chen, Rongrong Ji, Jiayi Ji, Xiaoshuai Sun, Baochang Zhang 0001, Xuri Ge, Yongjian Wu 0001, Feiyue Huang, Yan Wang 0059 |
NeurIPS | 8 |
| 2019 | Joint face alignment and segmentation via deep multi-task learning
Fan Tang, Weiming Dong, Feiyue Huang, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Harnessing Synthesized Abstraction Images to Improve Facial Attribute RecognitionabstractFacial attribute recognition is an important and yet challenging research topic. Different from most previous approaches which predict attributes only based on the whole images, this paper leverages facial parts locations for better attribute prediction. A facial abstraction image which contains both local facial parts and facial texture information is introduced. This abstraction image is generated by a Generative Adversarial Network (GAN). Then we build a dual-path facial attribute recognition network to utilize features from the original face images and facial abstraction images. Empirically, the features of facial abstraction images are complementary to features of original face images. With the facial parts localized by the abstraction images, our method improves facial attributes recognition, especially the attributes located on small face regions. Extensive evaluations conducted on CelebA and LFWA benchmark datasets show that state-of-the-art performance is achieved. Keke He, Yanwei Fu 0001, Wuhao Zhang, Chengjie Wang 0001, Yu-Gang Jiang 0001, Feiyue Huang, Xiangyang Xue 0001 |
IJCAI | 6 |
| 2018 | Accelerating Convolutional Networks via Global & Dynamic Filter PruningabstractAccelerating convolutional neural networks has recently received ever-increasing research focus. Among various approaches proposed in the literature, filter pruning has been regarded as a promising solution, which is due to its advantage in significant speedup and memory reduction of both network model and intermediate feature maps. To this end, most approaches tend to prune filters in a layer-wise fixed manner, which is incapable to dynamically recover the previously removed filter, as well as jointly optimize the pruned network across layers. In this paper, we propose a novel global & dynamic pruning (GDP) scheme to prune redundant filters for CNN acceleration. In particular, GDP first globally prunes the unsalient filters across all layers by proposing a global discriminative function based on prior knowledge of filters. Second, it dynamically updates the filter saliency all over the pruned sparse network, and then recover the mistakenly pruned filter, followed by a retraining phase to improve the model accuracy. Specially, we effectively solve the corresponding non-convex optimization problem of the proposed GDP via stochastic gradient descent with greedy alternative updating. Extensive experiments show that, comparing to the state-of-the-art filter pruning methods, the proposed approach achieves superior performance to accelerate several cutting-edge CNNs on the ILSVRC 2012 benchmark. Shaohui Lin, Rongrong Ji, Yongjian Wu 0001, Feiyue Huang, Baochang Zhang 0001 |
IJCAI | 5 |
| 2018 | Adversarial Attribute-Image Person Re-identificationabstractWhile attributes have been widely used for person re-identification (Re-ID) which aims at matching the same person images across disjoint camera views, they are used either as extra features or for performing multi-task learning to assist the image-image matching task. However, how to find a set of person images according to a given attribute description, which is very practical in many surveillance applications, remains a rarely investigated cross-modality matching problem in person Re-ID. In this work, we present this challenge and leverage adversarial learning to formulate the attribute-image cross-modality person Re-ID model. By imposing a semantic consistency constraint across modalities as a regularization, the adversarial learning enables to generate image-analogous concepts of query attributes for matching the corresponding images at both global level and semantic ID level. We conducted extensive experiments on three attribute datasets and demonstrated that the regularized adversarial modelling is so far the most effective method for the attribute-image cross-modality person Re-ID problem. Zhou Yin, Wei-Shi Zheng 0001, Ancong Wu, Hong-Xing Yu, Hai Wan, Feiyue Huang, Jian-Huang Lai |
IJCAI | 7 |
| 2018 | Centralized Ranking Loss with Weakly Supervised Localization for Fine-Grained Object RetrievalabstractFine-grained object retrieval has attracted extensive research focus recently. Its state-of-the-art schemesare typically based upon convolutional neural network (CNN) features. Despite the extensive progress, two issues remain open. On one hand, the deep features are coarsely extracted at image level rather than precisely at object level, which are interrupted by background clutters. On the other hand, training CNN features with a standard triplet loss is time consuming and incapable to learn discriminative features. In this paper, we present a novel fine-grained object retrieval scheme that conquers these issues in a unified framework. Firstly, we introduce a novel centralized ranking loss (CRL), which achieves a very efficient (1,000times training speedup comparing to the triplet loss) and discriminative feature learning by a ?centralized? global pooling. Secondly, a weakly supervised attractive feature extraction is proposed, which segments object contours with top-down saliency. Consequently, the contours are integrated into the CNN response map to precisely extract features ?within? the target object. Interestingly, we have discovered that the combination of CRL and weakly supervised learning can reinforce each other. We evaluate the performance ofthe proposed scheme on widely-used benchmarks including CUB200-2011 and CARS196. We havereported significant gains over the state-of-the-art schemes, e.g., 5.4% over SCDA [Wei et al., 2017]on CARS196, and 3.7% on CUB200-2011. Xiawu Zheng, Rongrong Ji, Xiaoshuai Sun, Yongjian Wu 0001, Feiyue Huang, Yanhua Yang |
IJCAI | 5 |
| 2018 | Dense Auto-Encoder Hashing for Robust Cross-Modality RetrievalabstractCross-modality retrieval has been widely studied, which aims to search images as response to text queries or vice versa. When faced with large-scale dataset, cross-modality hashing serves as an efficient and effective solution, which learns binary codes to approximate the cross-modality similarity in the Hamming space. Most recent cross-modality hashing schemes focus on learning the hash functions from data instances with fully modalities. However, how to learn robust binary codes when facing incomplete modality (i.e., with one modality missed or partially observed), is left unexploited, which however widely occurs in real-world applications. In this paper, we propose a novel cross-modality hashing, termed Dense Auto-encoder Hashing (DAH), which can explicitly impute the missed modality and produce robust binary codes by leveraging the relatedness among different modalities. To that effect, we propose a novel Dense Auto-encoder Network (DAN) to impute the missing modalities, which densely connects each layer to every other layer in a feed-forward fashion. For each layer, a noisy auto-encoder block is designed to calculate the residue between the current prediction and original data. Finally, a hash-layer is added to the end of DAN, which serves as a special binary encoder model to deal with the incomplete modality input. Quantitative experiments on three cross-modality visual search benchmarks, i.e., the Wiki, NUS-WIDE, and FLICKR-25K, have shown that the proposed DAH has superior performance over the state-of-the-art approaches. Hong Liu 0009, Mingbao Lin, Shengchuan Zhang, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
ACM Multimedia | 5 |
| 2018 | Attention-based Multi-Patch Aggregation for Image Aesthetic AssessmentabstractAggregation structures with explicit information, such as image attributes and scene semantics, are effective and popular for intelligent systems for assessing aesthetics of visual data. However, useful information may not be available due to the high cost of manual annotation and expert design. In this paper, we present a novel multi-patch (MP) aggregation method for image aesthetic assessment. Different from state-of-the-art methods, which augment an MP aggregation network with various visual attributes, we train the model in an end-to-end manner with aesthetic labels only (i.e., aesthetically positive or negative). We achieve the goal by resorting to an attention-based mechanism that adaptively adjusts the weight of each patch during the training process to improve learning efficiency. In addition, we propose a set of objectives with three typical attention mechanisms (i.e., average, minimum, and adaptive) and evaluate their effectiveness on the Aesthetic Visual Analysis (AVA) benchmark. Numerical results show that our approach outperforms existing methods by a large margin. We further verify the effectiveness of the proposed attention-based objectives via ablation studies and shed light on the design of aesthetic assessment systems. Kekai Sheng, Weiming Dong, Chongyang Ma, Xing Mei, Feiyue Huang, Bao-Gang Hu |
ACM Multimedia | 5 |
| 2018 | Supervised latent Dirichlet allocation with a mixture of sparse softmax
Zhanyu Ma, Feiyue Huang, Xiaojie Wang 0006, Jun Guo 0002 |
Neurocomputing | 5 |
| 2018 | Corrigendum to "Supervised latent Dirichlet allocation with a mixture of sparse softmax" [Neurocomputing, volume 312, 27 October 2018, Pages 324-335]
Zhanyu Ma, Feiyue Huang, Xiaojie Wang 0006, Jun Guo 0002 |
Neurocomputing | 5 |
| 2018 | Unsupervised Domain Adaptation for Face Anti-SpoofingabstractFace anti-spoofing (a.k.a. presentation attack detection) has recently emerged as an active topic with great significance for both academia and industry due to the rapidly increasing demand in user authentication on mobile phones, PCs, tablets, and so on. Recently, numerous face spoofing detection schemes have been proposed based on the assumption that training and testing samples are in the same domain in terms of the feature space and marginal probability distribution. However, due to unlimited variations of the dominant conditions (illumination, facial appearance, camera quality, and so on) in face acquisition, such single domain methods lack generalization capability, which further prevents them from being applied in practical applications. In light of this, we introduce an unsupervised domain adaptation face anti-spoofing scheme to address the real-world scenario that learns the classifier for the target domain based on training samples in a different source domain. In particular, an embedding function is first imposed based on source and target domain data, which maps the data to a new space where the distribution similarity can be measured. Subsequently, the Maximum Mean Discrepancy between the latent features in source and target domains is minimized such that a more generalized classifier can be learned. State-of-the-art representations including both hand-crafted and deep neural network learned features are further adopted into the framework to quest the capability of them in domain adaptation. Moreover, we introduce a new database for face spoofing detection, which contains more than 4000 face samples with a large variety of spoofing types, capture devices, illuminations, and so on. Extensive experiments on existing benchmark databases and the new database verify that the proposed approach can gain significantly better generalization capability in cross-domain scenarios by providing consistently better anti-spoofing performance. Haoliang Li, Wen Li 0001, Shiqi Wang 0001, Feiyue Huang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2018 | Animated Construction of Chinese Brush PaintingsabstractIn this paper, we present a method for reconstructing the drawing process of Chinese brush paintings. We demonstrate the possibility of computing an artistically reasonable drawing order from a static brush painting that is consistent with the rules of art. We map the key principles of drawing composition to our computational framework, which first organizes the strokes in three stages and then optimizes stroke ordering with natural evolution strategies. Our system produces reasonable animated constructions of Chinese brush paintings with minimal or no user intervention. We test our algorithm on a range of input paintings with varying degrees of complexity and structure and then evaluate the results via a user study. We discuss the applications of the proposed system to painting instruction, painting animation, and image stylization, especially in the context of art teaching. Fan Tang, Weiming Dong, Yiping Meng, Xing Mei, Feiyue Huang, Xiaopeng Zhang 0001, Oliver Deussen |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2017 | ESPACE: Accelerating Convolutional Neural Networks via Eliminating Spatial and Channel RedundancyabstractRecent years have witnessed an extensive popularity of convolutional neural networks (CNNs) in various computer vision and artificial intelligence applications. However, the performance gains have come at a cost of substantially intensive computation complexity, which prohibits its usage inresource-limited applications like mobile or embedded devices. While increasing attention has been paid to the acceleration of internal network structure, the redundancy of visual input is rarely considered. In this paper, we make the first attempt of reducing spatial and channel redundancy directly from the visual input for CNNs acceleration. The proposed method, termed ESPACE (Elimination of SPAtial and Channel rEdundancy), works by the following three steps: First, the 3D channel redundancy of convolutional layers is reduced by a set of low-rank approximation of convolutional filters. Second, a novel mask based selective processing scheme is proposed, which further speedups the convolution operations via skipping unsalient spatial locations of the visual input. Third, the accelerated network is fine-tuned using the training data via back-propagation. The proposed method is evaluated on ImageNet 2012 with implementations on two widely adopted CNNs, i.e. AlexNet and GoogLeNet. In comparison to several recent methods of CNN acceleration, the proposed scheme has demonstrated new state-of-the-art acceleration performance by a factor of 5.48* and 4.12* speedup on AlexNet and GoogLeNet, respectively, with a minimal decrease in classification accuracy. Shaohui Lin, Rongrong Ji, Chao Chen 0026, Feiyue Huang |
AAAI | 4 |
| 2017 | Ordinal Constrained Binary Code Learning for Nearest Neighbor SearchabstractRecent years have witnessed extensive attention in binary code learning, a.k.a. hashing, for nearest neighbor search problems. It has been seen that high-dimensional data points can quantize into binary codes to give an efficient similarity approximation via Hamming distance. Among the existing schemes, ranking-based hashing is recent promising that targets at preserving ordinal relations of ranking in the Hamming space to minimize retrieval loss. However, the size of the ranking tuples that show the ordinal relations, is quadratic or cubic to the size of training samples. It is so very expensive to embed such ranking tuples in binary code learning, especially given a large-scale training data set. Besides, it remains difficult to build ranking tuples efficiently for most ranking-preserving hashing, which are deployed over an ordinal graph-based setting. To handle these problems, we propose a novel ranking-preserving hashing method, dubbed Ordinal Constraint Hashing (OCH), which efficiently learns the optimal hashing functions with a graph-based approximation to embed the ordinal relations. The core idea is to reduce the size of ordinal graph with ordinal constraint projection, which preserves the ordinal relations through a small data set (such as clusters or random samples). In particular, to learn such hash functions effectively, we further relax the discrete constraints and design a specific stochastic gradient decent algorithm for optimization. Experimental results on three large-scale visual search benchmark datasets, i.e. LabelMe, Tiny100K and GIST1M, show that the proposed OCH method can achieve superior performance over the state-of-the-arts approaches. Hong Liu 0009, Rongrong Ji, Yongjian Wu 0001, Feiyue Huang |
AAAI | 4 |
| 2017 | Cross-Modality Binary Code Learning via Fusion Similarity HashingabstractBinary code learning has been emerging topic in large-scale cross-modality retrieval recently. It aims to map features from multiple modalities into a common Hamming space, where the cross-modality similarity can be approximated efficiently via Hamming distance. To this end, most existing works learn binary codes directly from data instances in multiple modalities, which preserve both intra-and inter-modal similarities respectively. Few methods consider to preserve the fusion similarity among multi-modal instances instead, which can explicitly capture their heterogeneous correlation in cross-modality retrieval. In this paper, we propose a hashing scheme, termed Fusion Similarity Hashing (FSH), which explicitly embeds the graph-based fusion similarity across modalities into a common Hamming space. Inspired by the fusion by diffusion, our core idea is to construct an undirected asymmetric graph to model the fusion similarity among different modalities, upon which a graph hashing scheme with alternating optimization is introduced to learn binary codes that embeds such fusion similarity. Quantitative evaluations on three widely used benchmarks, i.e., UCI Handwritten Digit, MIR-Flickr25K and NUS-WIDE, demonstrate that the proposed FSH approach can achieve superior performance over the state-of-the-art methods. Hong Liu 0009, Rongrong Ji, Yongjian Wu 0001, Feiyue Huang, Baochang Zhang 0001 |
CVPR | 4 |
| 2017 | Fusing transcription results from polyphonic and monophonic audio for singing melody transcription in polyphonic musicabstractThis paper presents a new system for singing melody transcription from polyphonic songs. Instead of operating solely on polyphonic audio of each song to be processed (as most existing systems do), our system takes as inputs additionally multiple monophonic recordings of people singing the song. To transcribe the singing melody in a song, our system first tracks the singing pitch from polyphonic audio of the song by using a deep neural network (DNN)-based method, and then uses the estimated pitch series as reference to select the pitch sequences extracted from the multiple monophonic singing recordings. The selected monophonic pitch sequences, as well as the DNN pitch series from the polyphonic audio, are then transcribed separately, and their transcriptions results are fused to form the final note sequence. Experimental results show that, by introducing monophonic singings into transcription, the performance of singing melody transcription from polyphonic songs can be significantly improved. Bilei Zhu, Fuzhang Wu, Ke Li 0015, Yongjian Wu 0001, Feiyue Huang, Yunsheng Wu |
ICASSP | 5 |
| 2017 | Centroid-aware local discriminative metric learning in speaker verification
Kekai Sheng, Weiming Dong, Joseph Razik, Feiyue Huang, Bao-Gang Hu |
Pattern Recognit. | 5 |
| 2017 | Toward Optimal Manifold Hashing via Discrete Locally Linear EmbeddingabstractBinary code learning, also known as hashing, has received increasing attention in large-scale visual search. By transforming high-dimensional features to binary codes, the original Euclidean distance is approximated via Hamming distance. More recently, it is advocated that it is the manifold distance, rather than the Euclidean distance, that should be preserved in the Hamming space. However, it retains as an open problem to directly preserve the manifold structure by hashing. In particular, it first needs to build the local linear embedding in the original feature space, and then quantize such embedding to binary codes. Such a two-step coding is problematic and less optimized. Besides, the off-line learning is extremely time and memory consuming, which needs to calculate the similarity matrix of the original data. In this paper, we propose a novel hashing algorithm, termed discrete locality linear embedding hashing (DLLH), which well addresses the above challenges. The DLLH directly reconstructs the manifold structure in the Hamming space, which learns optimal hash codes to maintain the local linear relationship of data points. To learn discrete locally linear embeddingcodes, we further propose a discrete optimization algorithm with an iterative parameters updating scheme. Moreover, an anchor-based acceleration scheme, termed Anchor-DLLH, is further introduced, which approximates the large similarity matrix by the product of two low-rank matrices. Experimental results on three widely used benchmark data sets, i.e., CIFAR10, NUS-WIDE, and YouTube Face, have shown superior performance of the proposed DLLH over the state-of-the-art approaches. Rongrong Ji, Hong Liu 0009, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang |
IEEE Trans. Image Process. | 6 |
| 2017 | Data-Driven Synthesis of Cartoon Faces Using Different StylesabstractThis paper presents a data-driven approach for automatically generating cartoon faces in different styles from a given portrait image. Our stylization pipeline consists of two steps: an offline analysis step to learn about how to select and compose facial components from the databases; a runtime synthesis step to generate the cartoon face by assembling parts from a database of stylized facial components. We propose an optimization framework that, for a given artistic style, simultaneously considers the desired image-cartoon relationships of the facial components and a proper adjustment of the image composition. We measure the similarity between facial components of the input image and our cartoon database via image feature matching, and introduce a probabilistic framework for modeling the relationships between cartoon facial components. We incorporate prior knowledge about image-cartoon relationships and the optimal composition of facial components extracted from a set of cartoon faces to maintain a natural, consistent, and attractive look of the results. We demonstrate generality and robustness of our approach by applying it to a variety of portrait images and compare our output with stylized results created by artists via a comprehensive user study. Yong Zhang 0034, Weiming Dong, Chongyang Ma, Xing Mei, Ke Li 0015, Feiyue Huang, Bao-Gang Hu, Oliver Deussen |
IEEE Trans. Image Process. | 6 |
| 2016 | Measuring and Predicting Visual Importance of Similar ObjectsabstractSimilar objects are ubiquitous and abundant in both natural and artificial scenes. Determining the visual importance of several similar objects in a complex photograph is a challenge for image understanding algorithms. This study aims to define the importance of similar objects in an image and to develop a method that can select the most important instances for an input image from multiple similar objects. This task is challenging because multiple objects must be compared without adequate semantic information. This challenge is addressed by building an image database and designing an interactive system to measure object importance from human observers. This ground truth is used to define a range of features related to the visual importance of similar objects. Then, these features are used in learning-to-rank and random forest to rank similar objects in an image. Importance predictions were validated on 5,922 objects. The most important objects can be identified automatically. The factors related to composition (e.g., size, location, and overlap) are particularly informative, although clarity and color contrast are also important. We demonstrate the usefulness of similar object importance on various applications, including image retargeting, image compression, image re-attentionizing, image admixture, and manipulation of blindness images. Yan Kong, Weiming Dong, Xing Mei, Chongyang Ma, Tong-Yee Lee, Siwei Lyu, Feiyue Huang, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2016 | Similarity metric learning for face verification using sigmoid decision function
Xiao-Nan Hou, Shouhong Ding, Lizhuang Ma, Chengjie Wang 0001, Feiyue Huang |
Vis. Comput. | 6 |
| 2015 | Automatic script identification in the wildabstractWith the rapid increase of transnational communication and cooperation, people frequently encounter multilingual scenarios in various situations. In this paper, we are concerned with a relatively new problem: script identification at word or line levels in natural scenes. A large-scale dataset with a great quantity of natural images and 10 types of widely-used languages is constructed and released. In allusion to the challenges in script identification in real-world scenarios, a deep learning based algorithm is proposed. The experiments on the proposed dataset demonstrate that our algorithm achieves superior performance, compared with conventional image classification or script identification methods, including as the original CNN architecture, LLC and GLCM. Baoguang Shi, Cong Yao, Chengquan Zhang, Feiyue Huang, Xiang Bai |
ICDAR | 5 |
| 2015 | Evaluating the Quality of Face Alignment without Ground TruthabstractThe study of face alignment has been an area of intense research in computer vision, with its achievements widely used in computer graphics applications. The performance of various face alignment methods is often image-dependent or somewhat random because of their own strategy. This study aims to develop a method that can select an input image with good face alignment results from many results produced by a single method or multiple ones. The task is challenging because different face alignment results need to be evaluated without any ground truth. This study addresses this problem by designing a feasible feature extraction scheme to measure the quality of face alignment results. The feature is then used in various machine learning algorithms to rank different face alignment results. Our experiments show that our method is promising for ranking face alignment results and is able to pick good face alignment results, which can enhance the overall performance of a face alignment method with a random strategy. We demonstrate the usefulness of our ranking-enhanced face alignment algorithm in two practical applications: face cartoon stylization and digital face makeup. Kekai Sheng, Weiming Dong, Yan Kong, Xing Mei, Chengjie Wang 0001, Feiyue Huang, Bao-Gang Hu |
Comput. Graph. Forum | 7 |
| 2012 | A Novel Customized Recompression Framework for Massive Internet Images
Shouhong Ding, Feiyue Huang, Yongjian Wu 0001, Lizhuang Ma |
CVM | 2 |
| 2012 | Accurate Depth-of-Field Rendering Using Adaptive Bilateral Depth Filtering
Shang Wu 0001, Bin Sheng 0001, Feiyue Huang, Lizhuang Ma |
CVM | 4 |
| 2012 | A Customized Framework to Recompress Massive Internet Images
Shouhong Ding, Feiyue Huang, Yongjian Wu 0001, Bin Sheng 0001, Lizhuang Ma |
J. Comput. Sci. Technol. | 2 |
| 2010 | View and scale insensitive action representation and recognitionabstractIn this paper a view and scale insensitive action representation VSI-Surf is proposed. Scale invariant shape descriptor R-transform is used to extract compact 1D feature from view insensitive posture representation “Envelop shape” which uses only two orthogonal cameras without accurate calibration. Considering action is a posture sequence, to integrate temporal information, 1D posture feature is then extended in time dimension. Then we get an action representation insensitive to viewpoint and scale, which is called VSI-Surf. Actions recognition is processed in a hierarchical framework, in which body actions and gestures are recognized in different level. Encouraging recognition results have been demonstrated on the multi-view IXMAS action dataset. Yuanyuan Cao, Feiyue Huang, Linmi Tao, Guangyou Xu |
ICASSP | 2 |
| 2007 | Viewpoint Insensitive Action Recognition Using Envelop Shape
Feiyue Huang, Guangyou Xu |
ACCV (2) | 1 |
| 2004 | A multi-stream audio-video large-vocabulary Mandarin Chinese speech databaseabstractWe present the acquisition and content of a multi-stream audio-visual large-vocabulary database in Mandarin Chinese. The database consists of 17,000 utterances spoken by 225 people and captured by a set of seven cameras and 12 microphones. We also provide the label files that describe the endpoints of the utterances and the script files that represent the actual pronunciation of speech. The database can be used in audio-visual speech recognition (AVSR) for both large-vocabulary and small tasks, microphone array based speech recognition, audio-visual speaker identification and 3D face modeling. Luhong Liang, Feiyue Huang, Ara V. Nefian |
ICME | 3 |