Yunpeng Chen

dblp:156/7824 · DBLP profile ↗
← Back
47ranked-venue papers
9as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 28 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STFNet: A Knowledge-Guided Spatial-Temporal Fusion Network for Low-SNR Modulation Recognition
abstract
Automatic modulation classification (AMC) in complex electromagnetic environments is essential for ensuring reliable spectrum surveillance. However, most existing AMC algorithms primarily rely on the data engineer and focus on improving accuracy under high signal-to-noise ratio (SNR) conditions, making it difficult to maintain robust and accurate performance in real-world unstable SNR scenarios. In order to address this challenge, a knowledge-guided spatial-temporal fusion network for low-SNR modulation recognition is proposed, named STFNet. Firstly, aiming at the instability of the model classification caused by the single input mode, the enhanced constellation diagram and the I/Q signal are simultaneously introduced as the inputs of the STFNet. Furthermore, the STFNet is designed as a dual-path adaptive multimodal fusion architecture to simultaneously exploit the spatial features of the enhanced constellation diagram and the temporal features of the I/Q signal. Secondly, an AMC-specific transfer learning (AMC-TL) strategy is introduced to enhance the global robustness of spatial representations through contrastive learning. More importantly, a domain knowledge-guided mixture of experts (DKG-MoE) is proposed to incorporate traditional features into the expert routing process, which improves temporal recognition accuracy. Then, an adaptive modality attention (AMA) module is developed to balance the accuracy under high SNR and the robustness under low SNR. Finally, extensive experiments on four benchmark datasets demonstrate that the proposed method achieves a 4.13% accuracy improvement under low SNR conditions (-20 dB ∼ 0 dB) and outperforms all state-of-the-art (SOTA) methods in overall average accuracy (65.57%). The code is publicly available at https://github.com/yoho78/STFNet.
Hongyu Wei, Hanqian Mo, Yunpeng Chen, Pingfan Wu, Yaxin Peng, Hao Kong 0004
IEEE Internet Things J.3
2025 CharaConsist: Fine-Grained Consistent Character Generation
abstract
In text-to-image generation, producing a series of consistent contents that preserve the same identity is highly valuable for real-world applications. Although a few works have explored training-free methods to enhance the consistency of generated subjects, we observe that they suffer from the following problems. First, they fail to maintain consistent background details, which limits their applicability. Furthermore, when the foreground character undergoes large motion variations, inconsistencies in identity and clothing details become evident. To address these problems, we propose CharaConsist, which employs point-tracking attention and adaptive token merge along with decoupled control of the foreground and background. CharaConsist enables fine-grained consistency for both foreground and background, supporting the generation of one character in continuous shots within a fixed scene or in discrete shots across different scenes. Moreover, CharaConsist is the first consistent generation method tailored for text-to-image DiT model. Its ability to maintain fine-grained consistency, combined with the larger capacity of latest base model, enables it to produce high-quality visual outputs, broadening its applicability to a wider range of real-world scenarios. The source code has been released at https://github.com/Murray-Wang/CharaConsist
Mengyu Wang 0003, Henghui Ding, Jianing Peng, Yao Zhao 0001, Yunpeng Chen, Yunchao Wei
ICCV5
2025 An Open Dataset of Cyber Asset Graphs for Cybercrime Research
abstract
Cybercrime poses a severe threat to the entire Internet ecosystem. Various cyber assets, such as domain name, IP address, and security certificate, are staple infrastructures of cybercrime. A cyber asset graph (CAG) is a collection of closely related cyber assets held by a cybercrime gang to support online criminal activities. Analyzing CAGs provides rich data insights for cybercrime investigation and governance. This paper introduces an open dataset of CAGs comprised of 2.37 million nodes with eight types of cyber assets and 3.28 million edges with eleven types of relations. This paper introduces the dataset construction process, applied areas, and the experience of using the dataset in the ChinaVis Data Challenge 2022. This dataset contains numerous CAGs of cybercrime gangs in the real world, which is the first open dataset of CAGs for cybercrime research. This dataset can also support the development of other application-oriented areas, such as cyber asset management and cyber-physical-social system, and various graph-related research areas, such as graph theory, graph mining, and graph visualization.
Xin Zhao 0025, Shaolong Li, Ying Zhao 0001, Shuowen Fu, Yunpeng Chen, Zhuo Chen 0029
IEEE Trans. Big Data5
2025 Investigating Visual Perception of Degree Centrality in Graph Visualization
abstract
Degree centrality (DC) is a widely used metric that measures node importance in data space. A node-link diagram is a commonly used graph visualization to help viewers identify important nodes in visual space. Previous graph perception studies largely concentrated on revealing perception principles in visual space. However, they rarely investigated the intrinsic relations between computed and perceived important nodes by jointly using data and visual spaces, thereby hindering a deep integration of computational and interactive graph analytics. To address this gap, we adopted the visual perception of DC as a representative object to conduct a graph perception study by jointly using data and visual spaces. Two research questions were defined. (RQ1) Can viewers accurately estimate the relative DCs of the given nodes in a node-link diagram through visual perception? (RQ2) What visual factors influence viewers' visual estimation of relative DCs? A controlled user experiment was conducted to answer the questions. Results showed that: (1) The participants failed to estimate the relative DCs accurately, particularly when the DC differences between nodes were not great. (2) Seven visual factors influencing the tasks were summarized, such as the size of the visual receptive region of a node, the link and node densities in the visual receptive region of the node, and the wrapping angle of the node's neighbors. (3) The factors presented certain priorities in complicated situations. These findings provide rich implications for graph analytics, such as utilizing the findings to optimize graph visualizations to achieve the desired consistency between computed and perceived important nodes.
Xin Zhao 0025, Shuowen Fu, Yunpeng Chen, Ying Zhao 0001
IEEE Trans. Vis. Comput. Graph.5
2024 Visual Analysis of Money Laundering in Cryptocurrency Exchange
abstract
Blockchain-based cryptocurrencies, such as Bitcoin (BTC) and Ethereum (ETH), are newly emerging financial assets. Cryptocurrency exchanges are marketplaces for cryptocurrency circulation while becoming a new venue for money laundering. In this work, we cooperate with a cryptocurrency exchange to investigate new solutions for anti-money laundering in cryptocurrency exchanges. First, we learn the domain knowledge of cryptocurrency transactions and summarize data analytical requirements of transaction supervisors in their daily work of anti-money laundering. Then, we propose a visual analysis approach to support their daily work. The approach consists of a new algorithm that automatically detects suspicious money laundering accounts and a multiviewed user interface that visualizes the algorithm results and relevant transaction data. An abacus-inspired visualization is designed in the interface to depict transaction patterns contained in numerous cryptocurrency transactions, which can help supervisors find money laundering clues and deduce the trading tactic adopted by launderers. Finally, an algorithm performance experiment, a case study, and a field study are conducted with real-world data to demonstrate the effectiveness of our solution.
Yunpeng Chen, Chunyao Zhu, Lijia Jiang, Xincheng Liao, Zengsheng Zhong, Yi Chen 0007, Ying Zhao 0001
IEEE Trans. Comput. Soc. Syst.2
2024 Cross-Domain Few-Shot Classification via Dense-Sparse-Dense Regularization
abstract
This work addresses the problem of cross-domain few-shot classification which aims at recognizing novel categories in unseen domains with only a few labeled data samples. We think that the pre-trained model contains the redundant elements which are useless or even harmful for the downstream tasks. To remedy the drawback, we introduce an$L^{2}$-SP regularized dense-sparse-dense (DSD) fine-tuning flow for regularizing the capacity of pre-trained networks and achieving efficient few-shot domain adaptation. Given a pre-trained model from the source domain, we start by carrying out a conventional dense fine-tuning step using the target data. Then we execute a sparse pruning step that prunes the unimportant connections and fine-tunes the weights of sub-network. Finally, initialized with the fine-tuned sub-network, we retrain the original dense network as the output model for the target domain. The whole fine-tuning procedure is regularized by an$L^{2}$-SP term. In contrast to the existing methods that either tune the weights or prune the network structure for domain adaptation, our regularized DSD fine-tuning flow simultaneously exploits the benefits of sparsity regularity and dense network capacity to gain the best of both worlds. Our method can be applied in a plug-and-play manner to improve the existing fine-tuning methods. Extensive experimental results on benchmark datasets demonstrate that our method in many cases outperforms the existing cross-domain few-shot classification methods in significant margins. Our code will be released soon.
Fanfan Ji, Yunpeng Chen, Luoqi Liu, Xiao-Tong Yuan
IEEE Trans. Circuits Syst. Video Technol.2
2024 An Efficient Ring Oscillator PUF Using Programmable Delay Units on FPGA
abstract
The ring oscillator (RO) PUF can be implemented on different FPGA platforms with high uniqueness and reliability. To decrease the hardware cost of conventional RO PUFs, a new design using the programmable delay units is proposed, namely, PRO PUF. The programmable interconnect points (PIPs) of programmable delay units are used to enhance the configurability. The PUF cell of the proposed design has the ability to be efficiently programmed to an RO PUF at any stage by adjusting the propagation paths of the delay units. A significant number of responses can be generated by the proposed PRO PUF while consuming fewer hardware resources. To verify the performance, the proposed design has been implemented on Xilinx FPGAs and also simulated using a standard 40nm technology. The experimental results have shown that the proposed design achieves high uniqueness, reliability, and hardware efficiency. Moreover, the PRO PUF has been evaluated using a machine learning attack, the CMA-ES attack. The results have shown that the proposed structure is more resistant to common modeling attacks when compared to conventional RO-related PUF designs.
Yijun Cui, Jiang Li 0012, Yunpeng Chen, Chenghua Wang, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001
ACM Trans. Design Autom. Electr. Syst.3
2024 An open dataset of data lineage graphs for data governance research
abstract
Data have become valuable assets for enterprises. Data governance aims to manage and reuse data assets to facilitate enterprise management and product innovations. A data lineage graph (DLG) is an abstracted collection of data assets and their data lineages in data governance. Analyzing DLGs can provide rich data insights for data governance. However, the progress of data governance technologies is hindered by the shortage of available open datasets for DLGs. This paper introduces an open dataset of DLGs, including the DLG model, the dataset construction process, and applied areas. This real-world dataset is sourced from Huawei Cloud Computing Technology Company Limited, which contains 18 DLGs with three types of data assets and two types of relations. To the best of our knowledge, this dataset is the first open dataset of DLGs for data governance. This dataset can also support the development of other application areas, such as graph analytics and visualization.
Yunpeng Chen, Ying Zhao 0001, Xuanjing Li
Vis. Informatics1
2024 Corrigendum to "An open dataset of data lineage graphs for data governance research" [Vis. Inform. 8 (1) (2024) 1-5]
Yunpeng Chen, Ying Zhao 0001, Xuanjing Li
Vis. Informatics1
2023 Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models
abstract
Public large-scale text-to-image diffusion models, such as Stable Diffusion, have gained significant attention from the community. These models can be easily customized for new concepts using low-rank adaptations (LoRAs). However, the utilization of multiple-concept LoRAs to jointly support multiple customized concepts presents a challenge. We refer to this scenario as decentralized multi-concept customization, which involves single-client concept tuning and center-node concept fusion. In this paper, we propose a new framework called Mix-of-Show that addresses the challenges of decentralized multi-concept customization, including concept conflicts resulting from existing single-client LoRA tuning and identity loss during model fusion. Mix-of-Show adopts an embedding-decomposed LoRA (ED-LoRA) for single-client tuning and gradient fusion for the center node to preserve the in-domain essence of single concepts and support theoretically limitless concept fusion. Additionally, we introduce regionally controllable sampling, which extends spatially controllable sampling (e.g., ControlNet and T2I-Adapter) to address attribute binding and missing object problems in multi-concept sampling. Extensive experiments demonstrate that Mix-of-Show is capable of composing multiple customized concepts with high fidelity, including characters, objects, and scenes.
Yuchao Gu, Xintao Wang 0002, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao 0001, Shuning Chang, Weijia Wu 0001, Yixiao Ge, Ying Shan, Zheng Shou 0001
NeurIPS5
2023 Visual abstraction of dynamic network via improved multi-class blue noise sampling
Yanni Peng, Xiaoping Fan, Ziyao Yu, Yunpeng Chen, Ying Zhao 0001
Frontiers Comput. Sci.6
2022 Distribution-Aware Single-Stage Models for Multi-Person 3D Pose Estimation
abstract
In this paper, we present a novel Distribution-Aware Single-stage (DAS) model for tackling the challenging multi-person 3D pose estimation problem. Different from existing top-down and bottom-up methods, the proposed DAS model simultaneously localizes person positions and their corresponding body joints in the 3D camera space in a one-pass manner. This leads to a simplified pipeline with enhanced efficiency. In addition, DAS learns the true distribution of body joints for the regression of their positions, rather than making a simple Laplacian or Gaussian assumption as previous works. This provides valuable priors for model prediction and thus boosts the regression-based scheme to achieve competitive performance with volumetric-base ones. Moreover, DAS exploits a recur-sive update strategy for progressively approaching to regression target, alleviating the optimization difficulty and further lifting the regression performance. DAS is implemented with a fully Convolutional Neural Network and end-to-end learnable. Comprehensive experiments on benchmarks CMU Panoptic and MuPoTS-3D demonstrate the superior efficiency of the proposed DAS model, specifically 1.5x speedup over previous best model, and its stat-of-the-art accuracy for multi-person 3D pose estimation.
Zitian Wang, Xuecheng Nie, Xiaochao Qu, Yunpeng Chen, Si Liu 0001
CVPR4
2022 MorphMLP: An Efficient MLP-Like Backbone for Spatial-Temporal Representation Learning
Junhao Zhang 0001, Kunchang Li 0002, Yali Wang 0001, Yunpeng Chen, Shashwat Chandra, Yu Qiao 0001, Luoqi Liu, Zheng Shou 0001
ECCV (35)4
2022 SODAR: Exploring Locally Aggregated Learning of Mask Representations for Instance Segmentation
abstract
Recent state-of-the-art one-stage instance segmentation model SOLO divides the input image into a grid and directly predicts per grid cell object masks with fully-convolutional networks, yielding comparably good performance as traditional two-stage Mask R-CNN yet enjoying much simpler architecture and higher efficiency. We observe SOLO generates similar masks for an object at nearby grid cells, and these neighboring predictions can complement each other as some may better segment certain object part, most of which are however directly discarded by non-maximum-suppression. Motivated by the observed gap, we develop a novel learning-based aggregation method that improves upon SOLO by leveraging the rich neighboring information while maintaining the architectural efficiency. The resulting model is named SODAR. Unlike the original per grid cell object masks, SODAR is implicitly supervised to learn mask representations that encode geometric structure of nearby objects and complement adjacent representations with context. The aggregation method further includes two novel designs: 1) a mask interpolation mechanism that enables the model to generate much fewer mask representations by sharing neighboring representations among nearby grid cells, and thus saves computation and memory; 2) a deformable neighbour sampling mechanism that allows the model to adaptively adjust neighbor sampling locations thus gathering mask representations with more relevant context and achieving higher performance. SODAR significantly improves the instance segmentation performance, e.g., it outperforms a SOLO model with ResNet-101 backbone by 2.2 AP on COCO test set, with only about 3% additional computation. We further show consistent performance gain with the SOLOv2 model.
Tao Wang 0053, Jun Hao Liew, Yu Li 0016, Yunpeng Chen, Jiashi Feng
IEEE Trans. Image Process.4
2021 Continual Learning via Bit-Level Information Preserving
abstract
Continual learning tackles the setting of learning different tasks sequentially. Despite the lots of previous solutions, most of them still suffer significant forgetting or expensive memory cost. In this work, targeted at these problems, we first study the continual learning process through the lens of information theory and observe that forgetting of a model stems from the loss of information gain on its parameters from the previous tasks when learning a new task. From this viewpoint, we then propose a novel continual learning approach called Bit-Level Information Preserving (BLIP) that preserves the information gain on model parameters through updating the parameters at the bit level, which can be conveniently implemented with parameter quantization. More specifically, BLIP first trains a neural network with weight quantization on the new incoming task and then estimates information gain on each parameter provided by the task data to determine the bits to be frozen to prevent forgetting. We conduct extensive experiments ranging from classification tasks to reinforcement learning tasks, and the results show that our method produces better or on par results comparing to previous state-of-the-arts. Indeed, BLIP achieves close to zero forgetting while only requiring constant memory overheads throughout continual learning1.
Yujun Shi, Li Yuan 0007, Yunpeng Chen, Jiashi Feng
CVPR3
2021 Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
abstract
Transformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification. The ViT model splits each image into a sequence of tokens with fixed length and then applies multiple Transformer layers to model their global relation for classification. However, ViT achieves inferior performance to CNNs when trained from scratch on a midsize dataset like ImageNet. We find it is because: 1) the simple tokenization of input images fails to model the important local structure such as edges and lines among neighboring pixels, leading to low training sample efficiency; 2) the redundant attention backbone design of ViT leads to limited feature richness for fixed computation budgets and limited training samples. To overcome such limitations, we propose a new Tokens-To-Token Vision Transformer (T2T-VTT), which incorporates 1) a layer-wise Tokens-to-Token (T2T) transformation to progressively structurize the image to tokens by recursively aggregating neighboring Tokens into one Token (Tokens-to-Token), such that local structure represented by surrounding tokens can be modeled and tokens length can be reduced; 2) an efficient backbone with a deep-narrow structure for vision transformer motivated by CNN architecture design after empirical study. Notably, T2T-ViT reduces the parameter count and MACs of vanilla ViT by half, while achieving more than 3.0% improvement when trained from scratch on ImageNet. It also outperforms ResNets and achieves comparable performance with MobileNets by directly training on ImageNet. For example, T2T-ViT with comparable size to ResNet50 (21.5M parameters) can achieve 83.3% top1 accuracy in image resolution 384x384 on ImageNet.1
Li Yuan 0007, Yunpeng Chen, Tao Wang 0053, Weihao Yu 0001, Yujun Shi, Zihang Jiang, Francis E. H. Tay, Jiashi Feng, Shuicheng Yan
ICCV2
2021 PnP-DETR: Towards Efficient Visual Analysis with Transformers
abstract
Recently, DETR [3] pioneered the solution of vision tasks with transformers, it directly translates the image feature map into the object detection result. Though effective, translating the full feature map can be costly due to redundant computation on some area like the background. In this work, we encapsulate the idea of reducing spatial redundancy into a novel poll and pool (PnP) sampling module, with which we build an end-to-end PnP-DETR architecture that adaptively allocates its computation spatially to be more efficient. Concretely, the PnP module abstracts the image feature map into fine foreground object feature vectors and a small number of coarse background contextual feature vectors. The transformer models information interaction within the fine-coarse feature space and translates the features into the detection result. Moreover, the PnP-augmented model can instantly achieve various desired trade-offs between performance and computation with a single model by varying the sampled feature length, without requiring to train multiple models as existing methods. Thus it offers greater flexibility for deployment in diverse scenarios with varying computation constraint. We further validate the generalizability of the PnP module on panoptic segmentation and the recent transformer-based image recognition model ViT [7] and show consistent efficiency gain. We believe our method makes a step for efficient visual analysis with transformers, wherein spatial redundancy is commonly observed. Code and models will be available.
Tao Wang 0053, Li Yuan 0007, Yunpeng Chen, Jiashi Feng, Shuicheng Yan
ICCV3
2021 Dense Contrastive Visual-Linguistic Pretraining
abstract
Inspired by the success of BERT, several multimodal representation learning approaches have been proposed that jointly represent image and text. These approaches achieve superior performance by capturing high-level semantic information from large-scale multimodal pretraining. In particular, LXMERT and UNITER adopt visual region feature regression and label classification as pretext tasks. However, they tend to suffer from the problems of noisy labels and sparse semantic annotations, based on the visual features having been pretrained on a crowdsourced dataset with limited and inconsistent semantic labeling. To overcome these issues, we propose unbiased Dense Contrastive Visual-Linguistic Pretraining (DCVLP), which replaces the region regression and classification with cross-modality region contrastive learning that requires no annotations. Two data augmentation strategies (Mask Perturbation and Intra-Inter-Adversarial Perturbation) are developed to improve the quality of negative samples used in contrastive learning. Overall, DCVLP allows cross-modality dense region contrastive learning in a self-supervised setting independent of any object annotations. We compare our method against prior visual-linguistic pretraining frameworks to validate the superiority of dense contrastive learning on multimodal representation learning.
Lei Shi 0002, Kai Shuang, Shijie Geng, Peng Gao 0007, Zuohui Fu, Gerard de Melo, Yunpeng Chen, Sen Su
ACM Multimedia7
2021 Adversarial Domain Adaptation With Prototype-Based Normalized Output Conditioner
abstract
Domain adversarial training has become a prevailing and effective paradigm for unsupervised domain adaptation (UDA). To successfully align the multi-modal data structures across domains, the following works exploit discriminative information in the adversarial training process, e.g., using multiple class-wise discriminators and involving conditional information in the input or output of the domain discriminator. However, these methods either require non-trivial model designs or are inefficient for UDA tasks. In this work, we attempt to address this dilemma by devising simple and compact conditional domain adversarial training methods. We first revisit the simple concatenation conditioning strategy where features are concatenated with output predictions as the input of the discriminator. We find the concatenation strategy suffers from the weak conditioning strength. We further demonstrate that enlarging the norm of concatenated predictions can effectively energize the conditional domain alignment. Thus we improve concatenation conditioning by normalizing the output predictions to have the same norm of features, and term the derived method as Normalized OutpUt coNditioner (NOUN). However, conditioning on raw output predictions for domain alignment, NOUN suffers from inaccurate predictions of the target domain. To this end, we propose to condition the cross-domain feature alignment in the prototype space rather than in the output space. Combining the novel prototype-based conditioning with NOUN, we term the enhanced method as PROtotype-based Normalized OutpUt coNditioner (PRONOUN). Experiments on both object recognition and semantic segmentation show that NOUN can effectively align the multi-modal structures across domains and even outperform state-of-the-art domain adversarial training methods. Together with prototype-based conditioning, PRONOUN further improves the adaptation performance over NOUN on multiple object recognition benchmarks for UDA. Code is available at https://github.com/tim-learn/NOUN.
Dapeng Hu, Jian Liang 0001, Qibin Hou, Hanshu Yan, Yunpeng Chen
IEEE Trans. Image Process.5
2021 Visualization and visual analysis of vessel trajectory data: A survey
abstract
Maritime transports play a critical role in international trade and commerce. Massive vessels sailing around the world continuously generate vessel trajectory data that contain rich spatial–temporal patterns of vessel navigations. Analyzing and understanding these patterns are valuable for maritime traffic surveillance and management. As essential techniques in complex data analysis and understanding, visualization and visual analysis have been widely used in vessel trajectory data analysis. This paper presents a literature review on the visualization and visual analysis of vessel trajectory data. First, we introduce commonly used vessel trajectory data sets and summarize main operations in vessel trajectory data preprocessing. Then, we provide a taxonomy of visualization and visual analysis of vessel trajectory data based on existing approaches and introduce representative works in details. Finally, we expound on the prospects of the remaining challenges and directions for future research.
Yunpeng Chen, Ying Zhao 0001
Vis. Informatics5
2020 AdversarialNAS: Adversarial Neural Architecture Search for GANs
abstract
Neural Architecture Search (NAS) that aims to automate the procedure of architecture design has achieved promising results in many computer vision fields. In this paper, we propose an AdversarialNAS method specially tailored for Generative Adversarial Networks (GANs) to search for a superior generative model on the task of unconditional image generation. The AdversarialNAS is the first method that can search the architectures of generator and discriminator simultaneously in a differentiable manner. During searching, the designed adversarial search algorithm does not need to comput any extra metric to evaluate the performance of the searched architecture, and the search paradigm considers the relevance between the two network architectures and improves their mutual balance. Therefore, AdversarialNAS is very efficient and only takes 1 GPU day to search for a superior generative model in the proposed large search space. Experiments demonstrate the effectiveness and superiority of our method. The discovered generative model sets a new state-of-the-art FID score of 10.87 and highly competitive Inception Score of 8.74 on CIFAR-10. Its transferability is also proven by setting new state-of-the-art FID score of 26.98 and Inception score of 9.63 on STL-10. Code is at: https://github.com/chengaopro/AdversarialNAS.
Chen Gao 0005, Yunpeng Chen, Si Liu 0001, Zhenxiong Tan, Shuicheng Yan
CVPR2
2020 Highly Efficient Salient Object Detection with 100K Parameters
Shanghua Gao, Yong-Qiang Tan, Ming-Ming Cheng, Chengze Lu, Yunpeng Chen, Shuicheng Yan
ECCV (6)5
2020 Rethinking Bottleneck Structure for Efficient Mobile Network Design
Daquan Zhou, Qibin Hou, Yunpeng Chen, Jiashi Feng, Shuicheng Yan
ECCV (3)3
2020 Programmable Ring Oscillator PUF Based on Switch Matrix
abstract
Configurable ring oscillator (CRO) physical unclonable functions (PUFs) which can improve the uniqueness and reliability of conventional RO PUFs have been widely studied. Especially, the multiplier, XOR gate and tristate inverter based CRO PUFs can improve the uniqueness and reliability. However the efficiency is remain at the same level when compared with the conventional RO PUFs. In this paper, a programmable RO PUF (PRO PUF), which can be programmed to change the structure of a typical RO PUF, is proposed. The proposed PRO PUF design is implemented based on the switch matrix of an FPGA and can be programmed as a chained RO PUF or a random looped RO PUF. The proposed PRO PUF is implemented on Xilinx Spartan 6 FPGAs. Experimental results demonstrate that the proposed PRO PUF design has good uniqueness and reliability metrics as well as a high hardware efficiency.
Yijun Cui, Yunpeng Chen, Chenghua Wang, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001
ISCAS2
2020 Towards Accurate Human Pose Estimation in Videos of Crowded Scenes
abstract
Video-based human pose estimation in crowed scenes is a challenging problem due to occlusion, motion blur, scale variation and viewpoint change, etc. Prior approaches always fail to deal with this problem because of (1) lacking of usage of temporal information; (2) lacking of training data in crowded scenes. In this paper, we focus on improving human pose estimation in videos of crowded scenes from the perspectives of exploiting temporal context and collecting new data. In particular, we first follow the top-down strategy to detect persons and perform single-person pose estimation for each frame. Then, we refine the frame-based pose estimation with temporal contexts deriving from the optical-flow. Specifically, for one frame, we forward the historical poses from the previous frames and backward the future poses from the subsequent frames to current frame, leading to stable and accurate human pose estimation in videos. In addition, we mine new data of similar scenes to HIE dataset from the Internet for improving the diversity of training set. In this way, our model achieves best performance on 7 out of 13 videos and 56.33 average wAP on test dataset of HIE challenge.
Shuning Chang, Li Yuan 0007, Xuecheng Nie, Ziyuan Huang 0003, Yunpeng Chen, Jiashi Feng, Shuicheng Yan
ACM Multimedia6
2020 ConvBERT: Improving BERT with Span-based Dynamic Convolution
abstract
Pre-trained language models like BERT and its variants have recently achieved impressive performance in various natural language understanding tasks. However, BERT heavily relies on the global self-attention block and thus suffers large memory footprint and computation cost. Although all its attention heads query on the whole input sequence for generating the attention map from a global perspective, we observe some heads only need to learn local dependencies, which means existence of computation redundancy. We therefore propose a novel span-based dynamic convolution to replace these self-attention heads to directly model local dependencies. The novel convolution heads, together with the rest self-attention heads, form a new mixed attention block that is more efficient at both global and local context learning. We equip BERT with this mixed attention design and build a ConvBERT model. Experiments have shown that ConvBERT significantly outperforms BERT and its variants in various downstream tasks, with lower training cost and fewer model parameters. Remarkably, ConvBERTbase model achieves 86.4 GLUE score, 0.7 higher than ELECTRAbase, using less than 1/4 training cost. Code and pre-trained models will be released.
Zihang Jiang, Weihao Yu 0001, Daquan Zhou, Yunpeng Chen, Jiashi Feng, Shuicheng Yan
NeurIPS4
2020 Gcluster: a simple-to-use tool for visualizing and comparing genome contexts for numerous genomes
abstract
MOTIVATION: Comparing the organization of gene, gene clusters and their flanking genomic contexts is of critical importance to the determination of gene function and evolutionary basis of microbial traits. Currently, user-friendly and flexible tools enabling to visualize and compare genomic contexts for numerous genomes are still missing. RESULTS: We here present Gcluster, a stand-alone Perl tool that allows researchers to customize and create high-quality linear maps of the genomic region around the genes of interest across large numbers of completed and draft genomes. Importantly, Gcluster integrates homologous gene analysis, in the form of a built-in orthoMCL, and mapping genomes onto a given phylogeny to provide superior comparison of gene contexts. AVAILABILITY AND IMPLEMENTATION: Gcluster is written in Perl and released under GPLv3. The source code is freely available at https://github.com/Xiangyang1984/Gcluster and http://www.microbialgenomic.com/Gcluster_tool.html. Gcluster can also be installed through conda: 'conda install -c bioconda gcluster'. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yunpeng Chen
Bioinform.3
2019 Graph-Based Global Reasoning Networks
abstract
Globally modeling and reasoning over relations between regions can be beneficial for many computer vision tasks on both images and videos. Convolutional Neural Networks (CNNs) excel at modeling local relations by convolution operations, but they are typically inefficient at capturing global relations between distant regions and require stacking multiple convolution layers. In this work, we propose a new approach for reasoning globally in which a set of features are globally aggregated over the coordinate space and then projected to an interaction space where relational reasoning can be efficiently computed. After reasoning, relation-aware features are distributed back to the original coordinate space for down-stream tasks. We further present a highly efficient instantiation of the proposed approach and introduce the Global Reasoning unit (GloRe unit) that implements the coordinate-interaction space mapping by weighted global pooling and weighted broadcasting, and the relation reasoning via graph convolution on a small graph in interaction space. The proposed GloRe unit is lightweight, end-to-end trainable and can be easily plugged into existing CNNs for a wide range of tasks. Extensive experiments show our GloRe unit can consistently boost the performance of state-of-the-art backbone architectures, including ResNet, ResNeXt, SE-Net and DPN, for both 2D and 3D CNNs, on image classification, semantic segmentation and video action recognition task.
Yunpeng Chen, Marcus Rohrbach, Zhicheng Yan 0001, Shuicheng Yan, Jiashi Feng, Yannis Kalantidis
CVPR1
2019 Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave Convolution
abstract
In natural images, information is conveyed at different frequencies where higher frequencies are usually encoded with fine details and lower frequencies are usually encoded with global structures. Similarly, the output feature maps of a convolution layer can also be seen as a mixture of information at different frequencies. In this work, we propose to factorize the mixed feature maps by their frequencies, and design a novel Octave Convolution (OctConv) operation to store and process feature maps that vary spatially “slower” at a lower spatial resolution reducing both memory and computation cost. Unlike existing multi-scale methods, OctConv is formulated as a single, generic, plug-and-play convolutional unit that can be used as a direct replacement of (vanilla) convolutions without any adjustments in the network architecture. It is also orthogonal and complementary to methods that suggest better topologies or reduce channel-wise redundancy like group or depth-wise convolutions. We experimentally show that by simply replacing convolutions with OctConv, we can consistently boost accuracy for both image and video recognition tasks, while reducing memory and computational cost. An OctConv-equipped ResNet-152 can achieve 82.9% top-1 classification accuracy on ImageNet with merely 22.2 GFLOPs.
Yunpeng Chen, Haoqi Fan 0001, Zhicheng Yan 0001, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, Jiashi Feng
ICCV1
2019 Dynamic Feature Fusion for Semantic Edge Detection
abstract
Features from multiple scales can greatly benefit the semantic edge detection task if they are well fused. However, the prevalent semantic edge detection methods apply a fixed weight fusion strategy where images with different semantics are forced to share the same weights, resulting in universal fusion weights for all images and locations regardless of their different semantics or local context. In this work, we propose a novel dynamic feature fusion strategy that assigns different fusion weights for different input images and locations adaptively. This is achieved by a proposed weight learner to infer proper fusion weights over multi-level features for each location of the feature map, conditioned on the specific input. In this way, the heterogeneity in contributions made by different locations of feature maps and input images can be better considered and thus help produce more accurate and sharper edge predictions. We show that our model with the novel dynamic feature fusion is superior to fixed weight fusion and also the na\"ive location-invariant weight fusion methods, via comprehensive experiments on benchmarks Cityscapes and SBD. In particular, our method outperforms all existing well established methods and achieves new state-of-the-art.
Yuan Hu 0004, Yunpeng Chen, Xiang Li 0046, Jiashi Feng
IJCAI2
2019 Task Relation Networks
abstract
Multi-task learning is popular in machine learning and computer vision. In multitask learning, properly modeling task relations is important for boosting the performance of jointly learned tasks. Task covariance modeling has been successfully used to model the relations of tasks but is limited to homogeneous multi-task learning. In this paper, we propose a feature based task relation modeling approach, suitable for both homogeneous and heterogeneous multi-task learning. First, we propose a new metric to quantify the relations between tasks. Based on the quantitative metric, we then develop the task relation layer, which can be combined with any deep learning architecture to form task relation networks to fully exploit the relations of different tasks in an online fashion. Benefiting from the task relation layer, the task relation networks can better leverage the mutual information from the data. We demonstrate our proposed task relation networks are effective in improving the performance in both homogeneous and heterogeneous multi-task learning settings through extensive experiments on computer vision tasks.
Jianshu Li, Pan Zhou 0002, Yunpeng Chen, Jian Zhao 0006, Sujoy Roy, Shuicheng Yan, Jiashi Feng, Terence Sim
WACV3
2018 Multi-fiber Networks for Video Recognition
Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, Jiashi Feng
ECCV (1)1
2018 Sharing Residual Units Through Collective Tensor Factorization To Improve Deep Neural Networks
abstract
The residual unit and its variations are wildly used in building very deep neural networks for alleviating optimization difficulty. In this work, we revisit the standard residual function as well as its several successful variants and propose a unified framework based on tensor Block Term Decomposition (BTD) to explain these apparently different residual functions from the tensor decomposition view. With the BTD framework, we further propose a novel basic network architecture, named the Collective Residual Unit (CRU). CRU further enhances parameter efficiency of deep residual neural networks by sharing core factors derived from collective tensor factorization over the involved residual units. It enables efficient knowledge sharing across multiple residual units, reduces the number of model parameters, lowers the risk of over-fitting, and provides better generalization ability. Extensive experimental results show that our proposed CRU network brings outstanding parameter efficiency -- it achieves comparable classification performance with ResNet-200 while using a model size as small as ResNet-50 on the ImageNet-1k and Places365-Standard benchmark datasets.
Yunpeng Chen, Xiaojie Jin 0004, Bingyi Kang, Jiashi Feng, Shuicheng Yan
IJCAI1
2018 Multi-Human Parsing Machines
abstract
Human parsing is an important task in human-centric analysis. Despite the remarkable progress in single-human parsing, the more realistic case of multi-human parsing remains challenging in terms of the data and the model. Compared with the considerable number of available single-human parsing datasets, the datasets for multi-human parsing are very limited in number mainly due to the huge annotation effort required. Besides the data challenge to multi-human parsing, the persons in real-world scenarios are often entangled with each other due to close interaction and body occlusion, making it difficult to distinguish body parts from different person instances. In this paper we propose the Multi-Human Parsing Machines (MHPM) system, which contains an MHP Montage model and an MHP Solver, to address both challenges in multi-human parsing. Specifically, the MHP Montage model in MHPM generates realistic images with multiple persons together with the parsing labels. It intelligently composes single persons onto background scene images while maintaining the structural information between persons and the scene. The generated images can be used to train better multi-human parsing algorithms. On the other hand, the MHP Solver in MHPM solves the bottleneck of distinguishing multiple entangled persons with close interaction. It employs a Group-Individual Push and Pull (GIPP) loss function, which can effectively separate persons with close interaction. We experimentally show that the proposed MHPM can achieve state-of-the-art performance on the multi-human parsing benchmark and the person individualization benchmark, which distinguishes closely entangled person instances.
Jianshu Li, Jian Zhao 0006, Yunpeng Chen, Sujoy Roy, Shuicheng Yan, Jiashi Feng, Terence Sim
ACM Multimedia3
2018 A^2-Nets: Double Attention Networks
abstract
Learning to capture long-range relations is fundamental to image/video recognition. Existing CNN models generally rely on increasing depth to model such relations which is highly inefficient. In this work, we propose the “double attention block”, a novel component that aggregates and propagates informative global features from the entire spatio-temporal space of input images/videos, enabling subsequent convolution layers to access features from the entire space efficiently. The component is designed with a double attention mechanism in two steps, where the first step gathers features from the entire space into a compact set through second-order attention pooling and the second step adaptively selects and distributes features to each location via another attention. The proposed double attention block is easy to adopt and can be plugged into existing deep neural networks conveniently. We conduct extensive ablation studies and experiments on both image and video recognition tasks for evaluating its performance. On the image recognition task, a ResNet-50 equipped with our double attention blocks outperforms a much larger ResNet-152 architecture on ImageNet-1k dataset with over 40% less the number of parameters and less FLOPs. On the action recognition task, our proposed model achieves the state-of-the-art results on the Kinetics and UCF-101 datasets with significantly higher efficiency than recent works.
Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, Jiashi Feng
NeurIPS1
2017 Multi-Path Feedback Recurrent Neural Networks for Scene Parsing
abstract
In this paper, we consider the scene parsing problem and propose a novel Multi-Path Feedback recurrent neural network (MPF-RNN) for parsing scene images. MPF-RNN can enhance the capability of RNNs in modeling long-range context information at multiple levels and better distinguish pixels that are easy to confuse. Different from feedforward CNNs and RNNs with only single feedback, MPF-RNN propagates the contextual features learned at top layer through multiple weighted recurrent connections to learn bottom features. For better training MPF-RNN, we propose a new strategy that considers accumulative loss at multiple recurrent steps to improve performance of the MPF-RNN on parsing small objects. With these two novel components, MPF-RNN has achieved significant improvement over strong baselines (VGG16 and Res101) on five challenging scene parsing benchmarks, including traditional SiftFlow, Barcelona, CamVid, Stanford Background as well as the recently released large-scale ADE20K.
Xiaojie Jin 0004, Yunpeng Chen, Zequn Jie, Jiashi Feng, Shuicheng Yan
AAAI2
2017 Marginalized CNN: Learning Deep Invariant Representations
Jian Zhao 0006, Jianshu Li, Fang Zhao 0006, Xuecheng Nie, Yunpeng Chen, Shuicheng Yan, Jiashi Feng
BMVC5
2017 Video Scene Parsing with Predictive Feature Learning
abstract
Video scene parsing is challenging due to the following two reasons: firstly, it is non-trivial to learn meaningful video representations for producing the temporally consistent labeling map; secondly, such a learning process becomes more difficult with insufficient labeled video training data. In this work, we propose a unified framework to address the above two problems, which is to our knowledge the first model to employ predictive feature learning in the video scene parsing. The predictive feature learning is carried out in two predictive tasks: frame prediction and predictive parsing. It is experimentally proved that the learned predictive features in our model are able to significantly enhance the video parsing performance by combining with the standard image parsing network. Interestingly, the performance gain brought by the predictive learning is almost costless as the features are learned from a large amount of unlabeled video data in an unsupervised way. Extensive experiments over two challenging datasets, Cityscapes and Camvid, have demonstrated the effectiveness of our model by showing remarkable improvement over well-established baselines.
Xiaojie Jin 0004, Huaxin Xiao, Xiaohui Shen, Zhe Lin 0001, Jimei Yang, Yunpeng Chen, Jian Dong 0011, Luoqi Liu, Zequn Jie, Jiashi Feng, Shuicheng Yan
ICCV7
2017 Training Group Orthogonal Neural Networks with Privileged Information
abstract
Learning rich and diverse representations is critical for the performance of deep convolutional neural networks (CNNs). In this paper, we consider how to use privileged information to promote inherent diversity of a single CNN model such that the model can learn better representations and offer stronger generalization ability. To this end, we propose a novel group orthogonal convolutional neural network (GoCNN) that learns untangled representations within each layer by exploiting provided privileged information and enhances representation diversity effectively. We take image classification as an example where image segmentation annotations are used as privileged information during the training process. Experiments on two benchmark datasets – ImageNet and PASCAL VOC – clearly demonstrate the strong generalization ability of our proposed GoCNN model. On the ImageNet dataset, GoCNN improves the performance of state-of-the-art ResNet-152 model by absolute value of 1.2% while only uses privileged information of 10% of the training images, confirming effectiveness of GoCNN on utilizing available privileged knowledge to train better CNNs.
Yunpeng Chen, Xiaojie Jin 0004, Jiashi Feng, Shuicheng Yan
IJCAI1
2017 Dual Path Networks
abstract
In this work, we present a simple, highly efficient and modularized Dual Path Network (DPN) for image classification which presents a new topology of connection paths internally. By revealing the equivalence of the state-of-the-art Residual Network (ResNet) and Densely Convolutional Network (DenseNet) within the HORNN framework, we find that ResNet enables feature re-usage while DenseNet enables new features exploration which are both important for learning good representations. To enjoy the benefits from both path topologies, our proposed Dual Path Network shares common features while maintaining the flexibility to explore new features through dual path architectures. Extensive experiments on three benchmark datasets, ImagNet-1k, Places365 and PASCAL VOC, clearly demonstrate superior performance of the proposed DPN over state-of-the-arts. In particular, on the ImagNet-1k dataset, a shallow DPN surpasses the best ResNeXt-101(64x4d) with 26% smaller model size, 25% less computational cost and 8% lower memory consumption, and a deeper DPN (DPN-131) further pushes the state-of-the-art single model performance with about 2 times faster training speed. Experiments on the Places365 large-scale scene dataset, PASCAL VOC detection dataset, and PASCAL VOC segmentation dataset also demonstrate its consistently better performance than DenseNet, ResNet and the latest ResNeXt model over various applications.
Yunpeng Chen, Jianan Li 0001, Huaxin Xiao, Xiaojie Jin 0004, Shuicheng Yan, Jiashi Feng
NIPS1
2017 Predicting Scene Parsing and Motion Dynamics in the Future
abstract
It is important for intelligent systems, e.g. autonomous vehicles and robotics to anticipate the future in order to plan early and make decisions accordingly. Predicting the future scene parsing and motion dynamics helps the agents better understand the visual environment better as the former provides dense semantic segmentations, i.e. what objects will be present and where they will appear, while the latter provides dense motion information, i.e. how the objects move in the future. In this paper, we propose a novel model to predict the scene parsing and motion dynamics in unobserved future video frames simultaneously. Using history information (preceding frames and corresponding scene parsing results) as input, our model is able to predict the scene parsing and motion for arbitrary time steps ahead. More importantly, our model is superior compared to other methods that predict parsing and motion separately, as the complementary relationship between the two tasks are fully utilized in our model through joint learning. To our best knowledge, this is the first attempt in jointly predicting scene parsing and motion dynamics in the future frames. On the large-scale Cityscapes dataset, it is demonstrated that our model produces significantly better parsing and motion prediction results compared to well established baselines. In addition, we also show our model can be used to predict the steering angle of the vehicles, which further verifies the ability of our model to learn underlying latent parameters.
Xiaojie Jin 0004, Huaxin Xiao, Xiaohui Shen, Jimei Yang, Zhe Lin 0001, Yunpeng Chen, Zequn Jie, Jiashi Feng, Shuicheng Yan
NIPS6
2017 Learning to Segment Human by Watching YouTube
abstract
An intuition on human segmentation is that when a human is moving in a video, the video-context (e.g., appearance and motion clues) may potentially infer reasonable mask information for the whole human body. Inspired by this, based on popular deep convolutional neural networks (CNN), we explore a very-weakly supervised learning framework for human segmentation task, where only an imperfect human detector is available along with massive weakly-labeled YouTube videos. In our solution, the video-context guided human mask inference and CNN based segmentation network learning iterate to mutually enhance each other until no further improvement gains. In the first step, each video is decomposed into supervoxels by the unsupervised video segmentation. The superpixels within the supervoxels are then classified as human or non-human by graph optimization with unary energies from the imperfect human detection results and the predicted confidence maps by the CNN trained in the previous iteration. In the second step, the video-context derived human masks are used as direct labels to train CNN. Extensive experiments on the challenging PASCAL VOC 2012 semantic segmentation benchmark demonstrate that the proposed framework has already achieved superior results than all previous weakly-supervised methods with object class or bounding box annotations. In addition, by augmenting with the annotated masks from PASCAL VOC 2012, our method reaches a new state-of-the-art performance on the human segmentation task.
Xiaodan Liang, Yunchao Wei, Liang Lin 0004, Yunpeng Chen, Xiaohui Shen, Jianchao Yang, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 STC: A Simple to Complex Framework for Weakly-Supervised Semantic Segmentation
abstract
Recently, significant improvement has been made on semantic object segmentation due to the development of deep convolutional neural networks (DCNNs). Training such a DCNN usually relies on a large number of images with pixel-level segmentation masks, and annotating these images is very costly in terms of both finance and human effort. In this paper, we propose a simple to complex (STC) framework in which only image-level annotations are utilized to learn DCNNs for semantic segmentation. Specifically, we first train an initial segmentation network called Initial-DCNN with the saliency maps of simple images (i.e., those with a single category of major object(s) and clean background). These saliency maps can be automatically obtained by existing bottom-up salient object detection techniques, where no supervision information is needed. Then, a better network called Enhanced-DCNN is learned with supervision from the predicted segmentation masks of simple images based on the Initial-DCNN as well as the image-level annotations. Finally, more pixel-level segmentation masks of complex images (two or more categories of objects with cluttered background), which are inferred by using Enhanced-DCNN and image-level annotations, are utilized as the supervision information to learn the Powerful-DCNN for semantic segmentation. Our method utilizes 40K simple images from Flickr.com and 10K complex images from PASCAL VOC for step-wisely boosting the segmentation network. Extensive experimental results on PASCAL VOC 2012 segmentation benchmark well demonstrate the superiority of the proposed STC framework compared with other state-of-the-arts.
Yunchao Wei, Xiaodan Liang, Yunpeng Chen, Xiaohui Shen, Ming-Ming Cheng, Jiashi Feng, Yao Zhao 0001, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Collaborative Layer-Wise Discriminative Learning in Deep Neural Networks
Xiaojie Jin 0004, Yunpeng Chen, Jian Dong 0011, Jiashi Feng, Shuicheng Yan
ECCV (7)2
2016 Learning to segment with image-level annotations
Yunchao Wei, Xiaodan Liang, Yunpeng Chen, Zequn Jie, Yanhui Xiao, Yao Zhao 0001, Shuicheng Yan
Pattern Recognit.3
2015 Supervised feature learning via l2-norm regularized logistic regression for 3D object recognition
Fuhao Zou, Yang Yang 0002, Ke Zhou 0001, Yunpeng Chen, Jingkuan Song
Neurocomputing5
2015 Compact Image Fingerprint Via Multiple Kernel Hashing
abstract
Image fingerprinting is regarded as an alternative approach to watermarking in terms of near-duplicate detection application. It consists of feature extraction and feature indexing. Generally, the former is mainly related to discrimination, robustness , and security while the latter closely focuses on the efficiency of fingerprints search. To enable fast fingerprints searching over a very large database, we propose a new kernelized multiple feature hashing method to convert the real-value fingerprints into compact binary-value fingerprints. During the process of converting, the proposed hashing method jointly utilizes the kernel trick and multiple feature fusion strategy to map the image represented by multiple features into a compact binary code. With the help of the kernel function, the hashing method can be applied to any format (such as string, graph, set, and so on) as long as there is an associated kernel function available for similarity measurement. In addition, taking multiple features into account aims at improving the discriminability since these multiple evidences are complementary to each other. The extensive experimental results show that the proposed algorithm outperforms state-of-the-art kernelized hashing methods by up to 10 percent.
Fuhao Zou, Yunpeng Chen, Jingkuan Song, Ke Zhou 0001, Yang Yang 0002, Nicu Sebe
IEEE Trans. Multim.2