EDBT 2026 Demo / reviewers in the wild / expert
Yangyang Xu 0003
dblp:02/9364-3
· DBLP profile ↗
27ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0002-3383-4349ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 22 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Coherent Portrait-to-Anime Translation via Latent Cyclic TransformationabstractTranslating real portrait video into anime is an application of interest to both consumers and researchers. However, anime differs considerably from portraits, making portrait-to-anime translation challenging. Existing StyleGAN-based portrait stylization works assume that the portrait and stylized generators share the same latent space, but this assumption fails in the style of anime due to the large domain gap. Moreover, directly applying them to each video frame often leads to undesirable temporal inconsistencies. In this paper, we argue that two latent spaces with a large domain gap cannot be shared but can be related by a transformation, and develop a cyclic transformation network to connect the two spaces with two cycle constraints. This provides high-quality translation for each frame. We extend our framework to video transformation by proposing a novel frame interpolation constraint which ensures that in-between frames can be interpolated from their neighboring frames, guaranteeing temporal coherence across translated frames. Together with latent code smoothing regularization, this provides temporally coherent video-to-anime translation. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods both qualitatively and quantitatively. Yangyang Xu 0003, Shengfeng He, Kwan-Yee Kenneth Wong, Ping Luo 0002 |
Comput. Vis. Media | 1 |
| 2026 | Invert Your Prompt: Editing-Aware Diffusion Inversion
Yangyang Xu 0003, Wenqi Shao, Yong Du 0003, Haiming Zhu, Yang Zhou 0038, Jiayuan Xie, Ping Luo 0002, Shengfeng He |
Int. J. Comput. Vis. | 1 |
| 2026 | Zero-Shot Video Translation via Token WarpingabstractWith the revolution of generative AI, video-related tasks have been widely studied. However, current state-of-the-art video models still lag behind image models in visual quality and user control over generated content. In this paper, we introduce TokenWarping, a novel framework for temporally coherent video translation. Existing diffusion-based video editing approaches rely solely on key and value patches in self-attention to ensure temporal consistency, often sacrificing the preservation of local and structural regions. Critically, these methods overlook the significance of the query patches in achieving accurate feature aggregation and temporal coherence. In contrast, TokenWarping leverages complementary token priors by constructing temporal correlations across different frames. Our method begins by extracting optical flows from source videos. During the denoising process of the diffusion model, these optical flows are used to warp the previous frame's query, key, and value patches, aligning them with the current frame's patches. By directly warping the query patches, we enhance feature aggregation in self-attention, while warping the key and value patches ensures temporal consistency across frames. This token warping imposes explicit constraints on the self-attention layer outputs, effectively ensuring temporally coherent translation. Our framework does not require any additional training or fine-tuning and can be seamlessly integrated with existing text-to-image editing methods. We conduct extensive experiments on various video translation tasks, demonstrating that TokenWarping surpasses state-of-the-art methods both qualitatively and quantitatively. Video demonstrations are available in supplementary materials. Haiming Zhu, Yangyang Xu 0003, Jun Yu 0002, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Occlusion-Insensitive Talking Head Video Generation via Facelet CompensationabstractTalking head video generation involves animating a still face image using facial motion cues derived from a driving video to replicate target poses and expressions. Traditional methods often rely on the assumption that the relative positions of facial keypoints remain unchanged. However, this assumption fails when keypoints are occluded or when the head is in a profile pose, leading to inconsistencies in identity and blurring in certain facial regions. In this paper, we introduce Occlusion-Insensitive Talking Head Video Generation, a novel approach that eliminates the reliance on spatial correlation of keypoints and instead leverages semantic correlation. Our method transforms facial features into a facelet semantic bank, where each facelet token represents a specific facial semantic. This bank is devoid of spatial information, allowing it to compensate for any invisible or occluded face regions during motion warping. The facelet compensation module then populates the facelet tokens within the initially warped features by learning a correlation matrix between facial semantics and the facelet bank. This approach enables precise compensation for occlusions and pose changes, enhancing the fidelity of the generated videos. Extensive experiments demonstrate that our method achieves state-of-the-art results, preserving source identity, maintaining fine-grained facial details, and capturing nuanced facial expressions with remarkable accuracy. Yuhui Deng 0005, Yuqin Lu, Yangyang Xu 0003, Yongwei Nie, Shengfeng He |
AAAI | 3 |
| 2025 | PersonaMagic: Stage-Regulated High-Fidelity Face Customization with Tandem EquilibriumabstractPersonalized image generation has made significant strides in adapting content to novel concepts. However, a persistent challenge remains: balancing the accurate reconstruction of unseen concepts with the need for editability according to the prompt, especially when dealing with the complex nuances of facial features. In this study, we delve into the temporal dynamics of the text-to-image conditioning process, emphasizing the crucial role of stage partitioning in introducing new concepts. We present PersonaMagic, a stage-regulated generative technique designed for high-fidelity face customization. Using a simple MLP network, our method learns a series of embeddings within a specific timestep interval to capture face concepts. Additionally, we develop a Tandem Equilibrium mechanism that adjusts self-attention responses in the text encoder, balancing text description and identity preservation, improving both areas. Extensive experiments confirm the superiority of PersonaMagic over state-of-the-art methods in both qualitative and quantitative evaluations. Moreover, its robustness and flexibility are validated in non-facial domains, and it can also serve as a valuable plug-in for enhancing the performance of pretrained personalization models. Xinzhe Li 0003, Jiahui Zhan, Shengfeng He, Yangyang Xu 0003, Junyu Dong, Huaidong Zhang, Yong Du 0003 |
AAAI | 4 |
| 2025 | Cross-Subject Mind Decoding from Inaccurate Representations
Yangyang Xu 0003, Bangzhen Liu, Wenqi Shao, Yong Du 0003, Shengfeng He |
ICCV | 1 |
| 2025 | OmniVTON: Training-Free Universal Virtual Try-OnabstractImage-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve fine-grained texture retention. For precise pose alignment, we utilize DDIM inversion to capture structural cues while suppressing texture interference, ensuring accurate body alignment independent of the original image textures. By disentangling garment and pose constraints, OmniVTON eliminates the bias inherent in diffusion models when handling multiple conditions simultaneously. Experimental results demonstrate that OmniVTON achieves superior performance across diverse datasets, garment types, and application scenarios. Notably, it is the first framework capable of multi-human VTON, enabling realistic garment transfer across multiple individuals in a single scene. Code is available at https://github.com/Jerome-Young/OmniVTON Zhaotong Yang, Shengfeng He, Xinzhe Li 0003, Yangyang Xu 0003, Junyu Dong, Yong Du 0003 |
ICCV | 5 |
| 2025 | Stable Score DistillationabstractText-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieves cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's structure and ensures that editing trajectories are closely aligned with the source prompt, enabling smooth, prompt-specific modifications while maintaining coherence in surrounding regions. Additionally, SSD incorporates a prompt enhancement branch to boost editing strength, particularly for style transformations. Our method achieves state-of-the-art results in 2D and 3D editing tasks, including NeRF and text-driven style edits, with faster convergence and reduced complexity, providing a robust and efficient solution for text-guided editing. Haiming Zhu, Yangyang Xu 0003, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du 0003, Jun Yu 0002, Shengfeng He |
ICCV | 2 |
| 2025 | DiffusionMat: Alpha Matting as Deterministic Sequential Refinement Learning
Yangyang Xu 0003, Shengfeng He, Wenqi Shao, Yong Du 0003, Kwan-Yee Kenneth Wong, Yu Qiao 0001, Jun Yu 0002, Ping Luo 0002 |
ACM Multimedia | 1 |
| 2025 | RIGID: Recurrent GAN Inversion and Editing of Real Face Videos and Beyond
Yangyang Xu 0003, Shengfeng He, Kwan-Yee Kenneth Wong, Ping Luo 0002 |
Int. J. Comput. Vis. | 1 |
| 2025 | ContX: Scene context prediction via context bank and layout perception
Jingxin Liang, Yangyang Xu 0003, Haorui Song, Yuqin Lu, Yuhui Deng 0005, Yiyi Long, Shengxin Liu, Jianbo Jiao, Shengfeng He |
Pattern Recognit. | 2 |
| 2025 | Open-Set Mixed Domain Adaptation via Visual-Linguistic Focal EvolvingabstractWe introduce a new task, Open-set Mixed Domain Adaptation (OSMDA), which considers the potential mixture of multiple distributions in the target domains, thereby better simulating real-world scenarios. To tackle the semantic ambiguity arising from multiple domains, our key idea is that the linguistic representation can serve as a universal descriptor for samples of the same category across various domains. We thus propose a more practical framework for cross-domain recognition via visual-linguistic guidance. On the other hand, the presence of multiple domains also poses a new challenge in classifying both known and unknown categories. To combat this issue, we further introduce a visual-linguistic focal evolving approach to gradually enhance the classification ability of a known/unknown binary classifier from two aspects. Specifically, we start with identifying highly confident focal samples to expand the pool of known samples by incorporating those from different domains. Then, we amplify the feature discrepancy between known and unknown samples through dynamic entropy evolving via an adaptive entropies min/max game, enabling us to accurately identify possible unknown samples in a gradual manner. Extensive experiments demonstrate our method’s superiority against the state-of-the-arts in both open-set and open-set mixed domain adaptation. Bangzhen Liu, Yangyang Xu 0003, Xuemiao Xu, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | DreamAnime: Learning Style-Identity Textual Disentanglement for Anime and BeyondabstractText-to-image generation models have significantly broadened the horizons of creative expression through the power of natural language. However, navigating these models to generate unique concepts, alter their appearance, or reimagine them in unfamiliar roles presents an intricate challenge. For instance, how can we exploit language-guided models to transpose an anime character into a different art style, or envision a beloved character in a radically different setting or role? This paper unveils a novel approach named DreamAnime, designed to provide this level of creative freedom. Using a minimal set of 2-3 images of a user-specified concept such as an anime character or an art style, we teach our model to encapsulate its essence through novel "words" in the embedding space of a pre-existing text-to-image model. Crucially, we disentangle the concepts of style and identity into two separate "words", thus providing the ability to manipulate them independently. These distinct "words" can then be pieced together into natural language sentences, promoting an intuitive and personalized creative process. Empirical results suggest that this disentanglement into separate word embeddings successfully captures a broad range of unique and complex concepts, with each word focusing on style or identity as appropriate. Comparisons with existing methods illustrate DreamAnime's superior capacity to accurately interpret and recreate the desired concepts across various applications and tasks. Chenshu Xu, Yangyang Xu 0003, Huaidong Zhang, Xuemiao Xu, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | RIGID: Recurrent GAN Inversion and Editing of Real Face VideosabstractGAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a unified recurrent framework, named Recurrent vIdeo GAN Inversion and eDiting (RIGID), to explicitly and simultaneously enforce temporally coherent GAN inversion and facial editing of real videos. Our approach models the temporal relations between current and previous frames from three aspects. To enable a faithful real video reconstruction, we first maximize the inversion fidelity and consistency by learning a temporal compensated latent code. Second, we observe incoherent noises lie in the high-frequency domain that can be disentangled from the latent space. Third, to remove the inconsistency after attribute manipulation, we propose an in-between frame composition constraint such that the arbitrary frame must be a direct composite of its neighboring frames. Our unified framework learns the inherent coherence between input frames in an end-to-end manner, and therefore it is agnostic to a specific attribute and can be applied to arbitrary editing of the same video without re-training. Extensive experiments demonstrate that RIGID outperforms state-of-the-art methods qualitatively and quantitatively in both inversion and editing tasks. The deliverables can be found in https://cnnlstm.github.io/RIGID. Yangyang Xu 0003, Shengfeng He, Kwan-Yee Kenneth Wong, Ping Luo 0002 |
ICCV | 1 |
| 2023 | Deep Texture-Aware Features for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging task that aims to identify objects having similar texture to the surroundings. This paper presents to amplify the subtle texture difference between camouflaged objects and the background for camouflaged object detection by formulating multiple texture-aware refinement modules to learn the texture-aware features in a deep convolutional neural network. The texture-aware refinement module computes the biased co-variance matrices of feature responses to extract the texture information, adopts an affinity loss to learn a set of parameter maps that help to separate the texture between camouflaged objects and the background, and leverages a boundary-consistency loss to explore the structures of object details. We evaluate our network on the benchmark datasets for camouflaged object detection both qualitatively and quantitatively. Experimental results show that our approach outperforms various state-of-the-art methods by a large margin. Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Yangyang Xu 0003, Weiming Wang 0002, Zijun Deng, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Representative Feature Alignment for Adaptive Object DetectionabstractUnsupervised domain adaptation for object detection aims to generalize the object detector trained on the label-rich source domain to the unlabeled target domain. Recently, existing works adopt the instance-level alignment or pixel-level alignment to perform domain transfer, which can effectively avoid the negative transfer due to the diverse background between domains. However, we find that they treat all the regions of an instance feature equally without suppressing background area. They do not segment the specific texture and discriminative regions of objects, which are transferable during adaptation. We call the features that combine the local structure feature and semantic discriminant features as representative features. We propose a novel Representative Feature Alignment (RFA) model to align the features extracted from representative patterns of objects, i.e. representative features, for domain adaptation. Specifically, the representative features are extracted by the Representative Feature Extraction (RFE) submodules. The RFE submodules take the features extracted from different intermediate layers of the detector as input, and filter out the representative features layer-by-layer via integrating class weighting generator, category selection and class activation mapping. Then the representative features from multi-layers are further adaptively aggregated to obtain the final representative features, which are utilized to conduct feature alignment in a class-aware manner. Our representative features are free of untransferable regions and background areas, which leads to better feature alignment. Extensive experimental results show that the proposed model outperforms state-of-the-art methods on a few benchmark datasets. Shan Xu 0006, Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Yangyang Xu 0003, Liangui Dai, Kup-Sze Choi, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Parsing-Conditioned Anime Translation: A New Dataset and MethodabstractAnime is an abstract art form that is substantially different from the human portrait, leading to a challenging misaligned image translation problem that is beyond the capability of existing methods. This can be boiled down to a highly ambiguous unconstrained translation between two domains. To this end, we design a new anime translation framework by deriving the prior knowledge of a pre-trained StyleGAN model. We introduce disentangled encoders to separately embed structure and appearance information into the same latent code, governed by four tailored losses. Moreover, we develop a FaceBank aggregation method that leverages the generated data of the StyleGAN, anchoring the prediction to produce in-domain animes. To empower our model and promote the research of anime translation, we propose the first anime portrait parsing dataset, Danbooru-Parsing , containing 4,921 densely labeled images across 17 classes. This dataset connects the face semantics with appearances, enabling our new constrained translation setting. We further show the editability of our results, and extend our method to manga images, by generating the first manga parsing pseudo data. Extensive experiments demonstrate the values of our new dataset and method, resulting in the first feasible solution on anime translation. Zhansheng Li, Yangyang Xu 0003, Nanxuan Zhao, Yang Zhou 0007, Yongtuo Liu, Dahua Lin, Shengfeng He |
ACM Trans. Graph. | 2 |
| 2022 | High-resolution Face Swapping via Latent Semantics DisentanglementabstractWe present a novel high-resolution face swapping method using the inherent prior knowledge of a pre-trained GAN model. Although previous research can leverage generative priors to produce high-resolution results, their quality can suffer from the entangled semantics of the latent space. We explicitly disentangle the latent semantics by utilizing the progressive nature of the generator, deriving structure at-tributes from the shallow layers and appearance attributes from the deeper ones. Identity and pose information within the structure attributes are further separated by introducing a landmark-driven structure transfer latent direction. The disentangled latent code produces rich generative features that incorporate feature blending to produce a plausible swapping result. We further extend our method to video face swapping by enforcing two spatio-temporal constraints on the latent space and the image space. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art image/video face swapping methods in terms of hallucination quality and consistency. Code can be found at: https://github.com/cnnlstm/FSLSD_HiRes. Yangyang Xu 0003, Bailin Deng, Junle Wang, Yanqing Jing, Jia Pan 0001, Shengfeng He |
CVPR | 1 |
| 2022 | Background Matting via Recursive ExcitationabstractWe propose a simple yet effective technique that significantly improves the performance of the current state-of-the-art background matting model without compromising its original speed. We achieve this by carefully exciting the proper neural activations using an excitation map in the training phase and performing recursive inference in the testing phase. To avoid being over-reliant on perfect excitations, we follow the idea of curriculum learning to divide the training phase into three easy-to-hard stages and gradually shift the excitation map from GT alpha matte to pseudo GT alpha matte. In the testing phase, we propose a recursive inference mechanism that uses the output alpha matte as the excitation map to further refine the output alpha matte. Our method is a simple plug-in for arbitrary matting models. Compared with the original ones, the enhanced models alleviate the problem of performance degradation with complex background and thus boosts the matting accuracy. Junjie Deng, Yangyang Xu 0003, Shengfeng He |
ICME | 2 |
| 2022 | Self-Supervised Matting-Specific Portrait Enhancement and GenerationabstractWe resolve the ill-posed alpha matting problem from a completely different perspective. Given an input portrait image, instead of estimating the corresponding alpha matte, we focus on the other end, to subtly enhance this input so that the alpha matte can be easily estimated by any existing matting models. This is accomplished by exploring the latent space of GAN models. It is demonstrated that interpretable directions can be found in the latent space and they correspond to semantic image transformations. We further explore this property in alpha matting. Particularly, we invert an input portrait into the latent code of StyleGAN, and our aim is to discover whether there is an enhanced version in the latent space which is more compatible with a reference matting model. We optimize multi-scale latent vectors in the latent spaces under four tailored losses, ensuring matting-specificity and subtle modifications on the portrait. We demonstrate that the proposed method can refine real portrait images for arbitrary matting models, boosting the performance of automatic alpha matting by a large margin. In addition, we leverage the generative property of StyleGAN, and propose to generate enhanced portrait data which can be treated as the pseudo GT. It addresses the problem of expensive alpha matte annotation, further augmenting the matting performance of existing models. Yangyang Xu 0003, Shengfeng He |
IEEE Trans. Image Process. | 1 |
| 2022 | Pro-PULSE: Learning Progressive Encoders of Latent Semantics in GANs for Photo UpsamplingabstractThe state-of-the-art photo upsampling method, PULSE, demonstrates that a sharp, high-resolution (HR) version of a given low-resolution (LR) input can be obtained by exploring the latent space of generative models. However, mapping an extreme LR input (162) directly to an HR image (10242) is too ambiguous to preserve faithful local facial semantics. In this paper, we propose an enhanced upsampling approach, Pro-PULSE, that addresses the issues of semantic inconsistency and optimization complexity. Our idea is to learn an encoder that progressively constructs the HR latent codes in the extended$\mathcal {W}+$latent space of StyleGAN. This design divides the complex$64\times $upsampling problem into several steps, and therefore small-scale facial semantics can be inherited from one end to the other. In particular, we train two encoders, the base encoder maps latent vectors in$\mathcal {W}$space and serves as a foundation of the HR latent vector, while the second scale-specific encoder performed in$\mathcal {W}+$space gradually replaces the previous vector produced by the base encoder at each scale. This process produces intermediate side-outputs, which injects deep supervision into the training of encoder. Extensive experiments demonstrate superiorities over the latest latent space exploration methods, in terms of efficiency, quantitative quality metrics, and qualitative visual results. Yang Zhou 0038, Yangyang Xu 0003, Yong Du 0003, Shengfeng He |
IEEE Trans. Image Process. | 2 |
| 2021 | From Continuity to Editability: Inverting GANs with Consecutive ImagesabstractExisting GAN inversion methods are stuck in a paradox that the inverted codes can either achieve high-fidelity reconstruction, or retain the editing capability. Having only one of them clearly cannot realize real image editing. In this paper, we resolve this paradox by introducing consecutive images (e.g., video frames or the same person with different poses) into the inversion process. The rationale behind our solution is that the continuity of consecutive images leads to inherent editable directions. This inborn property is used for two unique purposes: 1) regularizing the joint inversion process, such that each of the inverted codes is semantically accessible from one of the other and fastened in an editable domain; 2) enforcing inter-image coherence, such that the fidelity of each inverted code can be maximized with the complement of other images. Extensive experiments demonstrate that our alternative significantly outperforms state-of-the-art methods in terms of reconstruction fidelity and editability on both the real image dataset and synthesis dataset. Furthermore, our method provides the first support of video-based GAN inversion and an interesting application of unsupervised semantic transfer from consecutive images. Source code can be found at: https://github.com/cnnlstm/InvertingGANs_with_ConsecutiveImgs. Yangyang Xu 0003, Yong Du 0003, Wenpeng Xiao, Xuemiao Xu, Shengfeng He |
ICCV | 1 |
| 2021 | Multi-View Face Synthesis via Progressive Face FlowabstractExisting GAN-based multi-view face synthesis methods rely heavily on "creating" faces, and thus they struggle in reproducing the faithful facial texture and fail to preserve identity when undergoing a large angle rotation. In this paper, we combat this problem by dividing the challenging large-angle face synthesis into a series of easy small-angle rotations, and each of them is guided by a face flow to maintain faithful facial details. In particular, we propose a Face Flow-guided Generative Adversarial Network (FFlowGAN) that is specifically trained for small-angle synthesis. The proposed network consists of two modules, a face flow module that aims to compute a dense correspondence between the input and target faces. It provides strong guidance to the second module, face synthesis module, for emphasizing salient facial texture. We apply FFlowGAN multiple times to progressively synthesize different views, and therefore facial features can be propagated to the target view from the very beginning. All these multiple executions are cascaded and trained end-to-end with a unified back-propagation, and thus we ensure each intermediate step contributes to the final result. Extensive experiments demonstrate the proposed divide-and-conquer strategy is effective, and our method outperforms the state-of-the-art on four benchmark datasets qualitatively and quantitatively. Yangyang Xu 0003, Xuemiao Xu, Jianbo Jiao, Shengfeng He |
IEEE Trans. Image Process. | 1 |
| 2021 | Erratum to "Multi-View Face Synthesis via Progressive Face Flow"
Yangyang Xu 0003, Xuemiao Xu, Jianbo Jiao, Shengfeng He |
IEEE Trans. Image Process. | 1 |
| 2021 | Transductive Zero-Shot Action Recognition via Visually Connected Graph Convolutional NetworksabstractWith the explosive growth of action categories, zero-shot action recognition aims to extend a well-trained model to novel/unseen classes. To bridge the large knowledge gap between seen and unseen classes, in this brief, we visually associate unseen actions with seen categories in a visually connected graph, and the knowledge is then transferred from the visual features space to semantic space via the grouped attention graph convolutional networks (GAGCNs). In particular, we extract visual features for all the actions, and a visually connected graph is built to attach seen actions to visually similar unseen categories. Moreover, the proposed grouped attention mechanism exploits the hierarchical knowledge in the graph so that the GAGCN enables propagating the visual-semantic connections from seen actions to unseen ones. We extensively evaluate the proposed method on three data sets: HMDB51, UCF101, and NTU RGB + D. Experimental results show that the GAGCN outperforms state-of-the-art methods. Yangyang Xu 0003, Chu Han, Harry Qin, Xuemiao Xu, Guoqiang Han 0002, Shengfeng He |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Invertible Grayscale with Sparsity Enforcing PriorsabstractColor dimensionality reduction is believed as a non-invertible process, as re-colorization results in perceptually noticeable and unrecoverable distortion. In this article, we propose to convert a color image into a grayscale image that can fully recover its original colors, and more importantly, the encoded information is discriminative and sparse, which saves storage capacity. Particularly, we design an invertible deep neural network for color encoding and decoding purposes. This network learns to generate a residual image that encodes color information, and it is then combined with a base grayscale image for color recovering. In this way, the non-differentiable compression process (e.g., JPEG) of the base grayscale image can be integrated into the network in an end-to-end manner. To further reduce the size of the residual image, we present a specific layer to enhance Sparsity Enforcing Priors (SEP), thus leading to negligible storage space. The proposed method allows color embedding on a sparse residual image while keeping a high, 35dB PSNR on average. Extensive experiments demonstrate that the proposed method outperforms state-of-the-arts in terms of image quality and tolerability to compression. Yong Du 0003, Yangyang Xu 0003, Taizhong Ye, Chu-Feng Xiao 0001, Junyu Dong, Guoqiang Han 0002, Shengfeng He |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Unsupervised Domain Adaptation via Importance SamplingabstractUnsupervised domain adaptation aims to generalize a model from the label-rich source domain to the unlabeled target domain. Existing works mainly focus on aligning the global distribution statistics between source and target domains. However, they neglect distractions from the unexpected noisy samples in domain distribution estimation, leading to domain misalignment or even negative transfer. In this paper, we present an importance sampling method for domain adaptation (ISDA), to measure sample contributions according to their “informative” levels. In particular, informative samples, as well as outliers, can be effectively modeled using feature-norm and prediction entropy of the network. The importance of information is further formulated as the importance sampling losses in features and label spaces. In this way, the proposed model mitigates the noisy outliers while enhancing the important samples during domain alignment. In addition, our model is easy to implement yet effective, and it does not introduce any extra parameters. Extensive experiments on several benchmark datasets show that our method outperforms state-of-the-art methods under both the standard and partial domain adaptation settings. Xuemiao Xu, Hai He, Huaidong Zhang, Yangyang Xu 0003, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 4 |