EDBT 2026 Demo / reviewers in the wild / expert
Congying Han
dblp:07/2808
· DBLP profile ↗
42ranked-venue papers
0as first author
36since 2021 · last 2026
0000-0002-3445-4620ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Theory of computation · 5 · 5 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerated variance-reduced random reshuffling gradient descent algorithm for nonconvex finite-sum optimization
Longhui Liu, Congying Han, Tiande Guo |
Expert Syst. Appl. | 2 |
| 2026 | Stochastic ADMM with variance-reduced recursive momentum and its accelerated variant for nonconvex nonsmooth optimization
Feiyu Long, Congying Han, Tiande Guo, Shichen Liao |
J. Glob. Optim. | 2 |
| 2025 | StyO: Stylize Your Face in Only One-ShotabstractThis paper focuses on face stylization with a single artistic target. Existing works for this task often fail to retain the source content while achieving geometry variation. Here, we present a novel StyO model, i.e., Stylize the face in only One-shot, to solve the above problem. In particular, StyO exploits a disentanglement and recombination strategy. It first disentangles the content and style of source and target images into identifiers, which are then recombined in a cross manner to derive the stylized face image. In this way, StyO decomposes complex images into independent and specific attributes, and simplifies one-shot face stylization as the combination of different attributes from input images, thus producing results better matching face geometry of target image and content of source one. StyO is implemented with latent diffusion models (LDM) and composed of two key modules: 1) Identifier Disentanglement Learner (IDL) for disentanglement phase. It represents identifiers as contrastive text prompts, i.e. positive and negative descriptions. And it introduces a novel triple reconstruction loss to fine-tune the pre-trained LDM for encoding style and content into corresponding identifiers; 2) Fine-graind Content Controller (FCC) for recombination phase. It recombines disentangled identifiers from IDL to form an augmented text prompt for generating stylized faces. In addition, FCC also constrains the cross-attention maps of latent and text features to preserve source face details in results. The extensive evaluation shows that StyO produces high-quality images on numerous paintings of various styles and outperforms the current state-of-the-art. Bonan Li, Xuecheng Nie, Congying Han, Yinhan Hu, Xinmin Qiu, Tiande Guo |
AAAI | 4 |
| 2025 | DreamHA: Towards High-Quality Human Animation with Image-to-Video Diffusion ModelsabstractRecent diffusion models have made significant advancements in generating lifelike videos from driving signals, including a reference character and a skeleton sequence. Nevertheless, these models often struggle with maintaining fidelity, as the generated results frequently deviate in character features, e.g., appearance and identity from the reference. We attribute this issue to the use of driving signals from the same individual during the training process, which biases the model towards skeleton-based shape features and limits its capacity to fully exploit character-specific information, and propose DreamHA to address this issue. DreamHA incorporates diffusion models with Rigid Transformation Augmentation (RTAug), a simple yet effective technique to perturb the shape characteristics of training data, improving the capability of diffusion models in capturing basic appearance features. Additionally, we introduce Identity Keeper (IK) to provide fine-grained facial control and enhance identity consistency. Extensive experimental results demonstrate that our method outperforms state-of-the-art approaches, producing more faithful and consistent animations. Longran Shao, Bonan Li, Congying Han, Wenzhao Liu, Tiande Guo, Tianchi Xing, Xinmin Qiu |
ICASSP | 3 |
| 2025 | A-PSRO: A Unified Strategy Learning Method with Advantage Metric for Normal-form GamesabstractSolving the Nash equilibrium in normal-form games with large-scale strategy spaces presents significant challenges. Open-ended learning frameworks, such as PSRO and its variants, have emerged as effective solutions. However, these methods often lack an efficient metric for evaluating strategy improvement, which limits their effectiveness in approximating equilibria.
In this paper, we introduce a novel evaluative metric called Advantage, which possesses desirable properties inherently connected to the Nash equilibrium, ensuring that each strategy update approaches equilibrium.
Building upon this, we propose the Advantage Policy Space Response Oracle (A-PSRO), an innovative unified open-ended learning framework applicable to both zero-sum and general-sum games. A-PSRO leverages the Advantage as a refined evaluation metric, leading to a consistent learning objective for agents in normal-form games.
Experiments showcase that A-PSRO significantly reduces exploitability in zero-sum games and improves rewards in general-sum games, outperforming existing algorithms and validating its practical effectiveness. Yudong Hu, Haoran Li 0027, Congying Han, Tiande Guo, Bonan Li, Mingqiang Li |
ICML | 3 |
| 2025 | Feature out! Let Raw Image as Your Condition for Blind Face RestorationabstractBlind face restoration (BFR), which involves converting low-quality (LQ) images into high-quality (HQ) images, remains challenging due to complex and unknown degradations.
While previous diffusion-based methods utilize feature extractors from LQ images as guidance, using raw LQ images directly as the starting point for the reverse diffusion process offers a theoretically optimal solution.
In this work, we propose Pseudo-Hashing Image-to-image Schrödinger Bridge (P-I2SB), a novel framework inspired by optimal mass transport problems, which enhances the restoration potential of Schrödinger Bridge (SB) by correcting data distributions and effectively learning the optimal transport path between any two data distributions.
Notably, we theoretically explore and identify that existing methods are limited by the optimality and reversibility of solutions in SB, leading to suboptimal performance.
Our approach involves preprocessing HQ images during training by hashing them into pseudo-samples according to a rule related to LQ images, ensuring structural similarity in distribution.
This guarantees optimal and reversible solutions in SB, enabling the inference process to learn effectively and allowing P-I2SB to achieve state-of-the-art results in BFR, with more natural textures and retained inference speed compared to previous methods. Xinmin Qiu, Gege Chen, Bonan Li, Congying Han, Tiande Guo |
ICML | 4 |
| 2025 | MIRROR: Make Your Object-Level Multi-View Generation More Consistent with Training-Free RectificationabstractMulti-view Diffusion has greatly advanced the development of 3D content creation by generating multiple images from distinct views, achieving remarkable photorealistic results. However, existing works are still vulnerable to inconsistent 3D geometric structures (commonly known as Janus Problem) and severe artifacts. In this paper, we introduce MIRROR, a versatile plug-and-play method that rectifies such inconsistencies in a training-free manner, enabling the acquisition of high-fidelity, realistic structures without compromising diversity. Our key idea focuses on tracing the motion trajectory of physical points across adjacent viewpoints, enabling rectifications based on neighboring observations of the same region. Technically, MIRROR comprises two core modules: Trajectory Tracking Module (TTM) for pixel-wise trajectory tracking that labels identical points across views, and Feature Rectification Module (FRM) for explicitly adjustment of each pixel embedding on noisy synthesized images by minimizing the distance to corresponding block features in neighboring views, thereby achieving consistent outputs. Extensive evaluations demonstrate that MIRROR can seamlessly integrate with a diverse range of off-the-shelf object-level multi-view diffusion models, significantly enhancing both the consistency and the fidelity in an efficient way. Tianchi Xing, Bonan Li, Congying Han, Xinmin Qiu, Tiande Guo |
ICML | 3 |
| 2025 | Understanding Oversmoothing in Diffusion-Based GNNs From the Perspective of Operator Semigroup TheoryabstractThis paper presents an analytical study of the oversmoothing issue in diffusion-based Graph Neural Networks (GNNs). Generalizing beyond extant approaches grounded in random walk analysis or particle systems, we approach this problem through operator semigroup theory. This theoretical framework allows us to rigorously prove that oversmoothing is intrinsically linked to the ergodicity of the diffusion operator. Relying on semigroup method, we can quantitatively analyze the dynamic of graph diffusion and give a specific mathematical form of the smoothing feature by ergodicity and invariant measure of operator, which improves previous works only show existence of oversmoothing. This finding further poses a general and mild ergodicity-breaking condition, encompassing the various specific solutions previously offered, thereby presenting a more universal and theoretically grounded approach to relieve oversmoothing in diffusion-based GNNs. Additionally, we offer a probabilistic interpretation of our theory, forging a link with prior works and broadening the theoretical horizon. Our experimental results reveal that this ergodicity-breaking term effectively mitigates oversmoothing measured by Dirichlet energy, and simultaneously enhances performance in node classification tasks. Chenguang Wang 0001, Xinyan Wang 0004, Congying Han, Tiande Guo, Tianshu Yu 0001 |
KDD (1) | 4 |
| 2025 | Purity Law for Neural Routing Problem Solvers with Enhanced GeneralizabilityabstractAchieving generalization in neural approaches across different scales and distributions remains a significant challenge for routing problems. A key obstacle is that neural networks often fail to learn robust principles for identifying universal patterns and deriving optimal solutions from diverse instances. In this paper, we first uncover Purity Law, a fundamental structural principle for optimal solutions of routing problems, defining that edge prevalence grows exponentially with the sparsity of surrounding vertices. Statistically and theoretically validated across diverse instances, Purity Law reveals a consistent bias toward local sparsity in global optima. Building on this insight, we propose Purity Policy Optimization (PUPO), a novel training paradigm that explicitly aligns characteristics of neural solutions with Purity Law during the solution construction process to enhance generalization. Extensive experiments demonstrate that PUPO can be seamlessly integrated with popular neural solvers, significantly enhancing their generalization performance without incurring additional computational overhead during inference. Wenzhao Liu, Haoran Li 0027, Congying Han, Tiande Guo |
NeurIPS | 3 |
| 2025 | Momentum-based variance-reduced stochastic Bregman proximal gradient methods for nonconvex nonsmooth optimization
Shichen Liao, Yan Liu 0092, Congying Han, Tiande Guo |
Expert Syst. Appl. | 3 |
| 2025 | An inertial stochastic Bregman generalized alternating direction method of multipliers for nonconvex and nonsmooth optimization
Longhui Liu, Congying Han, Tiande Guo, Shichen Liao |
Expert Syst. Appl. | 2 |
| 2025 | Preference-based opponent shaping in differentiable games
Xinyu Qiao, Yudong Hu, Congying Han, Weiyan Wu, Tiande Guo |
Mach. Learn. | 3 |
| 2024 | Learning Dynamic Tetrahedra for High-Quality Talking Head SynthesisabstractRecent works in implicit representations, such as Neural Radiance Fields (NeRF), have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters, since the lack of explicit geometric constraints poses a fundamental challenge in accurately modeling complex facial deformations. In this paper, we introduce Dynamic Tetrahedra (DynTet), a novel hybrid representation that encodes explicit dynamic meshes by neural networks to ensure geometric consistency across various motions and viewpoints. DynTet is parameterized by the coordinate-based networks which learn signed distance, deformation, and material texture, anchoring the training data into a predefined tetrahedra grid. Leveraging Marching Tetrahedra, DynTet efficiently decodes textured meshes with a consistent topology, enabling fast rendering through a differentiable rasterizer and supervision via a pixel loss. To enhance training efficiency, we incorporate classical 3D Morphable Models to facilitate geometry learning and define a canonical space for simplifying texture learning. These advantages are readily achievable owing to the effective geometric representation employed in DynTet. Compared with prior works, DynTet demonstrates significant improvements in fidelity, lip synchronization, and real-time performance according to various metrics. Beyond producing stable and visually appealing synthesis videos, our method also outputs the dynamic meshes which is promising to enable many emerging applications. Code is available at https://github.com/zhangzc21/DynTet. Ruobing Zheng, Bonan Li, Congying Han, Tiande Guo, Jingdong Chen, Ziwen Liu 0001, Ming Yang 0007 |
CVPR | 4 |
| 2024 | BlazeBVD: Make Scale-Time Equalization Great Again for Blind Video Deflickering
Xinmin Qiu, Congying Han, Bonan Li, Tiande Guo, Pingyu Wang, Xuecheng Nie |
ECCV (17) | 2 |
| 2024 | Towards Optimal Adversarial Robust Q-learning with Bellman Infinity-errorabstractEstablishing robust policies is essential to counter attacks or disturbances affecting deep reinforcement learning (DRL) agents. Recent studies explore state-adversarial robustness and suggest the potential lack of an optimal robust policy (ORP), posing challenges in setting strict robustness constraints. This work further investigates ORP: At first, we introduce a consistency assumption of policy (CAP) stating that optimal actions in the Markov decision process remain consistent with minor perturbations, supported by empirical and theoretical evidence. Building upon CAP, we crucially prove the existence of a deterministic and stationary ORP that aligns with the Bellman optimal policy. Furthermore, we illustrate the necessity of $L^{\infty}$-norm when minimizing Bellman error to attain ORP. This finding clarifies the vulnerability of prior DRL algorithms that target the Bellman optimal policy with $L^{1}$-norm and motivates us to train a Consistent Adversarial Robust Deep Q-Network (CAR-DQN) by minimizing a surrogate of Bellman Infinity-error. The top-tier performance of CAR-DQN across various benchmarks validates its practical effectiveness and reinforces the soundness of our theoretical analysis. Haoran Li 0027, Congying Han, Yudong Hu, Tiande Guo, Shichen Liao |
ICML | 4 |
| 2024 | Fusion of Multi-level Information: Solve Large-Scale Traveling Salesman Problem with an Efficient Framework
Wenzhao Liu, Congying Han, Tiande Guo, Haoran Li 0027 |
ICONIP (4) | 2 |
| 2024 | Subspace Newton method for sparse group ℓ 0 optimization problem
Shichen Liao, Congying Han, Tiande Guo, Bonan Li |
J. Glob. Optim. | 2 |
| 2024 | Stochastic linearized generalized alternating direction method of multipliers: Expected convergence rates and large deviation propertiesabstractAbstract Alternating direction method of multipliers (ADMM) receives much attention in the field of optimization and computer science, etc. The generalized ADMM (G-ADMM) proposed by Eckstein and Bertsekas incorporates an acceleration factor and is more efficient than the original ADMM. However, G-ADMM is not applicable in some models where the objective function value (or its gradient) is computationally costly or even impossible to compute. In this paper, we consider the two-block separable convex optimization problem with linear constraints, where only noisy estimations of the gradient of the objective function are accessible. Under this setting, we propose a stochastic linearized generalized ADMM (called SLG-ADMM) where two subproblems are approximated by some linearization strategies. And in theory, we analyze the expected convergence rates and large deviation properties of SLG-ADMM. In particular, we show that the worst-case expected convergence rates of SLG-ADMM are $\mathcal{O}\left( {{N}^{-1/2}}\right)$ and $\mathcal{O}\left({\ln N} \cdot {N}^{-1}\right)$ for solving general convex and strongly convex problems, respectively, where N is the iteration number, similarly hereinafter, and with high probability, SLG-ADMM has $\mathcal{O}\left ( \ln N \cdot N^{-1/2} \right ) $ and $\mathcal{O}\left ( \left ( \ln N \right )^{2} \cdot N^{-1} \right ) $ constraint violation bounds and objective error bounds for general convex and strongly convex problems, respectively. Tiande Guo, Congying Han |
Math. Struct. Comput. Sci. | 3 |
| 2024 | STDNet: Rethinking Disentanglement Learning With Information TheoryabstractDisentangled representation learning is typically achieved by a generative model, variational encoder (VAE). Existing VAE-based methods try to disentangle all the attributes simultaneously in a single hidden space, while the separation of the attribute from irrelevant information varies in complexity. Thus, it should be conducted in different hidden spaces. Therefore, we propose to disentangle the disentanglement itself by assigning the disentanglement of each attribute to different layers. To achieve this, we present a stair disentanglement net (STDNet), a stair-like structure network with each step corresponding to the disentanglement of an attribute. An information separation principle is employed to peel off the irrelevant information to form a compact representation of the targeted attribute within each step. Compact representations, thus, obtained together form the final disentangled representation. To ensure the final disentangled representation is compressed as well as complete with respect to the input data, we propose a variant of the information bottleneck (IB) principle, the stair IB (SIB) principle, to optimize a tradeoff between compression and expressiveness. In particular, for the assignment to the network steps, we define an attribute complexity metric to assign the attributes by the complexity ascending rule (CAR) that dictates a sequencing of the attribute disentanglement in ascending order of complexity. Experimentally, STDNet achieves state-of-the-art results in representation learning and image generation on multiple benchmarks, including Mixed National Institute of Standards and Technology database (MNIST), dSprites, and CelebA. Furthermore, we conduct thorough ablation experiments to show how the strategies employed here contribute to the performance, including neurons block, CAR, hierarchical structure, and variational form of SIB. Ziwen Liu 0001, Mingqiang Li, Congying Han, Siqi Tang, Tiande Guo |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | The Regularized Submodular Maximization via the Lyapunov Method
Congying Han, Dachuan Xu 0001, Yang Zhou 0018 |
COCOON (2) | 2 |
| 2023 | DropKey for Vision TransformerabstractIn this paper, we focus on analyzing and improving the dropout technique for self-attention layers of Vision Transformer, which is important while surprisingly ignored by prior works. In particular, we conduct researches on three core questions: First, what to drop in self-attention layers? Different from dropping attention weights in literature, we propose to move dropout operations forward ahead of attention matrix calculation and set the Key as the dropout unit, yielding a novel dropout-before-softmax scheme. We theoretically verify that this scheme helps keep both regularization and probability features of attention weights, alleviating the overfittings problem to specific patterns and enhancing the model to globally capture vital information; Second, how to schedule the drop ratio in consecutive layers? In contrast to exploit a constant drop ratio for all layers, we present a new decreasing schedule that gradually decreases the drop ratio along the stack of self-attention layers. We experimentally validate the proposed schedule can avoid overfittings in low-level features and missing in high-level semantics, thus improving the robustness and stableness of model training; Third, whether need to perform structured dropout operation as CNN? We attempt patch-based block-version of dropout operation and find that this useful trick for CNN is not essential for ViT. Given exploration on the above three questions, we present the novel Drop-Key method that regards Key as the drop unit and exploits decreasing schedule for drop ratio, improving ViTs in a general way. Comprehensive experiments demonstrate the effectiveness of DropKey for various ViT architectures, e.g. T2T, VOLO, CeiT and DeiT, as well as for various vision tasks, e.g., image classification, object detection, human-object interaction detection and human body shape recovery. Bonan Li, Yinhan Hu, Xuecheng Nie, Congying Han, Xiangjian Jiang, Tiande Guo, Luoqi Liu |
CVPR | 4 |
| 2023 | Transforming Radiance Field with Lipschitz Network for Photorealistic 3D Scene StylizationabstractRecent advances in 3D scene representation and novel view synthesis have witnessed the rise of Neural Radiance Fields (NeRFs). Nevertheless, it is not trivial to exploit NeRF for the photorealistic 3D scene stylization task, which aims to generate visually consistent and photorealistic stylized scenes from novel views. Simply coupling NeRF with photorealistic style transfer (PST) will result in cross-view inconsistency and degradation of stylized view syntheses. Through a thorough analysis, we demonstrate that this non-trivial task can be simplified in a new light: When transforming the appearance representation of a pre-trained NeRF with Lipschitz mapping, the consistency and photorealism across source views will be seamlessly encoded into the syntheses. That motivates us to build a concise and flexible learning framework namely LipRF, which upgrades arbitrary 2D PST methods with Lipschitz mapping tailored for the 3D scene. Technically, LipRF first pre-trains a radiance field to reconstruct the 3D scene, and then emulates the style on each view by 2D PST as the prior to learn a Lipschitz network to stylize the pre-trained appearance. In view of that Lipschitz condition highly impacts the expressivity of the neural network, we devise an adaptive regularization to balance the reconstruction and stylization. A gradual gradient aggregation strategy is further introduced to optimize LipRF in a cost-efficient manner. We conduct extensive experiments to show the high quality and robust performance of LipRF on both photorealistic 3D stylization and object appearance editing. Yinglu Liu, Congying Han, Yingwei Pan, Tiande Guo, Ting Yao 0003 |
CVPR | 3 |
| 2023 | DiffBFR: Bootstrapping Diffusion Model for Blind Face RestorationabstractBlind face restoration (BFR) is important while challenging. Prior works prefer to exploit GAN-based frameworks to tackle this task due to the balance of quality and efficiency. However, these methods suffer from poor stability and adaptability to long-tail distribution, failing to simultaneously retain source identity and restore detail. In this paper, we propose to introduce Diffusion Probabilistic Model (DPM) for BFR to tackle the above problem, given its superiority over GAN in aspects of avoiding training collapse and generating long-tail distribution. We name the proposed framework as DiffBFR. In particular, DiffBFR utilizes a two-step design, that first restores identity information from low-quality images and then enhances texture details according to the distribution of real faces. This design is implemented with two key components: 1) Identity Restoration Module (IRM) for preserving the face details in results. Instead of denoising from pure Gaussian random distribution with LQ images as the condition during the reverse process, we propose a novel truncated sampling method which starts from LQ images with part noise added. We theoretically prove that this change shrinks the evidence lower bound of DPM and then restores more original details. With theoretical proof, two cascade conditional DPMs with different input sizes are introduced to strengthen this sampling effect and reduce training difficulty in the high-resolution image generated directly. 2) Texture Enhancement Module (TEM) for polishing the texture of the image. Here an unconditional DPM, a LQ-free model, is introduced to further force the restorations to appear realistic. We theoretically proved that this unconditional DPM trained on pure HQ images contributes to justifying the correct distribution of inference images output from IRM in pixel-level space. Concretely, truncated sampling with fractional time step is utilized to polish pixel-level textures while preserving identity information. Our experiments demonstrated that the proposed DiffBFR achieves significantly superior results to state-of-the-art methods both quantitatively and qualitatively. Xinmin Qiu, Congying Han, Bonan Li, Tiande Guo, Xuecheng Nie |
ACM Multimedia | 2 |
| 2023 | Towards Consistent Video Editing with Text-to-Image Diffusion ModelsabstractExisting works have advanced Text-to-Image (TTI) diffusion models for video editing in a one-shot learning manner. Despite their low requirements of data and computation, these methods might produce results of unsatisfied consistency with text prompt as well as temporal sequence, limiting their applications in the real world. In this paper, we propose to address the above issues with a novel EI$^2$ model towards Enhancing vIdeo Editing consIstency of TTI-based frameworks. Specifically, we analyze and find that the inconsistent problem is caused by newly added modules into TTI models for learning temporal information. These modules lead to covariate shift in the feature space, which harms the editing capability. Thus, we design EI$^2$ to tackle the above drawbacks with two classical modules: Shift-restricted Temporal Attention Module (STAM) and Fine-coarse Frame Attention Module (FFAM). First, through theoretical analysis, we demonstrate that covariate shift is highly related to Layer Normalization, thus STAM employs a Instance Centering layer replacing it to preserve the distribution of temporal features. In addition, STAM employs an attention layer with normalized mapping to transform temporal features while constraining the variance shift. As the second part, we incorporate STAM with a novel FFAM, which efficiently leverages fine-coarse spatial information of overall frames to further enhance temporal consistency. Extensive experiments demonstrate the superiority of the proposed EI$^2$ model. Bonan Li, Xuecheng Nie, Congying Han, Tiande Guo, Luoqi Liu |
NeurIPS | 4 |
| 2023 | Modeling opponent learning in multiagent repeated gamesabstractAbstract Multiagent reinforcement learning (MARL) has been used extensively in the game environment. One of the main challenges in MARL is that the environment of the agent system is dynamic, and the other agents are also updating their strategies. Therefore, modeling the opponents’ learning process and adopting specific strategies to shape learning is an effective way to obtain better training results. Previous studies such as DRON, LOLA and SOS approximated the opponent’s learning process and gave effective applications. However, these studies modeled only transient changes in opponent strategies and lacked stability in the improvement of equilibrium efficiency. In this article, we design the MOL (modeling opponent learning) method based on the Stackelberg game. We use best response theory to approximate the opponents’ preferences for different actions and explore stable equilibrium with higher rewards. We find that MOL achieves better results in several games with classical structures (the Prisoner’s Dilemma, Stackelberg Leader game and Stag Hunt with 3 players), and in randomly generated bimatrix games. MOL performs well in competitive games played against different opponents and converges to stable points that score above the Nash equilibrium in repeated game environments. The results may provide a reference for the definition of equilibrium in multiagent reinforcement learning systems, and contribute to the design of learning objectives in MARL to avoid local disadvantageous equilibrium and improve general efficiency. Yudong Hu, Congying Han, Haoran Li 0027, Tiande Guo |
Appl. Intell. | 2 |
| 2023 | Solving uncapacitated P-Median problem with reinforcement learning assisted by graph attention networks
Chenguang Wang 0011, Congying Han, Tiande Guo, Man Ding |
Appl. Intell. | 2 |
| 2023 | Learning graph representation by aggregating subgraphs via mutual information maximization
Ziwen Liu 0001, Chenguang Wang 0011, Congying Han, Tiande Guo |
Neurocomputing | 3 |
| 2023 | A learnable sampling method for scalable graph neural networks
Tiande Guo, Xiaoxi Yu, Congying Han |
Neural Networks | 4 |
| 2022 | Shrinking Temporal Attention in Transformers for Video Action RecognitionabstractSpatiotemporal modeling in an unified architecture is key for video action recognition. This paper proposes a Shrinking Temporal Attention Transformer (STAT), which efficiently builts spatiotemporal attention maps considering the attenuation of spatial attention in short and long temporal sequences. Specifically, for short-term temporal tokens, query token interacts with them in a fine-grained manner in dealing with short-range motion. It then shrinks to a coarse attention in neighborhood for long-term tokens, to provide larger receptive field for long-range spatial aggregation. Both of them are composed in a short-long temporal integrated block to build visual appearances and temporal structure concurrently with lower costly in computation. We conduct thorough ablation studies, and achieve state-of-the-art results on multiple action recognition benchmarks including Kinetics400 and Something-Something v2, outperforming prior methods with 50% less FLOPs and without any pretrained model. Bonan Li, Pengfei Xiong, Congying Han, Tiande Guo |
AAAI | 3 |
| 2022 | PetsGAN: Rethinking Priors for Single Image GenerationabstractSingle image generation (SIG), described as generating diverse samples that have the same visual content as the given natural image, is first introduced by SinGAN, which builds a pyramid of GANs to progressively learn the internal patch distribution of the single image. It shows excellent performance in a wide range of image manipulation tasks. However, SinGAN has some limitations. Firstly, due to lack of semantic information, SinGAN cannot handle the object images well as it does on the scene and texture images. Secondly, the independent progressive training scheme is time-consuming and easy to cause artifacts accumulation. To tackle these problems, in this paper, we dig into the single image generation problem and improve SinGAN by fully-utilization of internal and external priors. The main contributions of this paper include: 1) We interpret single image generation from the perspective of the general generative task, that is, to learn a diverse distribution from the Dirac distribution composed of a single image. In order to solve this non-trivial problem, we construct a regularized latent variable model to formulate SIG. To the best of our knowledge, it is the first time to give a clear formulation and optimization goal of SIG, and all the existing methods for SIG can be regarded as special cases of this model. 2) We design a novel Prior-based end-to-end training GAN (PetsGAN), which is infused with internal prior and external prior to overcome the problems of SinGAN. For one thing, we employ the pre-trained GAN model to inject external prior for image generation, which can alleviate the problem of lack of semantic information and generate natural, reasonable and diverse samples, even for the object image. For another, we fully-utilize the internal prior by a differential Patch Matching module and an effective reconstruction network to generate consistent and realistic texture. 3) We construct abundant of qualitative and quantitative experiments on three datasets. The experimental results show our method surpasses other methods on both generated image quality, diversity, and training speed. Moreover, we apply our method to other image manipulation tasks (e.g., style transfer, harmonization) and the results further prove the effectiveness and efficiency of our method. Yinglu Liu, Congying Han, Hailin Shi, Tiande Guo |
AAAI | 3 |
| 2022 | DFS: A Diverse Feature Synthesis Model for Generalized Zero-Shot LearningabstractGenerative based strategy has shown great potential in the Generalized Zero-Shot Learning task. However, it suffers severe generalization problem due to lacking of feature diversity for unseen classes to train a good classifier. In this paper, we propose to enhance the generalizability of GZSL models via improving feature diversity of unseen classes. For this purpose, we present a novel Diverse Feature Synthesis (DFS) model. Different from prior works that solely utilize semantic knowledge in the generation process, DFS leverages visual knowledge with semantic one in a unified way, thus deriving class-specific diverse feature samples and leading to robust classifier for recognizing both seen and unseen classes in the testing phase. To simplify the learning, DFS represents visual and semantic knowledge in the aligned space, making it able to produce good feature samples with a low-complexity implementation. Accordingly, DFS is composed of two consecutive generators: an aligned feature generator, transferring semantic and visual representations into aligned features; a synthesized feature generator, producing diverse feature samples of unseen classes in the aligned space. We conduct comprehensive experiments to verify the efficacy of DFS. Results demonstrate its effectiveness to generate diverse features for unseen classes, leading to superior performance on multiple benchmarks. Code will be released upon acceptance. Bonan Li, Yinhan Hu, Congying Han, Tiande Guo |
ICPR | 3 |
| 2022 | Generalized One-shot Domain Adaptation of Generative Adversarial NetworksabstractThe adaptation of a Generative Adversarial Network (GAN) aims to transfer a pre-trained GAN to a target domain with limited training data. In this paper, we focus on the one-shot case, which is more challenging and rarely explored in previous works. We consider that the adaptation from a source domain to a target domain can be decoupled into two parts: the transfer of global style like texture and color, and the emergence of new entities that do not belong to the source domain. While previous works mainly focus on style transfer, we propose a novel and concise framework to address the \textit{generalized one-shot adaptation} task for both style and entity transfer, in which a reference image and its binary entity mask are provided. Our core idea is to constrain the gap between the internal distributions of the reference and syntheses by sliced Wasserstein distance. To better achieve it, style fixation is used at first to roughly obtain the exemplary style, and an auxiliary network is introduced to the generator to disentangle entity and style transfer. Besides, to realize cross-domain correspondence, we propose the variational Laplacian regularization to constrain the smoothness of the adapted generator. Both quantitative and qualitative experiments demonstrate the effectiveness of our method in various scenarios. Code is available at \url{https://github.com/zhangzc21/Generalized-One-shot-GAN-adaptation}. Yinglu Liu, Congying Han, Tiande Guo, Ting Yao 0003, Tao Mei 0001 |
NeurIPS | 3 |
| 2022 | Complexity Analysis of a Stochastic Variant of Generalized Alternating Direction Method of Multipliers
Tiande Guo, Congying Han |
TAMC | 3 |
| 2021 | ExSinGAN: Learning an Explainable Generative Model from a Single Image
Congying Han, Tiande Guo |
BMVC | 2 |
| 2021 | Disentangled features with direct sum decomposition for zero shot learning
Bonan Li, Congying Han, Tiande Guo, Tong Zhao 0004 |
Neurocomputing | 2 |
| 2021 | Facial depth descend: A generation paradigm for facial depth map
Congying Han, Hanqin Chen, Tiande Guo |
Neurocomputing | 2 |
| 2020 | A novel method based on deep learning for aligned fingerprints matching
Baicun Zhou, Congying Han, Tiande Guo |
Appl. Intell. | 3 |
| 2020 | Registration and matching method for directed point set with orientation attributes and local information
Congying Han, Tiande Guo, Tong Zhao 0004 |
Comput. Vis. Image Underst. | 2 |
| 2020 | Fast minutiae extractor using neural network
Baicun Zhou, Congying Han, Tiande Guo |
Pattern Recognit. | 2 |
| 2017 | Partial Fingerprint Matching via Phase-Only Correlation and Deep Convolutional Neural Network
Siqi Tang, Congying Han, Tiande Guo |
ICONIP (6) | 3 |
| 2016 | A novel fingerprint classification method based on deep learningabstractFingerprint classification is an effective technique for reducing the candidate numbers of fingerprints in the stage of matching in automatic fingerprint identification system (AFIS). In recent years, deep learning is an emerging technology which has achieved great success in many fields, such as image processing, computer vision. In this paper, we have a preliminary attempt on the traditional fingerprint classification problem based on the new depth neural network method. For the four-class problem, only choosing orientation field as the classification feature, we achieve 91.4% accuracy using the stacked sparse autoencoders (SAE) with three hidden layers in the NIST-DB4 database. And then two classification probabilities are used for fuzzy classification which can effectively enhance the accuracy of classification. By only adjusting the probability threshold, we get the accuracy of classification is 96.1% (setting threshold is 0.85), 97.2% (setting threshold is 0.90) and 98.0% (setting threshold is 0.95) with a single layer SAE. Applying the fuzzy method, we obtain higher accuracy. Congying Han, Tiande Guo |
ICPR | 2 |
| 2005 | The Automatic Generation of Basis Set of Path for Path TestingabstractBasis set of path is consisted of some of the program’s paths. The automatic generation method of basis set of path is discussed in this paper. It is built by searching the control flow graph of a program by depth-first searching method. In order to avoiding that the algorithm will never stop and reducing the searching procedure, the sub-path from the multi-indegree nodes to the end node of a program and the sub-path that contains a loop is recorded during the construction of a basis path. Some new basis paths can be constructed by merging these two kinds of sub-paths. Guangmei Zhang, Chen Rui, Xiaowei Li 0001, Congying Han |
Asian Test Symposium | 4 |