EDBT 2026 Demo / reviewers in the wild / expert
Yunlian Sun
dblp:03/3647
· DBLP profile ↗
21ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-4696-8848ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Security and privacy · 5 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene PerceptionabstractGenerating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and animation. Existing methods for human-object interaction often neglect scene context, leading to implausible penetrations, while human-scene interaction approaches struggle to coordinate fine-grained manipulations with long-range navigation. To address these limitations, we propose HOSIG, a novel framework for synthesizing full-body interactions through hierarchical scene perception. Our method decouples the task into three key components: 1) a scene-aware grasp pose generator that ensures collision-free whole-body postures with precise hand-object contact by integrating local geometry constraints, 2) a heuristic navigation algorithm that autonomously plans obstacle-avoiding paths in complex indoor environments via compressed 2D floor maps and dual-component spatial reasoning, and 3) a scene-guided motion diffusion model that generates trajectory-controlled, full-body motions with finger-level accuracy by incorporating spatial anchors and dual-space gradient-based guidance. Extensive experiments on the TRUMANS dataset demonstrate superior performance over state-of-the-art methods. Notably, our framework supports unlimited motion length through autoregressive generation and requires minimal manual intervention. This work bridges the critical gap between scene-aware navigation and dexterous object manipulation, advancing the frontier of embodied interaction synthesis. Yunlian Sun, Hongwen Zhang 0001, Yebin Liu, Jinhui Tang 0001 |
AAAI | 2 |
| 2026 | HiTMM: Generative Temporal Masked Modeling of Human Interactive MotionsabstractWe have recently seen some progress in the current field of human-human interaction generation. However, directly generating complex two-person interactive motions remains a significant challenge. Meanwhile, these models typically employ two independent timelines when generating motions for interactive scenarios involving two individuals. This design overlooks the temporal dependencies between motions at each timestep and fails to account for the roles of active and reactive participants during the generation process, often resulting in unrealistic and unnatural motions. In this work, we propose HiTMM, a novel framework for Human interaction generation based on Temporal Masked Modeling. HiTMM first decomposes the human interaction into two separate single-person motions. Individual motions within the interaction belong to the same type, enabling them to be mapped to a shared latent space through a coarse-to-fine approach that produces multi-layer discrete tokens. We then arrange all tokens of the two interacting individuals along a shared timeline. Subsequently, we employ a masked transformer and a residual transformer to model the base-layer and rest-layer motion tokens. Both the base-layer and rest-layer motion tokens are arranged along a single timeline, allowing the model to explicitly capture the temporal order and initiating role embedded in the sequence, where the first individual's motion initiates the interaction. Note that, our model utilizes a shared temporal representation, making it capable of performing temporal editing on specific regions within human interaction sequences. Experimental results show that our model achieves an FID of 5.017 on the InterHuman dataset, surpassing the current state-of-the-art model (vs 5.154 for InterMask), and an FID of 0.373 on the InterX dataset (vs 0.399 for InterMask). Zicheng Jiao, Yunlian Sun, Hongwen Zhang 0001, Jinhui Tang 0001, Massimo Tistarelli |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | EDMG: Towards Efficient Long Dance Motion Generation with Fundamental Movements from Dance GenresabstractDance is an important art form in human culture, but creating new dances can be both challenging and time-consuming. In this paper, we propose a novel dance choreography framework, EDMG, designed to efficiently generate creative and long-lasting dance sequences conditioning on music and dance descriptions. In the first stage, we propose a flexible dance diffusion method, combined with dance genre description and descriptions of fundamental movements to generate the dance sequences. To achieve high computational efficiency and inference speed, EDMG designs a lightweight denoising module by using selective parallel scanning algorithm from Mamba2. This Parallel Mamba Denoiser reduces significantly the number of parameters and accelerates remarkably both the learning and inference processes. In the second stage, by designing a smoothing module with a long receptive field, we mitigate joint error accumulation that causes jittering movements and foot sliding, thereby enhancing the fluency and visual appeal of the dance movements. Furthermore, we extend the AIST++ dataset by adding detailed descriptions of dance genres and fundamental movements, using the Large Language Model (LLM). These descriptions further improve the choreography generation. EDMG is validated through extensive experiments, demonstrating that our method can both effectively and efficiently generate long-term dances suitable for various dance genres. Project URL: https://github.com/neymar277/EDMG. Yunlian Sun, Hongwen Zhang 0001, Jinhui Tang 0001 |
ACM Multimedia | 2 |
| 2024 | FG-MDM: Towards Zero-Shot Human Motion Generation via ChatGPT-Refined Descriptions
Chuanchen Luo, Junran Peng, Hongwen Zhang 0001, Yunlian Sun |
ICPR (28) | 6 |
| 2024 | STAF: 3D Human Mesh Recovery From Video With Spatio-Temporal Alignment FusionabstractThe recovery of 3D human mesh from monocular images has significantly been developed in recent years. However, existing models usually ignore spatial and temporal information, which might lead to mesh and image misalignment and temporal discontinuity. For this reason, we propose a novel Spatio-Temporal Alignment Fusion (STAF) model. As a video-based model, it leverages coherence clues from human motion by an attention-based Temporal Coherence Fusion Module (TCFM). As for spatial mesh-alignment evidence, we extract fine-grained local information through predicted mesh projection on the feature maps. Based on the spatial features, we further introduce a multi-stage adjacent Spatial Alignment Fusion Module (SAFM) to enhance the feature representation of the target frame. In addition to the above, we propose an Average Pooling Module (APM) to allow the model to focus on the entire input sequence rather than just the target frame. This method can remarkably improve the smoothness of recovery results from video. Extensive experiments on 3DPW, MPII3D, and H36M demonstrate the superiority of STAF. We achieve a state-of-the-art trade-off between precision and smoothness. Our code and more video results are on the project pagehttps://yw0208.github.io/staf/. Hongwen Zhang 0001, Yunlian Sun, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | SI-Net: spatial interaction network for deepfake detection
Jian Wang 0129, Xiaoyu Du 0002, Yunlian Sun, Jinhui Tang 0001 |
Multim. Syst. | 4 |
| 2023 | Boosting Few-Shot Fine-Grained Recognition With Background Suppression and Foreground AlignmentabstractFew-shot fine-grained recognition (FS-FGR) aims to recognize novel fine-grained categories with the help of limited available samples. Undoubtedly, this task inherits the main challenges from both few-shot learning and fine-grained recognition. First, the lack of labeled samples makes the learned model easy to overfit. Second, it also suffers from high intra-class variance and low inter-class differences in the datasets. To address this challenging task, we propose a two-stage background suppression and foreground alignment framework, which is composed of a background activation suppression (BAS) module, a foreground object alignment (FOA) module, and a local-to-local (L2L) similarity metric. Specifically, the BAS is introduced to generate a foreground mask for localization to weaken background disturbance and enhance dominative foreground objects. The FOA then reconstructs the feature map of each support sample according to its correction to the query ones, which addresses the problem of misalignment between support-query image pairs. To enable the proposed method to have the ability to capture subtle differences in confused samples, we present a novel L2L similarity metric to further measure the local similarity between a pair of aligned spatial features in the embedding space. What’s more, considering that background interference brings poor robustness, we infer the pairwise similarity of feature maps using both the raw image and the refined image. Extensive experiments conducted on multiple popular fine-grained benchmarks demonstrate that our method outperforms the existing state of the art by a large margin. The source codes are available at:https://github.com/CSer-Tang-hao/BSFA-FSFG. Zican Zha, Hao Tang 0007, Yunlian Sun, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Depth-Based Ensemble Learning Network For Face Anti-SpoofingabstractAlthough significant progress has been made in face anti-spoofing, current methods can only achieve satisfactory results under intra-dataset settings. In other words, they tend to perform poorly when suffering from unseen attacks. Previous methods try to extract a common feature space from multiple domains, but this idea is inefficient due to the enormous distribution difference among training domains. Unlike previous methods, we assume that the data distribution of the target domain will be similar to one of the training domains. Based on this hypothesis, we draw on the idea of ensemble learning and propose a generalized framework with multiple domain-specific modules. Given a test sample, the proposed framework allows it to dynamically choose which module to use based on its similarity to the training domains. In addition, we employ GCBlock to better mine face depth information for auxiliary supervision. Since fake information is spread throughout the image, we further introduce DropBlock to avoid overfitting. Extensive experiments on four public datasets show that our approach is practical. Yunlian Sun |
ICASSP | 2 |
| 2022 | Ganet: Unary Attention Reaches Pairwise Attention Via Implicit Group Clustering in Light-Weight CNNsabstractThe attention mechanism has been widely explored to construct a long-range connection which is beyond the realm of convolutions. The two groups of attention, unary and pair-wise attention, seem like being incompatible as fire and water due to the completely different operations. In this paper, we propose a Group Attention (GA) block to bridge the gap between these two attentions and merely leverage unary attention to lightweightly reach the effect of pairwise attention, based on the implicit group clustering of light-weight CNNs. Compared with the conventional pairwise attention, i.e, Non-Local networks, our method artfully bypasses the burdensome pixel-pair calculation to save a huge computational cost, that is a big advantage of our work. Experiments on the task of image classification demontrate the effectiveness and efficiency of our GA block to enhance the light-weight models. Code will be released at https://github.com/ChiSuWq/GANet. Cheng Zhuang, Yunlian Sun |
ICASSP | 2 |
| 2022 | LiSiam: Localization Invariance Siamese Network for Deepfake DetectionabstractAdvances in facial manipulation technology have led to increasing indistinguishable and realistic face swap videos, which raises growing concerns about the security risk of deepfakes in the community. Although current deepfake detectors can gain promising performance when handling high-quality faces under within-database settings, most detectors suffer from performance degradation in cross-database evaluation. Moreover, when test faces’ quality is different from training faces, the performance degrades even under within-database settings. To this end, we propose a novel Localization invariance Siamese Network (LiSiam) to enforce localization invariance against different image degradation for deepfake detection. Specifically, our Siamese network-based feature extractor takes the original image and the corresponding quality-degraded image as pairwise inputs and outputs two segmentation maps. A localization invariance loss is further proposed to impose localization consistency between the two segmentation maps. In addition, we design a Mask-guided Transformer to capture the co-occurrence between the forgery region and its surroundings. Finally, a multi-task learning strategy is utilized to obtain a robust and discriminative feature representation and jointly optimize multiple objective functions (i.e., segmentation, classification, and localization invariance losses) in an end-to-end manner. Experimental results on two public datasets, i.e., FaceForensics++ and Celeb-DF, demonstrate the superior performance of our proposed method to state-of-the-art methods. Jian Wang 0129, Yunlian Sun, Jinhui Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Attention-aware conditional generative adversarial networks for facial age synthesis
Xiahui Chen, Yunlian Sun, Xiangbo Shu |
Neurocomputing | 2 |
| 2021 | Host-Parasite: Graph LSTM-in-LSTM for Group Activity RecognitionabstractThis article aims to tackle the problem of group activity recognition in the multiple-person scene. To model the group activity with multiple persons, most long short-term memory (LSTM)-based methods first learn the person-level action representations by several LSTMs and then integrate all the person-level action representations into the following LSTM to learn the group-level activity representation. This type of solution is a two-stage strategy, which neglects the "host-parasite" relationship between the group-level activity ("host") and person-level actions ("parasite") in spatiotemporal space. To this end, we propose a novel graph LSTM-in-LSTM (GLIL) for group activity recognition by modeling the person-level actions and the group-level activity simultaneously. GLIL is a "host-parasite" architecture, which can be seen as several person LSTMs (P-LSTMs) in the local view or a graph LSTM (G-LSTM) in the global view. Specifically, P-LSTMs model the person-level actions based on the interactions among persons. Meanwhile, G-LSTM models the group-level activity, where the person-level motion information in multiple P-LSTMs is selectively integrated and stored into G-LSTM based on their contributions to the inference of the group activity class. Furthermore, to use the person-level temporal features instead of the person-level static features as the input of GLIL, we introduce a residual LSTM with the residual connection to learn the person-level residual features, consisting of temporal features and static features. Experimental results on two public data sets illustrate the effectiveness of the proposed GLIL compared with state-of-the-art methods. Xiangbo Shu, Liyan Zhang 0001, Yunlian Sun, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | CAN-GAN: Conditioned-attention normalized GAN for face age synthesis
Chenglong Shi, Jiachao Zhang, Yazhou Yao, Yunlian Sun, Huaming Rao, Xiangbo Shu |
Pattern Recognit. Lett. | 4 |
| 2020 | Recursive Discriminative Subspace Learning With $\ell_{1}$ -Norm Distance ConstraintabstractIn feature learning tasks, one of the most enormous challenges is to generate an efficient discriminative subspace. In this paper, we propose a novel subspace learning method, named recursive discriminative subspace learning with an ℓ1-norm distance constraint (RDSL). RDSL can robustly extract features from the contaminated images and learn a discriminative subspace. With the use of an inequation-based ℓ1-norm distance metric constraint, the minimized ℓ1-norm distance metric objective function with slack variables induces samples in the same class to cluster as close as possible, meanwhile samples from different classes can be separated from each other as far as possible. By utilizing ℓ1-norm items in both the objective function and the constraint, RDSL can well handle the noisy data and outliers. In addition, the large margin formulation makes the proposed method insensitive to initializations. We describe two approaches to solve RDSL with a recursive strategy. Experimental results on six benchmark datasets, including the original data and the contaminated data, demonstrate that RDSL outperforms the state-of-the-art methods. Yunlian Sun, Qiaolin Ye, Jinhui Tang 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | Facial Age Synthesis With Label Distribution-Guided Generative Adversarial NetworkabstractThe existing research work on facial age synthesis has been mostly focused on long-term aging (e.g., over an age span of 10 years or more). In this paper, we employ generative adversarial networks (GANs) as a tool to investigate age synthesis over different age spans. Compared with long-term aging, short-term age synthesis suffers from the reduced amount of available training data, which can severely hinder the model training. We conduct a series of experiments to validate this. To facilitate short-term age synthesis, we further propose label distribution-guided generative adversarial network (ldGAN), where each sample is associated with an age label distribution (ALD) rather than a single age group. Accordingly, each sample can contribute not only to the learning of its own age group but also to neighbouring groups' learning. This is useful when addressing short-term aging to cope with the reduced amount of training data. In addition, unlike one-hot encoding which treats age groups as independent from one another, ldGAN can well capture the correlation among different age groups, so that smooth aging sequences can be achieved. The ALD model is integrated into GAN with a two-step process. Firstly, instead of the traditional one-hot encoding, ALD is applied as the condition of the generator. Secondly, we add a sequence of label distribution learners on top of several multi-scale discriminators, with the aim of minimizing the label distribution learning loss when optimizing both the generator and discriminators. Both qualitative and quantitative evaluations are conducted to assess ldGAN's ability in dealing with two core issues of face aging, i.e., aging effect generation and identity preservation. The obtained experimental results demonstrate the effectiveness of ldGAN in both learning short-term aging patterns and coping with the lack of training data. Yunlian Sun, Jinhui Tang 0001, Xiangbo Shu, Zhenan Sun, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Facial Age and Expression Synthesis Using Ordinal Ranking Adversarial NetworksabstractFacial image synthesis has been extensively studied, for a long time, in both computer graphics and computer vision. Particularly, the synthesis of face images with varying ages, expressions and poses has received an increasing attention owing to several real-world applications. In this paper, facial age and expression synthesis are addressed. While previous and current research papers on facial age synthesis mostly adopt an age span of 10 years, this paper investigates face aging with a shorter time span. For expression synthesis, given a neutral face, we work on synthesizing faces with varying expression intensities (e.g., from zero to high). Note that both human ages and expression intensities are inherently ordinal. To fully exploit this ordinal nature, we devise ordinal ranking generative adversarial networks (ranking GAN). For each face, a one-hot label is assigned to define its age range/expression intensity. By exploiting the relative order information among age ranges/expression intensities, a binary ranking vector is further computed for each face. In ranking GAN, one-hot labels are used as the condition of the generator for synthesizing faces with target age groups/expression intensities. Moreover, we add a sequence of cost-sensitive ordinal rankers on top of several multi-scale discriminators, with the aim of minimizing age/intensity rank estimation loss when optimizing both the generator and discriminators. In order to evaluate the proposed ranking GAN, extensive experiments are carried out on several public face databases. As demonstrated by the experimental testing, this ranking scheme performs well even when the amount of available labeled training data is limited. The reported experimental results well demonstrate the effectiveness of ranking GAN on synthesizing face aging sequences and faces with varying expression intensities. Yunlian Sun, Jinhui Tang 0001, Zhenan Sun, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Tracking the evolution of overlapping communities in dynamic social networks
Zechao Li, Guan Yuan, Yunlian Sun, Xiaobin Rui, Xinguang Xiang |
Knowl. Based Syst. | 4 |
| 2018 | Demographic Analysis from Biometric Data: Achievements, Challenges, and New FrontiersabstractBiometrics is the technique of automatically recognizing individuals based on their biological or behavioral characteristics. Various biometric traits have been introduced and widely investigated, including fingerprint, iris, face, voice, palmprint, gait and so forth. Apart from identity, biometric data may convey various other personal information, covering affect, age, gender, race, accent, handedness, height, weight, etc. Among these, analysis of demographics (age, gender, and race) has received tremendous attention owing to its wide real-world applications, with significant efforts devoted and great progress achieved. This survey first presents biometric demographic analysis from the standpoint of human perception, then provides a comprehensive overview of state-of-the-art advances in automated estimation from both academia and industry. Despite these advances, a number of challenging issues continue to inhibit its full potential. We second discuss these open problems, and finally provide an outlook into the future of this very active field of research by sharing some promising opportunities. Yunlian Sun, Man Zhang 0005, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Complementary Cohort Strategy for Multimodal Face Pair MatchingabstractFace pair matching is the task of determining whether two face images represent the same person. Due to the limited expressive information embedded in the two face images as well as various sources of facial variations, it becomes a quite difficult problem. Toward the issue of few available images provided to represent each face, we propose to exploit an extra cohort set (identities in the cohort set are different from those being compared) by a series of cohort list comparisons. Useful cohort coefficients are then extracted from both sorted cohort identities and sorted cohort images for complementary information. To augment its robustness to complicated facial variations, we further employ multiple face modalities owing to their complementary value to each other for the face pair matching task. The final decision is made by fusing the extracted cohort coefficients with the direct matching score for all the available face modalities. To investigate the capacity of each individual modality on matching faces, the cohort behavior, and the performance achieved using our complementary cohort strategy, we conduct a set of experiments on two recently collected multimodal face databases. It is shown that using different modalities leads to different face pair matching performance. For each modality, employing our cohort scheme significantly reduces the equal error rate. By applying the proposed multimodal complementary cohort strategy, we achieve the best performance on our face pair matching task. Yunlian Sun, Kamal Nasrollahi, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | On the Use of Discriminative Cohort Score Normalization for Unconstrained Face RecognitionabstractFacial imaging has been largely addressed for automatic personal identification, in a variety of different environments. However, automatic face recognition becomes very challenging whenever the acquisition conditions are unconstrained. In this paper, a picture-specific cohort normalization approach, based on polynomial regression, is proposed to enhance the robustness of face matching under challenging conditions. A careful analysis is presented to better understand the actual discriminative power of a given cohort set. In particular, it is shown that the cohort polynomial regression alone conveys some discriminative information on the matching face pair, which is just marginally worse than the raw matching score. The influence of the cohort set size in the matching accuracy is also investigated. Further, tests performed on the Face Recognition Grand Challenge ver 2 database and the labeled faces in the wild database allowed to determine the relation between the quality of the cohort samples and cohort normalization performance. Experimental results obtained from the LFW data set demonstrate the effectiveness of the proposed approach to improve the recognition accuracy in unconstrained face acquisition scenarios. Massimo Tistarelli, Yunlian Sun, Norman Poh |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2003 | Training integrate-and-fire neurons with the Informax principle IIabstractFor pt I see J. Phys. A, vol. 35, p. 2379-94 (2002).We develop neuron learning rules using the Informax principle together with the input-output relationship of the integrate-and-fire (IF) model with Poisson inputs. The learning rule is then tested with constant inputs, time-varying inputs and images. For constant inputs, it is found that, under the Informax principle, a network of IF models with initially all positive weights tends to disconnect some connections between neurons. For time-varying inputs and images, we perform signal separation tasks called independent component analysis. Numerical simulations indicate that some number of inhibitory inputs improves the performance of the system in both biological and engineering senses. Jianfeng Feng, Yunlian Sun, Hilary Buxton |
IEEE Trans. Neural Networks | 2 |