EDBT 2026 Demo / reviewers in the wild / expert
Yuyao Yan
dblp:191/0439
· DBLP profile ↗
22ranked-venue papers
3as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
Chaolong Yang, Yuyao Yan, Chenru Jiang, Weiguang Zhao, Jie Sun 0024, Bin Dong 0003, Kaizhu Huang |
Int. J. Comput. Vis. | 3 |
| 2026 | Point2pix-Zero: Point-driven refined diffusion for multi-object image editing
Yuyao Yan, Kaizhu Huang |
Pattern Recognit. | 2 |
| 2025 | From 2D Images to 3D Model: Weakly Supervised Multi-View Face Reconstruction with Deep FusionabstractWhile weakly supervised multi-view face reconstruction (MVR) is garnering increased attention, one critical issue still remains open: how to effectively interact and fuse multiple image information to reconstruct high-precision 3D models. In this regard, we propose a novel pipeline called Deep Fusion MVR (DF-MVR) to explore the feature correspondences between multi-view images and reconstruct high-precision 3D faces. Specifically, we present a novel multi-view feature fusion backbone that utilizes face masks to align features from multiple encoders and integrates one multi-layer attention mechanism to enhance feature interaction and fusion, resulting in one unified facial representation. Additionally, we develop one concise face mask mechanism that facilitates multi-view feature fusion and facial reconstruction by identifying common areas and guiding the network’s focus on critical facial features (e.g., eyes, brows, nose, and mouth). Experiments on Pixel-Face and Bosphorus datasets indicate the superiority of the proposed method. Without the 3D annotation, DF-MVR achieves relative 5.2% and 3.0% RMSE improvement over the existing weakly supervised MVRs, respectively, on Pixel-Face and Bosphorus datasets. Our code is available at https://github.com/weiguangzhao/DF_MVR. Weiguang Zhao, Chaolong Yang, Jianan Ye, Rui Zhang 0012, Yuyao Yan, Xi Yang 0008, Bin Dong 0003, Amir Hussain 0001, Kaizhu Huang |
ICME | 5 |
| 2025 | Stock Price Prediction with Attention-Based Framework by Integrating LLM-Generated Features
Yining Sun, Penglei Gao, Yuyao Yan, Xi Yang 0008 |
ICONIP (5) | 3 |
| 2025 | Entropy-Guided Distillation for Medical Image Segmentation Under Missing Modalities
Yuyao Yan, Xi Yang 0008, Kaizhu Huang |
ICONIP (5) | 2 |
| 2025 | Towards Training-Free Open-World Classification with 3D Generative Modelsabstract3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring robust subsequent knowledge adaptation capabilities. While current approaches predominantly rely on 2D pre-trained models through 3D-to-2D projection, their performance degrades severely under arbitrary object orientations. Unlike these present efforts, this work makes a pioneering exploration of 3D generative models for 3D open-world classification-specifically, leverageing the accumulated prior knowledge from these models to provide anchors for novel categories, while integrating a rotation-invariant feature extractor. This innovative synergy endows our pipeline with the advantages of being training-free and pose-invariant, thus well suited to adapt novel categories in 3D open-world classification. Extensive experiments on benchmark datasets demonstrate the potential of this pipeline, achieving state-of-the-art performance on ModelNet10‡ and McGill‡ with 32.7% and 8.7% overall accuracy improvement, respectively. The code is available in the supplementary materials. Xinzhe Xia, Weiguang Zhao, Yuyao Yan, Guanyu Yang 0002, Rui Zhang 0012, Kaizhu Huang, Xi Yang 0008 |
ACM Multimedia | 3 |
| 2025 | KDTalker++: Controllable Talking Portrait Generation with Audio, Text, and Expression EditingabstractThis work presents KDTalker++, a real-time system for generating talking portrait videos from a single image using audio or text input. Built on a keypoint-based spatiotemporal diffusion model, it adds voice cloning, background editing, and fine-grained expression control. The demo is available at https://kdtalker.com. A live presentation video is available at https://drive.google.com/file/d/1N4Ggu0Y32DTsV3mbKhXS4kGY6l2ZYKSp/view. Chaolong Yang, Yinuo Guo, Yuyao Yan, Jie Sun 0024, Kaizhu Huang |
ACM Multimedia | 4 |
| 2025 | DeepMethyGene: a deep-learning model to predict gene expression using DNA methylationsabstractGene expression is the basis for cells to achieve various functions, while DNA methylation constitutes a critical epigenetic mechanism governing gene expression regulation. Here we propose DeepMethyGene, an adaptive recursive convolutional neural network model based on ResNet that predicts gene expression using DNA methylation information. Our model transforms methylation Beta values to M values for Gaussian distributed data optimization, dynamically adjusts the output channels according to input dimension, and implements residual blocks to mitigate the problem of gradient vanishing when training very deep networks. Benchmarking against the state-of-the-art geneEXPLORE model (R2 = 0.449), DeepMethyGene (R2 = 0.640) demonstrated superior predictive performance. Further analysis revealed that the number of methylation sites and the average distance between these sites and gene transcription start sites (TSS) significantly affected the prediction accuracy. By exploring the complex relationship between methylation and gene expression, this study provides theoretical support for disease progression prediction and clinical intervention. Relevant data and code are available at https://github.com/yaoyao-11/DeepMethyGene . Yuyao Yan, Xinyi Chai, Wenran Li |
BMC Bioinform. | 1 |
| 2025 | Open-Pose 3D zero-shot learning: Benchmark and challenges
Weiguang Zhao, Guanyu Yang 0002, Rui Zhang 0012, Chenru Jiang, Chaolong Yang, Yuyao Yan, Amir Hussain 0001, Kaizhu Huang |
Neural Networks | 6 |
| 2024 | Enhancing Semantic Segmentation in Open Compound Domain Adaptation Through Mixed Image and Epistemic Uncertainty
Yiqun Ma, Siyuan Wang 0017, Xi Yang 0008, Yuyao Yan |
ICONIP (11) | 5 |
| 2024 | A deep top-down framework towards generalisable multi-view pedestrian detection
Ming Xu 0011, Yuchen Ling, Jeremy S. Smith, Yuyao Yan, Xinheng Wang 0001 |
Neurocomputing | 5 |
| 2024 | PPM: A boolean optimizer for data association in multi-view pedestrian detection
Ming Xu 0011, Yuyao Yan, Jeremy S. Smith, Yuchen Ling |
Pattern Recognit. | 3 |
| 2023 | Divide and Conquer: 3D Point Cloud Instance Segmentation With Point-Wise BinarizationabstractInstance segmentation on point clouds is crucially important for 3D scene understanding. Most SOTAs adopt distance clustering, which is typically effective but does not perform well in segmenting adjacent objects with the same semantic label (especially when they share neighboring points). Due to the uneven distribution of offset points, these existing methods can hardly cluster all instance points. To this end, we design a novel divide-and-conquer strategy named PBNet that binarizes each point and clusters them separately to segment instances. Our binary clustering divides offset instance points into two categories: high and low density points (HPs vs. LPs). Adjacent objects can be clearly separated by removing LPs, and then be completed and refined by assigning LPs via a neighbor voting method. To suppress potential over-segmentation, we propose to construct local scenes with the weight mask for each instance. As a plug-in, the proposed binary clustering can replace the traditional distance clustering and lead to consistent performance gains on many mainstream baselines. A series of experiments on ScanNetV2 and S3DIS datasets indicate the superiority of our model. In particular, PBNet ranks first on the ScanNetV2 official benchmark challenge, achieving the highest mAP. Code will be available publicly at https://github.com/weiguangzhao/PBNet. Weiguang Zhao, Yuyao Yan, Chaolong Yang, Jianan Ye, Xi Yang 0008, Kaizhu Huang |
ICCV | 2 |
| 2023 | Towards Deeper and Better Multi-view Feature Fusion for 3D Semantic Segmentation
Chaolong Yang, Yuyao Yan, Weiguang Zhao, Jianan Ye, Xi Yang 0008, Amir Hussain 0001, Bin Dong 0003, Kaizhu Huang |
ICONIP (15) | 2 |
| 2023 | Generalized image outpainting with U-transformer
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, John Yannis Goulermas, Yujie Geng, Yuyao Yan, Kaizhu Huang |
Neural Networks | 6 |
| 2023 | Semantic Similarity Distance: Towards better text-image consistency metric in text-to-image generation
Zhaorui Tan, Xi Yang 0008, Zihan Ye, Qiufeng Wang 0001, Yuyao Yan, Anh Nguyen 0003, Kaizhu Huang |
Pattern Recognit. | 5 |
| 2023 | Mind the Gap: Alleviating Local Imbalance for Unsupervised Cross-Modality Medical Image SegmentationabstractUnsupervised cross-modality medical image adaptation aims to alleviate the severe domain gap between different imaging modalities without using the target domain label. A key in this campaign relies upon aligning the distributions of source and target domain. One common attempt is to enforce the global alignment between two domains, which, however, ignores the fatal local-imbalance domain gap problem, i.e., some local features with larger domain gap are harder to transfer. Recently, some methods conduct alignment focusing on local regions to improve the efficiency of model learning. While this operation may cause a deficiency of critical information from contexts. To tackle this limitation, we propose a novel strategy to alleviate the domain gap imbalance considering the characteristics of medical images, namely Global-Local Union Alignment. Specifically, a feature-disentanglement style-transfer module first synthesizes the target-like source images to reduce the global domain gap. Then, a local feature mask is integrated to reduce the 'inter-gap' for local features by prioritizing those discriminative features with larger domain gap. This combination of global and local alignment can precisely localize the crucial regions in segmentation target while preserving the overall semantic consistency. We conduct a series of experiments with two cross-modality adaptation tasks, i,e. cardiac substructure and abdominal multi-organ segmentation. Experimental results indicate that our method achieves state-of-the-art performance in both tasks. Zixian Su, Xi Yang 0008, Qiufeng Wang 0001, Yuyao Yan, Jie Sun 0024, Kaizhu Huang |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | 3D Random Occlusion and Multi-layer Projection for Deep Multi-camera Pedestrian Localization
Ming Xu 0011, Yuyao Yan, Jeremy S. Smith, Xi Yang 0008 |
ECCV (10) | 3 |
| 2021 | Multicamera pedestrian detection using logic minimization
Yuyao Yan, Ming Xu 0011, Jeremy S. Smith, Mo Shen, Jin Xi |
Pattern Recognit. | 1 |
| 2020 | Moving shadow detection via binocular vision and colour clusteringabstractA pedestrian segmentation algorithm in the presence of cast shadows is presented in this study. The novelty of this algorithm lies in the fusion of multi‐view and multi‐plane homographic projections of foregrounds and the use of the fused data to guide colour clustering. This brings about an advantage over the existing binocular algorithms in that it can remove cast shadows while keeping pedestrians’ body parts, which occlude shadows. Phantom detection, which is inherent with the binocular method, is also investigated. Experimental results with real‐world videos have demonstrated the efficiency of this algorithm. Ming Xu 0011, Jeremy S. Smith, Yuyao Yan |
IET Comput. Vis. | 4 |
| 2019 | VSB-DVM: An End-to-End Bayesian Nonparametric Generalization of Deep Variational Mixture ModelabstractMixture of factor analyzers is a fundamental model in unsupervised learning, which is particularly useful for high dimensional data. Recent efforts on deep auto-encoding mixture models made a fruitful progress in clustering. However, in most cases, their performance depends highly on the results of pre-training. Moreover, they tend to ignore the prior information when making clustering assignment, leading to a less strict inference and consequently limiting the performance. In this paper, we propose an end-to-end Bayesian nonparametric generalization of deep mixture model with a Variational Auto-Encoder (VAE) framework. Specifically, we develop a novel model called VSB-DVM exploiting the Variational Stick-Breaking Process to design a Deep Variational Mixture Model. Distinct from the existing deep auto-encoding mixture models, this novel unsupervised deep generative model can learn low-dimensional representations and clustering simultaneously without pre-training. Importantly, a strict inference is proposed using weights of stick-breaking process in a variational way. Furthermore, able to capture the richer statistical structure of the data, VSB-DVM can also generate highly realistic samples for any specified cluster. A series of experiments are carried out, both qualitatively and quantitatively, on benchmark clustering and generation tasks. Comparative results show that the proposed model is able to generate diverse and high-quality samples of data, and also achieves encouraging clustering results outperforming the state-of-the-art algorithms on four real-world datasets. Xi Yang 0008, Yuyao Yan, Kaizhu Huang, Rui Zhang 0012 |
ICDM | 2 |
| 2017 | Multiview pedestrian localisation via a prime candidate chart based on occupancy likelihoodsabstractA sound way to localize occluded people is to project the foregrounds from multiple camera views to a reference view by homographies and find the foreground intersections. However, this may give rise to phantoms due to foreground intersections from different people. In this paper, each intersection region is warped back to the original camera view and is associated with a candidate box of the average size of pedestrians at that location. Then a joint occupancy likelihood is calculated for each intersection region. In the second step, essential candidate boxes are identified first, each of which covers at least a part of the foreground that is not covered by another candidate box. The non-essential candidate boxes are selected to cover the remaining foregrounds in the order of their joint occupancy likelihoods. Experiments on benchmark video datasets have demonstrated the good performance of our algorithm in comparison with other state-of-the-art methods. Yuyao Yan, Ming Xu 0011, Jeremy S. Smith |
ICIP | 1 |