EDBT 2026 Demo / reviewers in the wild / expert
Feihu Yan
dblp:194/2285
· DBLP profile ↗
21ranked-venue papers
2as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HGAvatar: One-shot high-quality 3D Gaussian head avatar
Feihu Yan, Guangzhe Zhao |
Comput. Vis. Image Underst. | 3 |
| 2026 | Progressive densification of 3D Gaussians for high-fidelity talking head synthesis
Xueni Guo, Feihu Yan, Guangzhe Zhao |
Frontiers Comput. Sci. | 3 |
| 2025 | An Unified Stochastic Fusion of Dual Diffusion Paths for Faithful Architectural Image SynthesisabstractArchitectural image generation aims to automate the creation of visually and semantically coherent architectural visuals based on inputs such as text prompts or sketches. While diffusion models have recently achieved remarkable success in this domain, they still struggle with preserving structural distortions and semantic misalignment under complex or multi-modal prompts. In this paper, we propose a lightweight and modular enhancement to existing diffusion pipelines that improves consistency and architectural plausibility without requiring external conditioning modules or fixed fine-tuning strategies. Our method introduces a text-image-image conditional generation framework, where two image generated from the same prompt are fused in latent space under unified stochastic control, enabling coherent sampling and minimizing semantic degradation. This approach is model-agnostic and compatible with optional adaptation modules, making it highly extensible and adaptable across different architectures and tasks. Experiments demonstrate that our method consistently produces images with improved semantic fidelity, structural consistency, and stylistic harmony. Furthermore, its modular design allows for efficient integration into existing workflows, offers a promising foundation for controllable and reliable architectural image synthesis. Guangzhe Zhao, Shilong Yang, Feihu Yan |
ICTAI | 3 |
| 2025 | Aligned and Detail Guided Retrieval: Multi-scale Fine-Grained Features Enhancement for End-to-End Person Search
Feihu Yan, Kunlin Zou, Zhong Zhou, Haiyong Chen |
ICXR | 2 |
| 2025 | Tri-Plane Dynamic Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisabstractABSTRACT Neural radiation field (NeRF) has been widely used in the field of talking portrait synthesis. However, the inadequate utilisation of audio information and spatial position leads to the inability to generate images with high audio‐lip consistency and realism. This paper proposes a novel tri‐plane dynamic neural radiation field (Tri‐NeRF) that employs an implicit radiation field to study the impacts of audio on facial movements. Specifically, Tri‐NeRF propose tri‐plane offset network (TPO‐Net) to offset spatial positions in three 2D planes guided by audio. This allows for sufficient learning of audio features from image features in a low dimensional state to generate more accurate lip movements. In order to better preserve facial texture details, we innovatively propose a new gated attention fusion module (GAF) to dynamically fuse features based on strong and weak correlation of cross‐modal features. Extensive experiments have demonstrated that Tri‐NeRF can generate talking portraits with audio‐lip consistency and realism. Xueni Guo, Feihu Yan, Guangzhe Zhao |
IET Image Process. | 5 |
| 2025 | StableID: Multimodal learning for stable identity in personalized Text-to-Face generation
Feihu Yan, Guangzhe Zhao |
Pattern Recognit. Lett. | 4 |
| 2025 | Syn-Net: A Synchronous Frequency-Perception Fusion Network for Breast Tumor Segmentation in Ultrasound ImagesabstractAccurate breast tumor segmentation in ultrasound images is a crucial step in medical diagnosis and locating the tumor region. However, segmentation faces numerous challenges due to the complexity of ultrasound images, similar intensity distributions, variable tumor morphology, and speckle noise. To address these challenges and achieve precise segmentation of breast tumors in complex ultrasound images, we propose a Synchronous Frequency-perception Fusion Network (Syn-Net). Initially, we design a synchronous dual-branch encoder to extract local and global feature information simultaneously from complex ultrasound images. Secondly, we introduce a novel Frequency- perception Cross-Feature Fusion (FrCFusion) Block, which utilizes Discrete Cosine Transform (DCT) to learn all-frequency features and effectively fuse local and global features while mitigating issues arising from similar intensity distributions. In addition, we develop a Full-Scale Deep Supervision method that not only corrects the influence of speckle noise on segmentation but also effectively guides decoder features towards the ground truth. We conduct extensive experiments on three publicly available ultrasound breast tumor datasets. Comparison with 14 state-of-the-art deep learning segmentation methods demonstrates that our approach exhibits superior sensitivity to different ultrasound images, variations in tumor size and shape, speckle noise, and similarity in intensity distribution between surrounding tissues and tumors. On the BUSI and Dataset B datasets, our method achieves better Dice scores compared to state-of-the-art methods, indicating superior performance in ultrasound breast tumor segmentation. Guangzhe Zhao, Xingguo Zhu, Feihu Yan, Maozu Guo 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Instance-level cross-attention learning for fine-grained customizable face generation
Feihu Yan, Guangzhe Zhao |
Vis. Comput. | 4 |
| 2024 | FTMSNet: Towards boundary-aware polyp segmentation framework based on hybrid Fourier Transform and Multi-scale SubtractionabstractColorectal cancer is one of the diseases with the highest incidence and mortality rates worldwide, posing a severe threat to human life and health. Polyps are the primary cause of colorectal cancer, Colonoscopy is the gold standard for diagnosing colorectal polyps. Accurate polyp segmentation is crucial for patient diagnosis, treatment, and prognosis. However, the irregular shapes and low contrast between colorectal polyps and normal tissue make polyp segmentation a challenging task. Although deep learning-based methods have achieved promising results in this task, few approaches focus on the boundary information of colorectal polyps. In this study, to address the issue of edge blurring in polyp segmentation, we propose a boundary-aware polyp segmentation method based on a hybrid Fourier Transform and Multi-scale Convolutional neural network, referred to as FTMSNet. We introduce the Fourier Transform Module (FTM), which utilizes the Fourier Transform to retain only high-frequency information (such as boundary) in the frequency domain. By leveraging boundary information, our method achieves more precise and clearer boundary delineation for colorectal polyp segmentation. Simultaneously, the Multi-scale Feature Denoising Decoder (MFDD) we introduced is devised to mitigate noise interference during multi-scale information fusion. We validated the performance of FTMSNet on five publicly available datasets. Extensive experimental results demonstrate that our approach surpasses the state-of-the-art polyp segmentation methods. Guangzhe Zhao, Feihu Yan, Maozu Guo 0001 |
BIBM | 4 |
| 2024 | YMamba: A Dual-Branch Network Fusing State Space Model and CNN for Medical Image SegmentationabstractIntelligent technologies like deep learning have significantly improved medical image segmentation, enhancing clinical decision-making and reducing healthcare costs. However, CNN-based methods face limitations due to constrained receptive fields and semantic information loss in deeper layers. Consequently, to address these issues, we innovatively propose the YMamba model which employs a parallel hybrid architecture of CNN and VMamba, enhancing the modeling capability of distant features while maintaining local feature detail textures in medical images without introducing additional parameters. Additionally, our proposed DBFM module employs an enhanced attention strategy to integrate the strengths of both methods more effectively, reinforcing image feature representation and mitigating background noise. Finally, the MCFFD module receives shallow and deep features from the fusion module, addressing the challenge of size variation in target segmentation regions within medical images. Extensive experiments demonstrate that YMamba achieves state-of-the-art results on four medical image datasets BUSI, DDTI, TN3K, and ISIC2016. Guangzhe Zhao, Feihu Yan, Maozu Guo 0001 |
BIBM | 3 |
| 2024 | ShardingSim: A Modular Committee-Based Sharding Blockchain SimulatorabstractBlockchain performance is crucial in research, with sharding emerging as an effective solution for scalability. By dividing the network into smaller shards, sharding facilitates faster transaction processing. However, there is currently a lack of effective simulators for modeling sharding blockchain performance and assessing shard load balancing. In this paper, we introduce ShardingSim, a modular, committee-based sharding blockchain simulator. ShardingSim simulates various sharding configurations and evaluates their performance across diverse network conditions and transaction datasets. We present a use case by modeling and simulating RapidChain with ShardingSim: simulation results show that RapidChain’s performance improves with more shards under historical Bitcoin transaction datasets; however, such enhancement is absent in scenarios with uneven transaction distributions, highlighting the need for additional load balancing methods. Through ShardingSim, we can simulate committee-based sharding blockchain, evaluateits performance, and identify potential load-balancing issues between shards. Yuehua Wu, Feihu Yan, Wenzhi Chen |
ICBC | 3 |
| 2024 | CMFF-Face: Attention-Based Cross-Modal Feature Fusion for High-Quality Audio-Driven Talking Face GenerationabstractAudio-driven talking face generation creates lip-synchronized and high-quality face videos from given audio and target face images, which is a challenging task due to the inherent modal gap between audio and face images. To address this issue, we propose an attention-based Cross-Modal Feature Fusion network for talking Face generation, called CMFF-Face. Specifically, we introduce a cross-modal feature fusion generator, which incorporates a fusion process in each convolutional encoder layer, allowing for layer-wise fusing of interactive audio and face features to generate high-quality talking faces. Additionally, a lip synchronization discriminator is designed to improve audio-lip synchronization, which uses a two-branch cross-attention mechanism to capture the associations between synchronized audio and face more effectively. Finally, we employ a CLIP-based audio-lip synchronization loss that helps distinguish between positive and negative sample pairs to enhance the lip synchronization. Comprehensive experiments on the LRS2 and LRW datasets demonstrate that our method outperforms the state-of-the-arts in terms of lip synchronization and visual quality. Guangzhe Zhao, Feihu Yan |
ICMR | 4 |
| 2024 | Expression-aware neural radiance fields for high-fidelity talking portrait synthesis
Xueni Guo, Jiahe Li 0007, Feihu Yan, Guangzhe Zhao, Caiyong Wang |
Image Vis. Comput. | 6 |
| 2024 | PMANet: Progressive multi-stage attention networks for skin disease classification
Guangzhe Zhao, Benwang Lin, Feihu Yan |
Image Vis. Comput. | 5 |
| 2022 | Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image RecognitionabstractPrevious works on multi-label image recognition (MLIR) usually use CNNs as a starting point for research. In this paper, we take pure Vision Transformer (ViT) as the research base and make full use of the advantages of Transformer with long-range dependency modeling to circumvent the disadvantages of CNNs limited to local receptive field. However, for multi-label images containing multiple objects from different categories, scales, and spatial relations, it is not optimal to use global information alone. Our goal is to leverage ViT's patch tokens and self-attention mechanism to mine rich instances in multi-label images, named diverse instance discovery (DiD). To this end, we propose a semantic category-aware module and a spatial relationship-aware module, respectively, and then combine the two by a re-constraint strategy to obtain instance-aware attention maps. Finally, we propose a weakly supervised object localization-based approach to extract multi-scale local features, to form a multi-view pipeline. Our method requires only weakly supervised information at the label level, no additional knowledge injection or other strongly supervised information is required. Experiments on three benchmark datasets show that our method significantly outperforms previous works and achieves state-of-the-art results under fair experimental comparisons. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Feihu Yan, Yuan He 0011, Hui Xue 0001 |
ICME | 6 |
| 2022 | Robust and efficient edge-based visual odometryabstractVisual odometry, which aims to estimate relative camera motion between sequential video frames, has been widely used in the fields of augmented reality, virtual reality, and autonomous driving. However, it is still quite challenging for state-of-the-art approaches to handle low-texture scenes. In this paper, we propose a robust and efficient visual odometry algorithm that directly utilizes edge pixels to track camera pose. In contrast to direct methods, we choose reprojection error to construct the optimization energy, which can effectively cope with illumination changes. The distance transform map built upon edge detection for each frame is used to improve tracking efficiency. A novel weighted edge alignment method together with sliding window optimization is proposed to further improve the accuracy. Experiments on public datasets show that the method is comparable to state-of-the-art methods in terms of tracking accuracy, while being faster and more robust. Feihu Yan, Zhaoxin Li, Zhong Zhou |
Comput. Vis. Media | 1 |
| 2022 | Subspace clustering by directly solving Discriminative K-means
Chenhui Gao, Wenzhi Chen, Feiping Nie 0001, Weizhong Yu, Feihu Yan |
Knowl. Based Syst. | 5 |
| 2022 | Attention-Guided Collaborative CountingabstractExisting crowd counting designs usually exploit multi-branch structures to address the scale diversity problem. However, branches in these structures work in a competitive rather than collaborative way. In this paper, we focus on promoting collaboration between branches. Specifically, we propose an attention-guided collaborative counting module (AGCCM) comprising an attention-guided module (AGM) and a collaborative counting module (CCM). The CCM promotes collaboration among branches by recombining each branch's output into an independent count and joint counts with other branches. The AGM capturing the global attention map through a transformer structure with a pair of foreground-background related loss functions can distinguish the advantages of different branches. The loss functions do not require additional labels and crowd division. In addition, we design two kinds of bidirectional transformers (Bi-Transformers) to decouple the global attention to row attention and column attention. The proposed Bi-Transformers are able to reduce the computational complexity and handle images in any resolution without cropping the image into small patches. Extensive experiments on several public datasets demonstrate that the proposed algorithm performs favorably against the state-of-the-art crowd counting methods. Hong Mo, Wenqi Ren, Feihu Yan, Zhong Zhou, Xiaochun Cao, Wei Wu 0008 |
IEEE Trans. Image Process. | 4 |
| 2022 | 3D scene graph prediction from point cloudsabstractIn this study, we propose a novel 3D scene graph prediction approach for scene understanding from point clouds. It can automatically organize the entities of a scene in a graph, where objects are nodes and their relationships are modeled as edges. More specifically, we employ the DGCNN to capture the features of objects and their relationships in the scene. A Graph Attention Network (GAT) is introduced to exploit latent features obtained from the initial estimation to further refine the object arrangement in the graph structure. A one loss function modified from cross entropy with a variable weight is proposed to solve the multi-category problem in the prediction of object and predicate. Experiments reveal that the proposed approach performs favorably against the state-of-the-art methods in terms of predicate classification and relationship prediction and achieves comparable performance on object classification prediction. The 3D scene graph prediction approach can form an abstract description of the scene space from point clouds. Fanfan Wu, Feihu Yan, Zhong Zhou |
Virtual Real. Intell. Hardw. | 2 |
| 2021 | Monocular Dense SLAM with Consistent Deep Depth Prediction
Feihu Yan, Jiawei Wen, Zhaoxin Li, Zhong Zhou |
CGI | 1 |
| 2019 | Handling pure camera rotation in semi-dense monocular SLAM
Feihu Yan, Zhong Zhou |
Vis. Comput. | 2 |