EDBT 2026 Demo / reviewers in the wild / expert
Yu Fang 0008
dblp:88/3790-8
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DCIM-AVSR: Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction ModuleabstractSpeech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recognition (AVSR) model enhances traditional speech recognition, particularly in noisy environments, by incorporating visual modalities like lip movements and facial expressions. While traditional AVSR models trained on large-scale datasets with numerous parameters can achieve remarkable accuracy, often surpassing human performance, they also come with high training costs and deployment challenges. To address these issues, we introduce an efficient AVSR model that reduces the number of parameters through the integration of a Dual Conformer Interaction Module (DCIM). In addition, we propose a pre-training method that optimizes model performance by fine-tuning. Unlike conventional models that require the system to independently learn the hierarchical relationship between audio and visual modalities, our approach incorporates this distinction directly into the model architecture. This design enhances both efficiency and performance, resulting in a more practical and effective solution for AVSR tasks. Haolin Huang, Yu Fang 0008, Mengjie Xu, Qian Wang 0001 |
ICASSP | 4 |
| 2025 | Visual-informed Silent Video Identity ConversionabstractConventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC. Demo page is available at https://pussycat0700.github.io/MuteSwap-Demo/. Yu Fang 0008, Zhouhan Lin |
ACM Multimedia | 2 |
| 2025 | Clinical knowledge-guided hybrid classification network for automatic periodontal disease diagnosis in X-ray image
Lanzhuju Mei, Zhiming Cui 0001, Yu Fang 0008, Yuan Liu 0025, Hongchang Lai, Maurizio Tonetti, Dinggang Shen |
Medical Image Anal. | 4 |
| 2025 | Learning contrast and content representations for synthesizing magnetic resonance image of arbitrary contrast
Honglin Xiong, Zhenrong Shen 0001, Kaicong Sun, Yu Fang 0008, Dinggang Shen, Qian Wang 0001 |
Medical Image Anal. | 5 |
| 2025 | Geometry-Aware Attenuation Learning for Sparse-View CBCT ReconstructionabstractCone Beam Computed Tomography (CBCT) plays a vital role in clinical imaging. Traditional methods typically require hundreds of 2D X-ray projections to reconstruct a high-quality 3D CBCT image, leading to considerable radiation exposure. This has led to a growing interest in sparse-view CBCT reconstruction to reduce radiation doses. While recent advances, including deep learning and neural rendering algorithms, have made strides in this area, these methods either produce unsatisfactory results or suffer from time inefficiency of individual optimization. In this paper, we introduce a novel geometry-aware encoder-decoder framework to solve this problem. Our framework starts by encoding multi-view 2D features from various 2D X-ray projections with a 2D CNN encoder. Leveraging the geometry of CBCT scanning, it then back-projects the multi-view 2D features into the 3D space to formulate a comprehensive volumetric feature map, followed by a 3D CNN decoder to recover 3D CBCT image. Importantly, our approach respects the geometric relationship between 3D CBCT image and its 2D X-ray projections during feature back projection stage, and enjoys the prior knowledge learned from the data population. This ensures its adaptability in dealing with extremely sparse view inputs without individual training, such as scenarios with only 5 or 10 X-ray projections. Extensive evaluations on two simulated datasets and one real-world dataset demonstrate exceptional reconstruction quality and time efficiency of our method. Yu Fang 0008, Changjian Li 0001, Han Wu 0007, Yuan Liu 0025, Dinggang Shen, Zhiming Cui 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Contrast Representation Learning from Imaging Parameters for Magnetic Resonance Image Synthesis
Honglin Xiong, Yu Fang 0008, Kaicong Sun, Xiaopeng Zong, Qian Wang 0001 |
MICCAI (7) | 2 |
| 2024 | DTR-Net: Dual-Space 3D Tooth Model Reconstruction From Panoramic X-Ray ImagesabstractIn digital dentistry, cone-beam computed tomography (CBCT) can provide complete 3D tooth models, yet suffers from a long concern of requiring excessive radiation dose and higher expense. Therefore, 3D tooth model reconstruction from 2D panoramic X-ray image is more cost-effective, and has attracted great interest in clinical applications. In this paper, we propose a novel dual-space framework, namely DTR-Net, to reconstruct 3D tooth model from 2D panoramic X-ray images in both image and geometric spaces. Specifically, in the image space, we apply a 2D-to-3D generative model to recover intensities of CBCT image, guided by a task-oriented tooth segmentation network in a collaborative training manner. Meanwhile, in the geometric space, we benefit from an implicit function network in the continuous space, learning using points to capture complicated tooth shapes with geometric properties. Experimental results demonstrate that our proposed DTR-Net achieves state-of-the-art performance both quantitatively and qualitatively in 3D tooth model reconstruction, indicating its potential application in dental practice. Lanzhuju Mei, Yu Fang 0008, Yue Zhao 0012, Xiang Sean Zhou, Zhiming Cui 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2023 | HC-Net: Hybrid Classification Network for Automatic Periodontal Disease Diagnosis
Lanzhuju Mei, Yu Fang 0008, Zhiming Cui 0001, Nizhuan Wang 0001, Xuming He 0001, Yiqiang Zhan, Xiang Sean Zhou, Maurizio Tonetti, Dinggang Shen |
MICCAI (6) | 2 |
| 2023 | Multi-view Vertebra Localization and Identification from CT Images
Han Wu 0007, Yu Fang 0008, Nizhuan Wang 0001, Zhiming Cui 0001, Dinggang Shen |
MICCAI (5) | 3 |
| 2022 | Curvature-Enhanced Implicit Function Network for High-quality Tooth Model Generation from CBCT Images
Yu Fang 0008, Zhiming Cui 0001, Lei Ma 0006, Lanzhuju Mei, Yue Zhao 0012, Zhihao Jiang 0001, Yiqiang Zhan, Yongsheng Pan, Dinggang Shen |
MICCAI (5) | 1 |