Yihang Zhou

dblp:117/0292 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0001-6354-1259ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Rapid spatio-temporal MR fingerprinting using physics-informed implicit neural representation
Chaoguang Gong, Lixian Zou, Peng Li 0063, Xingyang Wu, Yangzi Qiao, Zhanqi Hu, Yihang Zhou, Kai Wang 0099, Yue Hu 0003, Haifeng Wang 0003
Medical Image Anal.8
2025 HAVIR: Hierarchical Vision to Image Reconstruction Using Clip-Guided Versatile Diffusion
Dong Liang 0001, Hairong Zheng, Yihang Zhou
BIBM4
2025 Extreme Value Policy Optimization for Safe Reinforcement Learning
abstract
Ensuring safety is a critical challenge in applying Reinforcement Learning (RL) to real-world scenarios. Constrained Reinforcement Learning (CRL) addresses this by maximizing returns under predefined constraints, typically formulated as the expected cumulative cost. However, expectation-based constraints overlook rare but high-impact extreme value events in the tail distribution, such as black swan incidents, which can lead to severe constraint violations. To address this issue, we propose the Extreme Value policy Optimization (EVO) algorithm, leveraging Extreme Value Theory (EVT) to model and exploit extreme reward and cost samples, reducing constraint violations. EVO introduces an extreme quantile optimization objective to explicitly capture extreme samples in the cost tail distribution. Additionally, we propose an extreme prioritization mechanism during replay, amplifying the learning signal from rare but high-impact extreme samples. Theoretically, we establish upper bounds on expected constraint violations during policy updates, guaranteeing strict constraint satisfaction at a zero-violation quantile level. Further, we demonstrate that EVO achieves a lower probability of constraint violations than expectation-based methods and exhibits lower variance than quantile regression methods. Extensive experiments show that EVO significantly reduces constraint violations during training while maintaining competitive policy performance compared to baselines.
Shiqing Gao, Yihang Zhou, Haoyu Luo, Yiheng Bing, Jiaxin Ding 0001, Luoyi Fu, Xinbing Wang
ICML2
2025 COMAE: COMprehensive Attribute Exploration for Zero-shot Hashing
abstract
Zero-shot hashing (ZSH) has shown excellent success owing to its efficiency and generalization in large-scale retrieval scenarios. However, existing works ignore the locality relationships of representations and attributes, which have effective transferability between seeable classes and unseeable classes. Also, the continuous value attributes are not fully harnessed. In response, we conduct a COMprehensive Attribute Exploration for ZSH, named COMAE, which depicts the relationships from seen classes to unseen ones through three meticulously designed explorations, i.e., point-wise, pair-wise and class-wise consistency constraints. By regressing attributes from the proposed attribute prototype network, COMAE learns the local features that are relevant to the visual attributes. Then COMAE utilizes contrastive learning to comprehensively depict the context of attributes, rather than instance-independent optimization. Finally, the class-wise constraint is designed to cohesively learn the hash code, image representation, and visual attributes more effectively. Furthermore, theoretical analysis is provided to show the effectiveness of COMAE. Experimental results demonstrate that COMAE outperforms state-of-the-art hashing models, especially in scenarios with a larger number of unseen label classes.
Qingqing Long, Yihang Zhou, Ran Zhang 0008, Zhiyuan Ning 0001, Zhihong Zhu 0001, Yuanchun Zhou, Xuezhi Wang 0004, Meng Xiao 0001
ICMR3
2025 NeuroSwift: A Lightweight Cross-Subject Framework for fMRI Visual Reconstruction of Complex Scenes
abstract
Reconstructing visual information from brain activity via computer vision technology provides an intuitive understanding of visual neural mechanisms. Despite progress in decoding fMRI data with generative models, achieving accurate cross‑subject reconstruction of visual stimuli remains challenging and computationally demanding. This difficulty arises from inter‑subject variability in neural representations and the brain’s abstract encoding of core semantic features in complex visual inputs.
Dong Liang 0001, Yihang Zhou
MMAsia3
2025 EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation
abstract
Designing enzyme backbones with substrate-specific functionality is a critical challenge in computational protein engineering. Current generative models excel in protein design but face limitations in binding data, substrate-specific control, and flexibility for de novo enzyme backbone generation. To address this, we introduce **EnzyBind**, a dataset with 11,100 experimentally validated enzyme-substrate pairs specifically curated from PDBbind. Building on this, we propose **EnzyControl**, a method that enables functional and substrate-specific control in enzyme backbone generation. Our approach generates enzyme backbones conditioned on MSA-annotated catalytic sites and their corresponding substrates, which are automatically extracted from curated enzyme-substrate data. At the core of EnzyControl is **EnzyAdapter**, a lightweight, modular component integrated into a pretrained motif-scaffolding model, allowing it to become substrate-aware. A two-stage training paradigm further refines the model's ability to generate accurate and functional enzyme structures. Experiments show that our EnzyControl achieves the best performance across structural and functional metrics on EnzyBind and EnzyBench benchmarks, with particularly notable improvements of 13% in designability and 13% in catalytic efficiency compared to the baseline models. The code is released at https://github.com/Vecteur-libre/EnzyControl.
Jianyu Shi, Hui Yu 0011, Yihang Zhou
NeurIPS8
2025 ProteinConformers: Benchmark Dataset for Simulating Protein Conformational Landscape Diversity and Plausibility
abstract
Understanding the conformational landscape of proteins is essential for elucidating protein function and facilitating drug design. However, existing protein conformation benchmarks fail to capture the full energy landscape, limiting their ability to evaluate the diversity and physical plausibility of AI-generated structures. We introduce ProteinConformers, a large-scale benchmark dataset comprising over 381,000 physically realistic conformations for 87 CASP targets. These were derived from more than 40,000 structural decoys via extensive all-atom molecular dynamics simulations totaling over 6 million CPU hours. Using this dataset, we propose novel metrics to evaluate conformational diversity and plausibility, and systematically benchmark six protein conformation generative models. Our results highlight that leveraging large-scale protein sequence data can enhance a model’s ability to explore conformational space, potentially reducing reliance on MD-derived data. Additionally, we find that PDB and MD datasets influence model performance differently, current models perform well on inter-atomic distance prediction but struggle with inter-residue orientation generation. Overall, our dataset, evaluation metrics, and benchmarking results provide the first comprehensive foundation for assessing generative models in protein conformational modeling. Dataset and instructions are available at https://huggingface.co/ datasets/Jim990908/ProteinConformers/tree/main. Codes are stored at https://github.com/auroua/ProteinConformers. An interactive website locates at https://zhanggroup.org/ProteinConformers.
Yihang Zhou, Minghao Sun, Jin Song
NeurIPS1
2025 Spatio-temporal masked autoencoder-based phonetic segments classification from ultrasound
Xi Dan, Kele Xu, Yihang Zhou, Chuanguang Yang, Yutao Dou, Cheng Yang 0004
Speech Commun.3
2025 PEARL: Cascaded Self-Supervised Cross-Fusion Learning for Parallel MRI Acceleration
abstract
Supervised deep learning (SDL) methodology holds promise for accelerated magnetic resonance imaging (AMRI) but is hampered by the reliance on extensive training data. Some self-supervised frameworks, such as deep image prior (DIP), have emerged, eliminating the explicit training procedure but often struggling to remove noise and artifacts under significant degradation. This work introduces a novel self-supervised accelerated parallel MRI approach called PEARL, leveraging a multiple-stream joint deep decoder with two cross-fusion schemes to accurately reconstruct one or more target images from compressively sampled k-space. Each stream comprises cascaded cross-fusion sub-block networks (SBNs) that sequentially perform combined upsampling, 2D convolution, joint attention, ReLU activation and batch normalization (BN). Among them, combined upsampling and joint attention facilitate mutual learning between multiple-stream networks by integrating multi-parameter priors in both additive and multiplicative manners. Long-range unified skip connections within SBNs ensure effective information propagation between distant cross-fusion layers. Additionally, incorporating dual-normalized edge-orientation similarity regularization into the training loss enhances detail reconstruction and prevents overfitting. Experimental results consistently demonstrate that PEARL outperforms the existing state-of-the-art (SOTA) self-supervised AMRI technologies in various MRI cases. Notably, 5-fold$\sim$6-fold accelerated acquisition yields a 1$\%$ $\sim$ 2$\%$ improvement in SSIM$_{\mathsf{ROI}}$ and a 3$\%$ $\sim$ 6$\%$ improvement in PSNR$_{\mathsf{ROI}}$, along with a significant 15$\%$ $\sim$ 20$\%$ reduction in RLNE$_{\mathsf{ROI}}$.
Qingyong Zhu, Zhuo-Xu Cui, Chentao Cao, Xiaomeng Yan, Yihang Zhou, Yanjie Zhu, Haifeng Wang 0003, Hongwu Zeng, Dong Liang 0001
IEEE J. Biomed. Health Informatics8
2025 Score-Based Diffusion Models With Self-Supervised Learning for Accelerated 3D Multi-Contrast Cardiac MR Imaging
abstract
Long scan time significantly hinders the widespread applications of three-dimensional multi-contrast cardiac magnetic resonance (3D-MC-CMR) imaging. This study aims to accelerate 3D-MC-CMR acquisition by a novel method based on score-based diffusion models with self-supervised learning. Specifically, we first establish a mapping between the undersampled k-space measurements and the MR images, utilizing a self-supervised Bayesian reconstruction network. Secondly, we develop a joint score-based diffusion model on 3D-MC-CMR images to capture their inherent distribution. The 3D-MC-CMR images are finally reconstructed using the conditioned Langenvin Markov chain Monte Carlo sampling. This approach enables accurate reconstruction without fully sampled training data. Its performance was tested on the dataset acquired by a 3D joint myocardial $ \text {T}_{{1}}$ and $ \text {T}_{{1}\rho }$ mapping sequence. The $ \text {T}_{{1}}$ and $ \text {T}_{{1}\rho }$ maps were estimated via a dictionary matching method from the reconstructed images. Experimental results show that the proposed method outperforms traditional compressed sensing and existing self-supervised deep learning MRI reconstruction methods. It also achieves high quality $ \text {T}_{{1}}$ and $ \text {T}_{{1}\rho }$ parametric maps close to the reference maps, even at a high acceleration rate of 14.
Zhuo-Xu Cui, Shucong Qin, Hairong Zheng, Haifeng Wang 0003, Yihang Zhou, Dong Liang 0001, Yanjie Zhu
IEEE Trans. Medical Imaging7
2024 Spatial-temporal Consistency Constraint for Depth and Ego-motion Estimation of Laparoscopic Images
abstract
Estimating depth and ego-motion are crucial tasks for laparoscopic navigation and robotic-assisted surgery. Most current self-supervised methods involve warping one frame onto an adjacent frame using the estimated depth and camera pose. The photometric loss between the estimated and original frames then serves as the training signal. However, these methods encounter major challenges due to non-Lambertian reflection regions and the textureless surfaces of organs, leading to significant performance degradation and scale ambiguity in monocular depth estimation. In this paper, we introduce a network that predicts depth and ego-motion using spatial-temporal consistency constraints. Spatial consistency is derived from the left and right views of the stereo laparoscopic image pairs, while temporal consistency comes from consecutive frames. To enhance the understanding of semantic information in surgical scenes, we employ the Swin Transformer as the encoder and decoder for depth estimation, due to its superior semantic segmentation capabilities. To address issues of illumination variance and scale ambiguity, we incorporate a SIFT loss term to eliminate oversaturated regions in laparoscopic images. Our method is evaluated on the SCARED dataset and shows remarkable results. The code is publicly available at https://github.com/nanasylum/spatialtemporal.
Xiangling Nan, Yingfang Fan, Yihang Zhou, Yaoqun Liu, Fucang Jia, Huoling Luo
BIBM4
2024 PIXEL: Prompt-based Zero-shot Hashing via Visual and Textual Semantic Alignment
abstract
Zero-Shot Hashing (ZSH) has aroused significant attention due to its efficiency and generalizability in multi-modal retrieval scenarios, which aims to encode semantic information into hash codes without needing unseen labeled training samples. In addition to commonly used visual images as visual semantics and class labels as global semantics, the corresponding attribute descriptions contain critical local semantics with detailed information. However, most existing methods focus on leveraging the extracted attribute numerical values, without exploring the textual semantics in attribute descriptions. To bridge this gap, in this paper, we propose Prompt-based zero-shot hashing via vIsual and teXtual sEmantic aLignment, namely PIXEL. Concretely, we design the attribute prompt template depending on attribute descriptions to make the model capture the corresponding local semantics. Then, achieving the textual embedding and visual embedding, we proposed an alignment module to model the intra- and inter-class contrastive distances. In addition, the attribute-wise constraint and class-wise constraint are utilized to collaboratively learn the hash code, image representation, and visual attributes more effectively. Finally, extensive experimental results demonstrate the superiority of PIXEL.
Zeyu Dong, Qingqing Long, Yihang Zhou, Pengfei Wang 0008, Zhihong Zhu 0001, Xiao Luo 0001, Yidong Wang 0003, Pengyang Wang, Yuanchun Zhou
CIKM3
2024 Online Relational Knowledge Distillation for Image Classification
abstract
Existing online Knowledge Distillation (KD) often perform probability-based predictions from independent data samples for knowledge transfer. However, these online KD methods neglect valuable relational information across multiple networks. To address this problem, we propose Online Relational Knowledge Distillation (ORKD). ORKD includes a discriminative loss to construct meaningful feature space and a relational distillation loss to guide structured knowledge transfer among multiple networks. Beyond feature-level distillation, we further construct an ensemble teacher by aggregating probability predictions from multiple networks. The virtual teacher is used to supervise a specific network to enhance its accuracy and avoid the cohort homogenization problem. Experimental results on CIFAR-100 and ImageNet classification demonstrate that ORKD achieves the best performance among state-of-the-art online KD methods over various network architectures. The qualitative visualization shows that ORKD can help the network to learn a more discriminative feature space, resulting in better classification performance.
Yihang Zhou, Chuanguang Yang, Libo Huang 0001, Zhulin An, Yongjun Xu 0001
CSCWD1
2024 MobileViT-FocR: MobileViT with Fixed-One-Centre Loss and Gradient Reversal for Generalised Fake Face Detection
Yihang Zhou, Yizhi Luo
MMM (3)2
2024 Physics-Informed DeepMRI: k-Space Interpolation Meets Heat Diffusion
abstract
Recently, diffusion models have shown considerable promise for MRI reconstruction. However, extensive experimentation has revealed that these models are prone to generating artifacts due to the inherent randomness involved in generating images from pure noise. To achieve more controlled image reconstruction, we reexamine the concept of interpolatable physical priors in k-space data, focusing specifically on the interpolation of high-frequency (HF) k-space data from low-frequency (LF) k-space data. Broadly, this insight drives a shift in the generation paradigm from random noise to a more deterministic approach grounded in the existing LF k-space data. Building on this, we first establish a relationship between the interpolation of HF k-space data from LF k-space data and the reverse heat diffusion process, providing a fundamental framework for designing diffusion models that generate missing HF data. To further improve reconstruction accuracy, we integrate a traditional physics-informed k-space interpolation model into our diffusion framework as a data fidelity term. Experimental validation using publicly available datasets demonstrates that our approach significantly surpasses traditional k-space interpolation methods, deep learning-based k-space interpolation techniques, and conventional diffusion models, particularly in HF regions. Finally, we assess the generalization performance of our model across various out-of-distribution datasets. Our code are available at https://github.com/ZhuoxuCui/Heat-Diffusion.
Zhuo-Xu Cui, Xiaohong Fan, Chentao Cao, Qingyong Zhu, Sen Jia 0005, Haifeng Wang 0003, Yanjie Zhu, Yihang Zhou, Jianping Zhang 0004, Qiegen Liu, Dong Liang 0001
IEEE Trans. Medical Imaging11
2023 Production Evaluation of Citrus Fruits based on the YOLOv5 compressed by Knowledge Distillation
abstract
Pre-harvest estimation of fruit production is crucial for fruit storage and price analysis in the planting of fruit trees. However, prior research consistently displays low accuracy because of problems with small objects, leaf occlusion, and fruit overlap, and they emphasize large networks that are unrealistic in the real world. In this study, we emphasize the use of smartphones to evaluate citrus fruit production. We suggest a simple method for detecting objections based on the YOLOv5 algorithm compressed by knowledge distillation. To extract the visual features, we first use mobilenetV2 as the foundation of YOLOv5. To learn the reliable detection features, we also incorporate an attention mechanism into YOLOv5. As such, we can obtain embedded Yolo served as a student model, which is learned via knowledge distillation. As such we can take the lightweight student model as the final detection model. Finally, we take the embedded Yolo to detect the citrus fruits and take a linear regression model to predict the number of counted fruits and the production is estimated. Experiments show that the proposed method can accurately count fruits and approximate the production.
Zirui Gong, Yihang Zhou, Yuting He 0007, Renjie Huang
CSCWD3
2023 Stay or Leave? Exploring Student Factors Associated with Dropout Patterns in Massive Open Online Courses
abstract
High dropout rates pose a problem for the improvement of massive open online courses (MOOCs). Previous research has suggested that student demographic information, past performance, and participation behaviors can predict dropout. In this study, we analyzed data from 32,593 students (31.2% dropout rate) enrolled in fully online courses at an open university. Using survival analysis, we examined the effect of demographic information, past performance, and participation behaviors on student dropout patterns in the prophase, metaphase, and anaphase of online courses. Our results showed that the median retention time for online learners was 241 days. The Cox proportional hazard model results indicated that student past performance and participation behaviors were significant predictors of dropout. However, certain persistent participation behaviors (i.e. active days and assignment submission) tended to become less important in predicting dropout as the course progressed. Based on these findings, we provide recommendations for identifying at-risk students in a fully online environment.
Yihang Zhou
ICALT2
2023 Dynamic Dual-Graph Fusion Convolutional Network for Alzheimer's Disease Diagnosis
abstract
In this paper, a dynamic dual-graph fusion convolutional network is proposed to improve Alzheimer’s disease (AD) diagnosis performance. The following are the paper’s main contributions: (a) propose a novel dynamic Graph Convolutional Network (GCN) architecture, which is an end-to-end pipeline for diagnosis of the AD task; (b) the proposed architecture can dynamically adjust the graph structure for GCN to produce better diagnosis outcomes by learning the optimal underlying latent graph; (c) incorporate feature graph learning and dynamic graph learning, giving those useful features of subjects more weight while decreasing the weights of other noise features. Experiments indicate that our model provides flexibility and stability while achieving excellent classification results in AD diagnosis.
Fanshi Li, Yanjie Zhu, Yihang Zhou, Dong Liang 0001, Haifeng Wang 0003
ICIP6
2022 PCE-RPM-NET: RPM-NET Based Video Sar Inter-Frame Registration Network
abstract
Due to the existence of system error or the difference in reference coordinates for imaging, there is a significantly spatial mismatch among video SAR frames, which will affect the performance of target localization and tracking greatly. To solve this problem, a PCE-RPM-Net that can achieve video SAR inter-frame images registration is proposed in this paper. PCE-RPM-Net is mainly composed of U-Net and RPM-Net, in which U-Net is adopted to extract features of SAR images to generate point clouds and RPM-Net is used to align the point clouds. Experimental results show that the PCE-RPM-Net has higher registration accuracy and faster speed compared with the classical traditional algorithm.
Zhikun Xie, Yuanyuan Zhou 0007, Yihang Zhou, Jun Shi 0002
IGARSS3