Yan Yang 0011

dblp:37/1091-11 · DBLP profile ↗
← Back
32ranked-venue papers
14as first author
31since 2021 · last 2026
0000-0002-6246-1748ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SZCo: Self-supervised zero-shot co-segmentation with region-text alignment learning
Xin Duan, Yan Yang 0011, Liyuan Pan, Xiabi Liu, Mingyang Gong
Pattern Recognit.2
2026 Toward Trustworthy Multi-View Representation With Fine-Grained Explainability Embeddings
abstract
Multiomics co-learning is a powerful analytical paradigm that has benefited biomedical studies substantially. However, due to the diverse information and complex relationships of multiomics data, naive multi-view learning methods usually run into spurious correlations and biased signatures irrelevant to the diseases of interest. Therefore, the learned representations and cross-omics associations cannot translate into clinical knowledge for disease prediction. This issue becomes particularly severe when clinical data are limited and scarce. To handle this issue, we propose a novel and powerful scheme, referred to as the Causality-driven Trustworthy Multi-View maPping approach (Cad-TMVP). Specifically, we design a fined multi-directional mapping module to extract co-expression patterns across different modalities and capture fine-grained interpretability factors. We also meticulously design dynamic mechanisms to facilitate adaptive loss-term reweighting and trustworthy integration of multiple modalities. Cad-TMVP enhances downstream tasks by developing a cooperative learning module that simultaneously performs automated diagnosis and result interpretation. Furthermore, we develop an efficient search strategy and support computation to reduce the high computational burden, making our approach practicable. We conduct extensive experiments on different types of multiomics data. The proposed method establishes new state-of-the-art results in various settings while maintaining excellent interpretability. Thus, it sets a potentially newparadigm in trustworthy multi-modal learning and verifies its flexibility and versatility in real biomedical applications.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
IEEE Trans. Medical Imaging2
2025 Predicting Brain Age Based on Neuroimaging and DNA Methylation via A Multiomics Attention-Based Variational Autoencoder
abstract
Brain age (BA) is recognized as a significant biomarker for health, closely associated with brain aging, and has been proposed to correlate with the progression of neurodegenerative diseases. Previous studies on brain age prediction predominantly relied on single omics data, limiting the integration of cooperative information from multiple omics data. Additionally, existing brain age prediction methods are susceptible to intermediate features since not all these features are related to brain age. To address these limitations, the study introduces a multiomics attention-based VAE method to predict brain age by integrating neuroimaging data and DNA methylation (DNAm) data, aiming to identify biologically meaningful features truly associated with brain aging. Experimental results show that our method predicts brain age with the MAE of 2.08 years, outperforming the state-of-the-art methods under the same experimental conditions. Furthermore, we estimated the brain age gap in patients with Mild Cognitive Impairment (MCI) and Alzheimer's Disease (AD). The results of both MCI and AD patients exhibited a larger brain age gap compared to the Cognitively Normal (CN) group, indicating the model's discriminative capacity across different diagnostic groups. These findings can assist in the early diagnosis of AD and the formulation of early treatment strategies.
Wenrui Cui, Hong Pang, Yan Yang 0011, Muheng Shang, Hongdong Li, Lei Du 0001
BIBM3
2025 A Data-Driven Brain Imaging Quantitative Trait Mediation Effect Identification Method for Alzheimer's Disease
abstract
Alzheimer's disease (AD) is a severe degenerative disease and finding its causal factors of high-risk are very important. Mediation analysis has been a powerful tool to elucidate the underlying mechanisms of trait of interest using genetic variations as instrumental variable (IVs). However, most current mediation analysis method can only work on a limited number of suspected traits and genetic variations, which demands extensive prior knowledge. In this study, we proposed a datadriven learning method to identify potential mediation effects of brain imaging quantitative traits (QTs). Our method couples the three single models of mediation model and treats it as a multiobjective learning problem. The proposed method can work on brain-wide and genome-wide brain imaging and genetic data without providing candidate imaging QTs and genetic IVs. We applied the proposed method to PET imaging QTs of whole brain and genetic IV of whole genome. The results showed that our method successfully identified multiple PET imaging mediators linking genetic variations to AD. Thus, our method can serve as a powerful screening tool for large-scale mediation analysis.
Yan Yang 0011, Muheng Shang, Hongdong Li, Lei Du 0001
BIBM1
2025 Predicting MCI Conversion Status Using Baseline Neuroimaging Scans and Genetics Variations
abstract
Mild cognitive impairment (MCI) is a prodromal stage of Alzheimer's disease (AD), but not all MCI subjects develop into AD finally. Therefore, distinguishing progressive MCI (pMCI) subjects from stable MCI (sMCI) subjects is an area of intense interest, which may provide targeted treatments for at-risk individuals. On this account, building an MCI conversion prediction model at the early stage is particularly important. The neuroimaging data, especially multi-modal ones, has proven to be a great alternative in predicting MCIs' conversion. In addition, genetic variations such as Single Nucleotide Polymorphism (SNP) can also imply the conversion risk of an individual. The neuroimaging data represents the current status, while SNPs convey the inherited risk of an individual. In this paper, we propose a deep representative fusion method that combines multi-modal baseline neuroimaging data and genetic variations. It can predict the progressive status of MCIs over the following two years, three years and four years, respectively. Experimental results from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database demonstrate that the proposed method has better prediction capability than comparison methods. Moreover, findings show that the stability of the default mode network (DMN) and ventral attention network (VAN) are correlated with the MCI conversion and the learned imaging representations are related to minimental state examination (MMSE) scores which are associated with AD progression.
Yan Yang 0011, Muheng Shang, Jin Zhang 0023, Hongdong Li, Lei Du 0001
BIBM1
2025 EZSR: Event-based Zero-Shot Recognition
abstract
This paper studies zero-shot object recognition using event camera data. Guided by CLIP, which is pre-trained on RGB images, existing approaches achieve zero-shot object recognition by optimizing embedding similarities between event data and RGB images respectively encoded by an event encoder and the CLIP image encoder. Alternatively, several methods learn RGB frame reconstructions from event data for the CLIP image encoder. However, they often result in suboptimal zero-shot performance.This study develops an event encoder without relying on additional reconstruction networks. We theoretically analyze the performance bottlenecks of previous approaches: the embedding optimization objectives are prone to suffer from the spatial sparsity of event data, causing semantic misalignments between the learned event embedding space and the CLIP text embedding space. To mitigate the issue, we explore a scalar-wise modulation strategy. Furthermore, to scale up the number of events and RGB data pairs for training, we also study a pipeline for synthesizing event data from static RGB images in mass.Experimentally, we demonstrate an attractive scaling property in the number of parameters and synthesized data. We achieve superior zero-shot object recognition performance on extensive standard benchmark datasets, even compared with past supervised learning approaches. For example, our model with a ViT/B-16 backbone achieves 47.84% zero-shot accuracy on the N-ImageNet dataset.
Yan Yang 0011, Liyuan Pan, Dongxu Li 0003, Liu Liu 0009
CVPR1
2025 Storyboard-guided Alignment for Fine-grained Video Action Recognition
abstract
Fine-grained video action recognition can be formulated as a video–text matching problem. Previous approaches primarily rely on global video semantics to consolidate video embeddings, often leading to misaligned video–text pairs due to inaccurate atomic-level action understanding. This inaccuracy arises due to i) videos with distinct global semantics may share similar atomic actions or visual appearances, and ii) atomic actions can be momentary, gradual, or not directly aligned with overarching video semantics. Inspired by storyboarding, where a script is segmented into individual shots, we propose a multi-granularity framework, SFAR. SFAR generates fine-grained descriptions of common atomic actions for each global semantic using a large language model. Unlike existing works that refine global semantics with auxiliary video frames, SFAR introduces a filtering metric to ensure correspondence between the descriptions and the global semantics, eliminating the need for direct video involvement and thereby enabling more nuanced recognition of subtle actions. By leveraging both global semantics and fine-grained descriptions, our SFAR effectively identifies prominent frames within videos, thereby improving the accuracy of embedding aggregation. Extensive experiments on various video action recognition datasets demonstrate the competitive performance of our SFAR in supervised, few-shot, and zero-shot settings.
Enqi Liu, Liyuan Pan, Yan Yang 0011, Yiran Zhong, Zhijing Wu 0001, Xinxiao Wu, Liu Liu 0009
NeurIPS3
2025 Flowering Time Prediction of Wheat From DIA-MS Data
abstract
Traditional methods utilising data-independent acquisition mass spectrometry (DIA-MS) data for predictions depend on database searches against predefined spectral libraries for characterisation and quantification of the proteomes, limiting scalability and adaptability across various applications. However, directly applying existing networks on DIA-MS data represented as images for end-to-end predictions struggles to mine a predictive pattern, due to non-uniform region importance across the image and divergences exhibited among different regions of the image. To overcome these limitations, we propose a new frame-work with two modules: i) a dynamic sampling module that identifies regions of interest from the DIA-MS image, constraining the network to focus on the most informative regions of the image only; ii) a mixture of experts module that sparsely routes the regions of interest to related expert networks, facilitating adaptive computation of region features. The region features are then fused for predictions. Experimentally, to benchmark our method, we collected a large DIA-MS dataset of wheat for flowering time prediction, and our approach significantly outperforms previous end-to-end methods, i. e., 0.171 R2 improvements.
Yan Yang 0011, Utpal Bose, James Broadbent, Sally Stockwell, Keren Byrne, Eric A. Stone, Shannon Dillon
WACV1
2025 Mutual-assistance learning for trustworthy biomarker discovery and disease prediction
abstract
Integrating and analyzing multiple omics datasets, such as genomics, environmental influences, and imaging endophenotypes, has yielded an abundance of candidate biomarkers. However, translating such findings into beneficial clinical knowledge for disease prediction remains challenging. This becomes even more challenging when studying interpretable high-order feature interactions such as gene-environment interaction (G$\times $E) to understand the etiology. To fill this gap, we draw on the idea of mutual-assistance (MA) learning and accordingly propose a fresh and powerful scheme, referred to as mutual-assistance causal biomarker discovery and stable disease prediction approach (MA-CBxDP). Specifically, we design an interpretable bi-directional mapping framework, integrated with a causal feature interaction module, to extract co-expression patterns across different modalities and identify trustworthy biomarkers including G$\times $E. A cooperative prediction module is further incorporated to ensure accurate diagnosis and identification of causal effects for pathogenesis. Importantly, biomarker discovery and disease prediction can mutually reinforce each other, helping to provide novel insights into chronic diseases. Furthermore, in light of the large computational burden incurred by the high-dimensional interactions, we devise a rapid strategy and extend it to a more practical but challenging chromosome-wide setting. We conduct extensive experiments on two databases under three tasks, i.e. multimodal correlation, disease diagnosis, and trait prediction. MA-CBxDP establishes new state-of-the-art results in predicting clinical scores and disease status classification, while maintaining exceptional interpretability, verifying its flexibility and versatility in practical applications.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
Briefings Bioinform.2
2025 Trustworthy causal biomarker discovery: a multiomics brain imaging genetics-based approach
abstract
MOTIVATION: Discovering genetic variations underpinning brain disorders is important to understand their pathogenesis. Indirect associations or spurious causal relationships pose a threat to the reliability of biomarker discovery for brain disorders, potentially misleading or incurring bias in subsequent decision-making. Unfortunately, the stringent selection of reliable biomarker candidates for brain disorders remains a predominantly unexplored challenge. RESULTS: In this article, to fill this gap, we propose a fresh and powerful scheme, referred to as the Causality-aware Genotype intermediate Phenotype Correlation Approach (Ca-GPCA). Specifically, we design a bidirectional association learning framework, integrated with a parallel causal variable decorrelation module and sparse variable regularizer module, to identify trustworthy causal biomarkers. A disease diagnosis module is further incorporated to ensure accurate diagnosis and identification of causal effects for pathogenesis. Additionally, considering the large computational burden incurred by high-dimensional genotype-phenotype covariances, we develop a fast and efficient strategy to reduce the runtime and prompt practical availability and applicability. Extensive experimental results on four simulation data and real neuroimaging genetic data clearly show that Ca-GPCA outperforms state-of-the-art methods with excellent built-in interpretability. This can provide novel and reliable insights into the underlying pathogenic mechanisms of brain disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/ZJ-Techie/Ca-GPCA.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
Bioinform.2
2025 LCCo: Lending CLIP to co-segmentation
Xin Duan, Yan Yang 0011, Liyuan Pan, Xiabi Liu
Pattern Recognit.2
2024 LDP: Language-driven Dual-Pixel Image Defocus Deblurring Network
abstract
Recovering sharp images from dual-pixel (DP) pairs with disparity-dependent blur is a challenging task. Existing blur map-based deblurring methods have demonstrated promising results. In this paper, we propose, to the best of our knowledge, the first framework that introduces the contrastive language-image pre-training framework (CLIP) to accurately estimate the blur map from a DP pair unsu-pervisedly. To achieve this, we first carefully design text prompts to enable CLIP to understand blur-related geo-metric prior knowledge from the DP pair. Then, we pro-pose a format to input a stereo DP pair to CLIP without any fine-tuning, despite the fact that CLIP is pre-trained on monocular images. Given the estimated blur map, we intro-duce a blur-prior attention block, a blur-weighting loss, and a blur-aware loss to recover the all-in-focus image. Our method achieves state-of-the-art performance in extensive experiments (see Fig. 1).
Hao Yang 0040, Liyuan Pan, Yan Yang 0011, Richard I. Hartley, Miaomiao Liu 0001
CVPR3
2024 Language-driven All-in-one Adverse Weather Removal
abstract
All-in-one (AiO) frameworks restore various adverse weather degradations with a single set of networks jointly. To handle various weather conditions, an AiO framework is expected to adaptively learn weather-specific knowledge for different degradations and shared knowledge for common patterns. However, existing methods: 1) rely on extra su-pervision signals, which are usually unknown in real-world applications; 2) employ fixed network structures, which re-strict the diversity of weather-specific knowledge. In this paper, we propose a Language-driven Restoration frame-work (LDR) to alleviate the aforementioned issues. First, we leverage the power of pre-trained vision-language (PVL) models to enrich the diversity of weather-specific knowl-edge by reasoning about the occurrence, type, and severity of degradation, generating description-based degradation priors. Then, with the guidance of degradation prior, we sparsely select restoration experts from a candidate list dy-namically based on a Mixture-of-Experts (MoE) structure. This enables us to adaptively learn the weather-specific and shared knowledge to handle various weather conditions (e.g., unknown or mixed weather). Experiments on exten-sive restoration scenarios show our superior performance.
Hao Yang 0040, Liyuan Pan, Yan Yang 0011, Wei Liang 0008
CVPR3
2024 Event Camera Data Dense Pre-training
Yan Yang 0011, Liyuan Pan, Liu Liu 0009
ECCV (43)1
2024 Event-based Few-shot Fine-grained Human Action Recognition
abstract
Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as high-dynamic or low lighting conditions. Event cameras can independently and sparsely capture brightness changes in a scene at microsecond resolution and high dynamic range, which offer a promising solution. However, the modality differences between events and RGB frames, and the lack of paired fine-grained data hinder the development of event-based FGH action recognition. Therefore, in this paper, we introduce the first Event Camera Fine-grained Human Action (E-FAction) dataset. This dataset comprises 3304 paired ‘event stream and RGB sequence’, covering 15 coarse action classes and 128 fine-grained actions. Then, we develop a versatile event feature extractor. Considering the spatial sparsity of event stream, we design two modules to mine the temporal motion and semantic features under the guidance of paired RGB frames, facilitating robust weight initialization for the feature extractor in few-shot FGH action recognition. We conduct extensive experiments on both published and our built synthetic and real datasets, and consistently achieve state-of-the-art performance compared to existing baselines. Code and dataset will be available at link.
Zonglin Yang 0002, Yan Yang 0011, Yuheng Shi, Hao Yang 0040, Ruikun Zhang, Liu Liu 0009, Xinxiao Wu, Liyuan Pan
IROS2
2024 Spatial Transcriptomics Analysis of Zero-Shot Gene Expression Prediction
Yan Yang 0011, Xuesong Li 0001, Shafin Rahman, Eric A. Stone
MICCAI (4)1
2024 Disease Progression Prediction Incorporating Genotype-Environment Interactions: A Longitudinal Neurodegenerative Disorder Study
Jin Zhang 0023, Muheng Shang, Yan Yang 0011, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
MICCAI (3)3
2024 SpikMamba: When SNN meets Mamba in Event-based Human Action Recognition
Yan Yang 0011, Shizhuo Deng, Da Teng, Liyuan Pan
MMAsia2
2024 LMHaze: Intensity-aware Image Dehazing with a Large-scale Multi-intensity Real Haze Dataset
abstract
Image dehazing has drawn a significant attention in recent years.Learning-based methods usually require paired hazy and corresponding ground truth (haze-free) images for training.However, it is difficult to collect real-world image pairs, which prevents developments of existing methods.Although several works partially alleviate this issue by using synthetic datasets or small-scale real datasets.The haze intensity distribution bias and scene homogeneity in existing datasets limit the generalization ability of these methods, particularly when encountering images with previously unseen haze intensities.In this work, we present LMHaze, a large-scale, high-quality real-world dataset.LMHaze comprises paired hazy and haze-free images captured in diverse indoor and outdoor environments, spanning multiple scenarios and haze intensities.It contains over 5K high-resolution image pairs, surpassing the size of the biggest existing real-world dehazing dataset by over 25 times.Meanwhile, to better handle images with different haze intensities, we propose a mixture-of-experts model based on Mamba (MoE-Mamba) for dehazing, which dynamically adjusts the model parameters according to the haze intensity.Moreover, with our proposed dataset, we conduct a new large multimodal model (LMM)-based benchmark study to simulate human perception for evaluating dehazed images.Experiments demonstrate that LMHaze dataset improves the dehazing performance in real scenarios and our dehazing method provides better results compared to state-of-the-art methods.The dataset and code are available at our project page.
Ruikun Zhang, Hao Yang 0040, Yan Yang 0011, Ying Fu 0003, Liyuan Pan
MMAsia3
2024 Convolutional Masked Image Modeling for Dense Prediction Tasks on Pathology Images
abstract
This paper studies a convolutional masked image modeling approach for boosting downstream dense prediction tasks on pathology images. Our method is self-supervised, and entails two strategies in sequence. Considering features contained in the pathology images usually have a large spatial span, e.g., glands, we insert [MASK] tokens to the masked regions after the stem layer of the convolutional network for encoding unmasked pixels, which facilitates information propagation through masked regions for reconstructing unmasked pixels. Furthermore, the pathology images contain features that are represented in diverse affine shapes and color spaces. We, therefore, enforce the network to learn the affine and color invariant embedding by imposing transformation constraints between the unmasked image-encoded embedding and reconstruction targets. Our approach is simple but effective. With extensive experiments on standard benchmark datasets, we demonstrate superior transfer learning performance on downstream tasks over past state-of-the-art approaches.
Yan Yang 0011, Liyuan Pan, Liu Liu 0009, Eric A. Stone
WACV1
2024 Modeling genotype-protein interaction and correlation for Alzheimer's disease: a multi-omics imaging genetics study
abstract
Integrating and analyzing multiple omics data sets, including genomics, proteomics and radiomics, can significantly advance researchers' comprehensive understanding of Alzheimer's disease (AD). However, current methodologies primarily focus on the main effects of genetic variation and protein, overlooking non-additive effects such as genotype-protein interaction (GPI) and correlation patterns in brain imaging genetics studies. Importantly, these non-additive effects could contribute to intermediate imaging phenotypes, finally leading to disease occurrence. In general, the interaction between genetic variations and proteins, and their correlations are two distinct biological effects, and thus disentangling the two effects for heritable imaging phenotypes is of great interest and need. Unfortunately, this issue has been largely unexploited. In this paper, to fill this gap, we propose $\textbf{M}$ulti-$\textbf{T}$ask $\textbf{G}$enotype-$\textbf{P}$rotein $\textbf{I}$nteraction and $\textbf{C}$orrelation disentangling method ($\textbf{MT-GPIC}$) to identify GPI and extract correlation patterns between them. To ensure stability and interpretability, we use novel and off-the-shelf penalties to identify meaningful genetic risk factors, as well as exploit the interconnectedness of different brain regions. Additionally, since computing GPI poses a high computational burden, we develop a fast optimization strategy for solving MT-GPIC, which is guaranteed to converge. Experimental results on the Alzheimer's Disease Neuroimaging Initiative data set show that MT-GPIC achieves higher correlation coefficients and classification accuracy than state-of-the-art methods. Moreover, our approach could effectively identify interpretable phenotype-related GPI and correlation patterns in high-dimensional omics data sets. These findings not only enhance the diagnostic accuracy but also contribute valuable insights into the underlying pathogenic mechanisms of AD.
Jin Zhang 0023, Zikang Ma, Yan Yang 0011, Lei Guo 0002, Lei Du 0001
Briefings Bioinform.3
2024 Weakly-Supervised Depth Estimation and Image Deblurring via Dual-Pixel Sensors
abstract
Dual-pixel (DP) imaging sensors are getting more popularly adopted by modern cameras. A DP camera captures a pair of images in a single snapshot by splitting each pixel in half. Several previous studies show how to recover depth information by treating the DP pair as an approximate stereo pair. However, dual-pixel disparity occurs only in image regions with defocus blur which is unlike classic stereo disparity. Heavy defocus blur in DP pairs affects the performance of depth estimation approaches based on matching. Therefore, we treat the blur removal and the depth estimation as a joint problem. We investigate the formation of the DP pair, which links the blur and depth information, rather than blindly removing the blur effect. We propose a mathematical DP model that can improve depth estimation by the blur. This exploration motivated us to propose our previous work, an end-to-end DDDNet (DP-based Depth and Deblur Network), which jointly estimates depth and restores the image in a supervised fashion. However, collecting the ground-truth (GT) depth map for the DP pair is challenging and limits the depth estimation potential of the DP sensor. Therefore, we propose an extension of the DDDNet, called WDDNet (Weakly-supervised Depth and Deblur Network), which includes an efficient reblur solver that does not require GT depth maps for training. To achieve this, we convert all-in-focus images into supervisory signals for unsupervised depth estimation in our WDDNet. We jointly estimate an all-in-focus image and a disparity map, then use a Reblur and Fstack module to regularize the disparity estimation and image restoration. We conducted extensive experiments on synthetic and real data to demonstrate the competitive performance of our method when compared to state-of-the-art (SOTA) supervised approaches.
Liyuan Pan, Richard I. Hartley, Liu Liu 0009, Shah Ariful Hoque Chowdhury, Yan Yang 0011, Hongdong Li, Miaomiao Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Spatial transcriptomics analysis of gene expression prediction using exemplar guided graph neural network
abstract
Spatial transcriptomics (ST) is essential for understanding diseases and developing novel treatments. It measures the gene expression of each fine-grained area (i.e., different windows) in the tissue slide with low throughput. This paper proposes an exemplar guided graph network dubbed EGGN to accurately and efficiently predict gene expression from each window of a tissue slide image. We apply exemplar learning to dynamically boost gene expression prediction from nearest/similar exemplars of a given tissue slide image window. Our framework has three main components connected in a sequence: (i) an extractor to structure a feature space for exemplar retrievals; (ii) a graph construction strategy to connect windows and exemplars as a graph; (iii) a graph convolutional network backbone to process window and exemplar features, and a graph exemplar bridging block to adaptively revise the window features using its exemplars. Finally, we complete the gene expression prediction task with a simple attention-based prediction block. Experiments on standard benchmark datasets indicate the superiority of our approach when compared with past state-of-the-art methods. We release our code at https://github.com/Yan98/EGN.
Yan Yang 0011, Eric A. Stone, Shafin Rahman
Pattern Recognit.1
2023 Identifying Disease-related Brain Imaging Quantitative Traits and Related Genetic Variations via A Bidirectional Association Learning Method
abstract
Discovering critical genetic biomarkers of Alzheimer’s disease (AD) by detecting the complex associations between genotypes (i.e. single nucleotide polymorphism, SNP) and phenotypes (i.e. quantitative trait, QT) is a long-standing and beneficial task for the diagnosis and the follow-up treatment of patients. The function of genes and their relationships with phenotypes are extremely complex. A lot of imaging genetic methods have been designed to uncover the association between brain imaging QTs and SNPs. However, most of them are focused on the effect of a single SNP, which may have limited ability due to the oligogenic or polygenic characteristic of AD. In this paper, we propose a deep reconstruction bidirectional association with feature selection (DRBA-FS) method to explore the multi-SNPmulti-QT associations. In this method, the co-effect of multiple AD-related genetic variations is identified and aggregated, and their high-level genetic associations to brain imaging QTs are jointly modeled. Experiment results on real neuroimaging genetic data from Alzheimer’s Disease Neuroimaging Initiative (ADNI) show that the identified biomarkers are all related to AD. Interestingly, our method can learn the joint effect of multiple AD-related genetic variations across the genome, and thus has significant potential in understanding the genetic mechanism of AD.
Muheng Shang, Yan Yang 0011, Minjianan Zhang, Jin Zhang 0023, Duo Xi, Lei Guo 0002, Lei Du 0001
BIBM2
2023 K3DN: Disparity-Aware Kernel Estimation for Dual-Pixel Defocus Deblurring
abstract
The dual-pixel (DP) sensor captures a two-view image pair in a single snapshot by splitting each pixel in half. The disparity occurs in defocus blurred regions between the two views of the DP pair, while the in-focus sharp regions have zero disparity. This motivates us to propose a K3DN framework for DP pair deblurring, and it has three modules: i) a disparity-aware deblur module. It estimates a disparity feature map, which is used to query a trainable kernel set to estimate a blur kernel that best describes the spatially-varying blur. The kernel is constrained to be symmetrical per the DP formulation. A simple Fourier transform is performed for deblurring that follows the blur model; ii) a reblurring regularization module. It reuses the blur kernel, performs a simple convolution for reblurring, and regularizes the estimated kernel and disparity feature unsupervisedly, in the training stage; iii) a sharp region preservation module. It identifies in-focus regions that correspond to areas with zero disparity between DP images, aims to avoid the introduction of noises during the deblurring process, and improves image restoration performance. Experiments on four standard DP datasets show that the proposed K3DN outperforms state-of-the-art methods, with fewer parameters and flops at the same time.
Yan Yang 0011, Liyuan Pan, Liu Liu 0009, Miaomiao Liu 0001
CVPR1
2023 Event Camera Data Pre-training
abstract
This paper proposes a pre-trained neural network for handling event camera data. Our model is a self-supervised learning framework, and uses paired event camera data and natural RGB images for training. Our method contains three modules connected in a sequence: i) a family of event data augmentations, generating meaningful event images for self-supervised training; ii) a conditional masking strategy to sample informative event patches from event images, encouraging our model to capture the spatial layout of a scene and accelerating training; iii) a contrastive learning approach, enforcing the similarity of embeddings between matching event images, and between paired event and RGB images. An embedding projection loss is proposed to avoid the model collapse when enforcing the event image embedding similarities. A probability distribution alignment loss is proposed to encourage the event image to be consistent with its paired RGB image in the feature space. Transfer learning performance on downstream tasks shows the superiority of our method over state-of-the-art methods. For example, we achieve top-1 accuracy at 64.83% on the N-ImageNet dataset. Our code is available at https://github.com/Yan98/Event-Camera-Data-Pre-training.
Yan Yang 0011, Liyuan Pan, Liu Liu 0009
ICCV1
2023 Exemplar Guided Deep Neural Network for Spatial Transcriptomics Analysis of Gene Expression Prediction
abstract
Spatial transcriptomics (ST) is essential for understanding diseases and developing novel treatments. It measures gene expression of each fine-grained area (i.e., different windows) in the tissue slide with low throughput. This paper proposes an Exemplar Guided Network (EGN) to accurately and efficiently predict gene expression directly from each window of a tissue slide image. We apply exemplar learning to dynamically boost gene expression prediction from nearest/similar exemplars of a given tissue slide image window. Our EGN framework composes of three main components: 1) an extractor to structure a representation space for unsupervised exemplar retrievals; 2) a vision transformer (ViT) backbone to progressively extract representations of the input window; and 3) an Exemplar Bridging (EB) block to adaptively revise the intermediate ViT representations by using the nearest exemplars. Finally, we complete the gene expression prediction task with a simple attention-based prediction block. Experiments on standard benchmark datasets indicate the superiority of our approach when comparing with the past state-of-the-art (SOTA) methods.
Yan Yang 0011, Eric A. Stone, Shafin Rahman
WACV1
2022 Less is More: Facial Landmarks can Recognize a Spontaneous Smile
Md. Tahrim Faroque, Yan Yang 0011, Sheikh Motahar Naim, Nabeel Mohammed, Shafin Rahman
BMVC2
2022 ISG: I can See Your Gene Expression
Yan Yang 0011, Liyuan Pan, Liu Liu 0009, Eric A. Stone
BMVC1
2022 S2FGAN: Semantically Aware Interactive Sketch-to-Face Translation
abstract
Interactive facial image manipulation attempts to edit single and multiple face attributes using a photo-realistic face and/or semantic mask as input. In the absence of the photo-realistic image (only sketch/mask available), previous methods only retrieve the original face but ignore the potential of aiding model controllability and diversity in the translation process. This paper proposes a sketch-to-image generation framework called S2FGAN, aiming to improve users’ ability to interpret and flexibility of face attribute editing from a simple sketch. First, to restore a vivid face from a sketch, we propose semantic level perceptual loss to increase the translation quality. Second, we dedicate the theoretic analysis of attribute editing and build attribute mapping networks with latent semantic loss to modify latent space semantics of Generative Adversarial Networks (GANs). The users can command the model to retouch the generated images by involving the semantic information in the generation process. In this way, our method can manipulate single or multiple face attributes by only specifying attributes to be changed. Extensive experimental results on the CelebAMask-HQ dataset empirically show our superior performance and effectiveness on this task. Our method successfully outperforms state-of-the-art sketch-to-image generation and attribute manipulation methods by exploiting greater control of attribute intensity.
Yan Yang 0011, Tom Gedeon, Shafin Rahman
WACV1
2021 FSE: a Powerful Feature Augmentation Technique for Classification Task
Yaozhong Liu, Yan Yang 0011
ICONIP (2)2
2020 RealSmileNet: A Deep End-to-End Network for Spontaneous and Posed Smile Recognition
Yan Yang 0011, Tom Gedeon, Shafin Rahman
ACCV (5)1