Yixin Wang 0003

dblp:32/6839-3 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0002-8062-0765ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Tabula: A Tabular Self-Supervised Foundation Model for Single-Cell Transcriptomics
abstract
Foundation models (FMs) have shown great promise in single-cell genomics, yet current approaches, such as scGPT, Geneformer, and scFoundation, rely on centralized training and language modeling objectives that overlook the tabular nature of single-cell data and raise significant privacy concerns. We present TABULA, a foundation model designed for single-cell transcriptomics, which integrates a novel tabular modeling objective and federated learning framework to enable privacy-preserving pretraining across decentralized datasets. TABULA directly models the cell-by-gene expression matrix through column-wise gene reconstruction and row-wise cell contrastive learning, capturing both gene-level relationships and cell-level heterogeneity without imposing artificial gene sequence order. Extensive experiments demonstrate the effectiveness of TABULA: despite using only half the pretraining data, TABULA achieves state-of-the-art performance across key tasks, including gene imputation, perturbation prediction, cell type annotation, and multi-omics integration. It is important to note that as public single-cell datasets continue to grow, TABULA provides a scalable and privacy-aware foundation that not only validates the feasibility of federated tabular modeling but also establishes a generalizable framework for training future models under similar privacy-preserving settings.
Jiayuan Ding, Jianhui Lin, Yixin Wang 0003, Ziyang Miao, Zhaoyu Fang, Jiliang Tang, Xiaojie Qiu
NeurIPS4
2024 Trust it or not: Confidence-guided automatic radiology report generation
Yixin Wang 0003, Zihao Lin 0003, Zhe Xu 0012, Jie Luo 0003, Jiang Tian, Zhongchao Shi, Lifu Huang, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002
Neurocomputing1
2024 Deep Learning in Single-cell Analysis
abstract
Single-cell technologies are revolutionizing the entire field of biology. The large volumes of data generated by single-cell technologies are high dimensional, sparse, and heterogeneous and have complicated dependency structures, making analyses using conventional machine learning approaches challenging and impractical. In tackling these challenges, deep learning often demonstrates superior performance compared to traditional machine learning methods. In this work, we give a comprehensive survey on deep learning in single-cell analysis. We first introduce background on single-cell technologies and their development, as well as fundamental concepts of deep learning including the most popular deep architectures. We present an overview of the single-cell analytic pipeline pursued in research applications while noting divergences due to data sources or specific applications. We then review seven popular tasks spanning different stages of the single-cell analysis pipeline, including multimodal integration, imputation, clustering, spatial domain identification, cell-type deconvolution, cell segmentation, and cell-type annotation. Under each task, we describe the most recent developments in classical and deep learning methods and discuss their advantages and disadvantages. Deep learning tools and benchmark datasets are also summarized for each task. Finally, we discuss the future directions and the most recent challenges. This survey will serve as a reference for biologists and computer scientists, encouraging collaborations.
Dylan Molho, Jiayuan Ding, Wenzhuo Tang, Zhaoheng Li, Hongzhi Wen, Yixin Wang 0003, Julian Venegas, Wei Jin 0009, Renming Liu, Runze Su, Patrick Danaher, Robert Yang, Yu L. Lei, Yuying Xie 0001, Jiliang Tang
ACM Trans. Intell. Syst. Technol.6
2024 A Survey of Visual Transformers
abstract
Transformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP). Inspired by such significant achievements, some pioneering works have recently been done on employing Transformer-liked architectures in the computer vision (CV) field, which have demonstrated their effectiveness on three fundamental CV tasks (classification, detection, and segmentation) as well as multiple sensory data stream (images, point clouds, and vision-language data). Because of their competitive modeling capabilities, the visual Transformers have achieved impressive performance improvements over multiple benchmarks as compared with modern convolution neural networks (CNNs). In this survey, we have reviewed over 100 of different visual Transformers comprehensively according to three fundamental CV tasks and different data stream types, where taxonomy is proposed to organize the representative methods according to their motivations, structures, and application scenarios. Because of their differences on training settings and dedicated vision tasks, we have also evaluated and compared all these existing visual Transformers under different configurations. Furthermore, we have revealed a series of essential but unexploited aspects that may empower such visual Transformers to stand out from numerous architectures, e.g., slack high-level semantic embeddings to bridge the gap between the visual Transformers and the sequential ones. Finally, two promising research directions are suggested for future investment. We will continue to update the latest articles and their released source codes at https://github.com/liuyang-ict/awesome-visual-transformers.
Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
IEEE Trans. Neural Networks Learn. Syst.3
2023 SAP-DETR: Bridging the Gap Between Salient Points and Queries-Based Transformer Detector for Fast Model Convergency
abstract
Recently, the dominant DETR-based approaches apply central-concept spatial prior to accelerating Transformer detector convergency. These methods gradually refine the reference points to the center of target objects and imbue object queries with the updated central reference information for spatially conditional attention. However, centralizing reference points may severely deteriorate queries' saliency and confuse detectors due to the indiscriminative spatial prior. To bridge the gap between the reference points of salient queries and Transformer detectors, we propose SAlient Point-based DETR (SAP-DETR) by treating object detection as a transformation from salient points to instance objects. Concretely, we explicitly initialize a query-specific reference point for each object query, gradually aggregate them into an instance object, and then predict the distance from each side of the bounding box to these points. By rapidly attending to query-specific reference regions and the conditional box edges, SAP-DETR can effectively bridge the gap between the salient point and the query-based Transformer detector with a significant convergency speed. Experimentally, SAP-DETR achieves 1.4× convergency speed with competitive performance and stably promotes the SoTA approaches by ∼1.0 AP. Based on ResNet-DC-101, SAP-DETR achieves 46.9 AP. The code will be released at https://github.com/liuyang-ict/SAP-DETR.
Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
CVPR3
2023 Long-Tailed Recognition with Causal Invariant Transformation
abstract
Standard classification models rely on the assumption that all the classes of interest are equally represented in training datasets. However, visual phenomena exhibit a long-tailed distribution, such that many standard approaches fail to properly model and result in a considerable degeneration on accuracy. The recent methods have produced encouraging results, but their efforts only seek to simulate the statistical relationship between data and labels and compensate for imbalanced data-related issues, without addressing the underlying causal mechanisms. In this paper, a comprehensive structural causal model is developed to excavate the intrinsic causal mechanism between data and labels. Specifically, we assume that each input is constructed from a mix of causal factors and non-causal factors, and only the causal factors cause the classification judgments. In order to extract such causal factors from inputs and then reconstruct the invariant causal mechanisms, we propose a Causal Invariant Transformation algorithm for Long-tailed recognition (CITL), which generates diverse data to avoid the over-fitting on the tail classes and enforces the learnt representations to maintain the causal factors and eliminate the non-causal factors. Our extensive experimental results on several widely used datasets have demonstrated the effectiveness of our proposed CITL approach.
Yahong Zhang, Sheng Shi, Yixin Wang 0003, Wenli Ouyang, WeiFan, Jianping Fan 0007
ICASSP4
2023 Learnable Query Guided Representation Learning for Treatment Effect Estimation
abstract
The estimation of Individual Treatment Effect (ITE) is a challenging problem in causal inference, due to the missing counterfactual data and the selection bias. In this paper, we propose a novel representation learning framework via the Learnable Query based transformer for Treatment Effect Estimation (LQTEE). A certain number of queries are learned in a sample-agnostic way and extract global critical features from covariate and treatment data separately. We also propose the hierarchical propensity score regularized adversarial loss to obtain balanced covariate representations, and the mutual orthogonal constraint to force queries to focus on diverse parts of covariates, thus the impact of instrumental variables can be adaptively reduced. Treatment representation learning enables our estimator to support general-purpose treatments, and more importantly, it can reveal the underlying patterns of data-generation process efficiently. Extensive experiments show that our ITE estimator significantly outperforms the state-of-the-art methods.
Yixin Wang 0003, Yahong Zhang, Wenli Ouyang, Sheng Shi, Jianping Fan 0007
IJCNN2
2023 Towards Expert-Amateur Collaboration: Prototypical Label Isolation Learning for Left Atrium Segmentation with Mixed-Quality Labels
Zhe Xu 0012, Jiangpeng Yan, Donghuan Lu, Yixin Wang 0003, Jie Luo 0003, Yefeng Zheng 0001, Raymond Kai-Yu Tong
MICCAI (7)4
2023 Ambiguity-selective consistency regularization for mean-teacher semi-supervised medical image segmentation
Zhe Xu 0012, Yixin Wang 0003, Donghuan Lu, Xiangde Luo, Jiangpeng Yan, Yefeng Zheng 0001, Raymond Kai-Yu Tong
Medical Image Anal.2
2022 Cross-Domain Few-Shot Learning for Rare-Disease Skin Lesion Segmentation
abstract
Recently, deep learning (DL)-based skin lesion segmentation in dermoscopic images has advanced the efficient diagnosis of skin diseases. Commonly, most of the DL-based methods require a large amount of training data and can only perform accurate predictions on pre-defined classes. However, there exist some rare skin diseases with very limited labeled samples, which poses great challenges to typical DL-based methods. Few-shot learning (FSL) technique, which aims to train models with abundant seen classes and then generalizes to related unseen classes, is promising in addressing a similar problem. Unfortunately, simply borrowing the typical FSL is infeasible since collecting such abundant seen-class data (common skin diseases), is also difficult. In this paper, we propose a cross-domain few-shot segmentation (CD-FSS) framework, which enables the model to leverage the learning ability obtained from the natural domain, to facilitate rare-disease skin lesion segmentation with limited data of common diseases. Specifically, the framework consists of two processes, i.e., specific learning and generic learning, which are alternately optimized in a meta-training manner. A specific learner and a generic learner are tailored to build relationships between both processes. Experimental results demonstrate that our framework significantly improves the generalization ability from natural domain to unseen medical domain.
Yixin Wang 0003, Zhe Xu 0012, Jiang Tian, Jie Luo 0003, Zhongchao Shi, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002
ICASSP1
2022 On the Dataset Quality Control for Image Registration Evaluation
Jie Luo 0003, Guangshen Ma, Nazim Haouchine, Zhe Xu 0012, Yixin Wang 0003, Tina Kapur, Lipeng Ning, William M. Wells III, Sarah F. Frisken
MICCAI (6)5
2022 Denoising for Relaxing: Unsupervised Domain Adaptive Fundus Image Segmentation Without Source Data
Zhe Xu 0012, Donghuan Lu, Yixin Wang 0003, Jie Luo 0003, Dong Wei 0004, Yefeng Zheng 0001, Raymond Kai-Yu Tong
MICCAI (5)3
2022 All-Around Real Label Supervision: Cyclic Prototype Consistency Learning for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning has substantially advanced medical image segmentation since it alleviates the heavy burden of acquiring the costly expert-examined annotations. Especially, the consistency-based approaches have attracted more attention for their superior performance, wherein the real labels are only utilized to supervise their paired images via supervised loss while the unlabeled images are exploited by enforcing the perturbation-based "unsupervised" consistency without explicit guidance from those real labels. However, intuitively, the expert-examined real labels contain more reliable supervision signals. Observing this, we ask an unexplored but interesting question: can we exploit the unlabeled data via explicit real label supervision for semi-supervised training? To this end, we discard the previous perturbation-based consistency but absorb the essence of non-parametric prototype learning. Based on the prototypical networks, we then propose a novel cyclic prototype consistency learning (CPCL) framework, which is constructed by a labeled-to-unlabeled (L2U) prototypical forward process and an unlabeled-to-labeled (U2L) backward process. Such two processes synergistically enhance the segmentation network by encouraging morediscriminative and compact features. In this way, our framework turns previous "unsupervised" consistency into new "supervised" consistency, obtaining the "all-around real label supervision" property of our method. Extensive experiments on brain tumor segmentation from MRI and kidney segmentation from CT images show that our CPCL can effectively exploit the unlabeled data and outperform other state-of-the-art semi-supervised medical image segmentation methods.
Zhe Xu 0012, Yixin Wang 0003, Donghuan Lu, Lequan Yu, Jiangpeng Yan, Jie Luo 0003, Kai Ma 0002, Yefeng Zheng 0001, Raymond Kai-Yu Tong
IEEE J. Biomed. Health Informatics2
2022 Anti-Interference From Noisy Labels: Mean-Teacher-Assisted Confident Learning for Medical Image Segmentation
abstract
Manually segmenting medical images is expertise-demanding, time-consuming and laborious. Acquiring massive high-quality labeled data from experts is often infeasible. Unfortunately, without sufficient high-quality pixel-level labels, the usual data-driven learning-based segmentation methods often struggle with deficient training. As a result, we are often forced to collect additional labeled data from multiple sources with varying label qualities. However, directly introducing additional data with low-quality noisy labels may mislead the network training and undesirably offset the efficacy provided by those high-quality labels. To address this issue, we propose a Mean-Teacher-assisted Confident Learning (MTCL) framework constructed by a teacher-student architecture and a label self-denoising process to robustly learn segmentation from a small set of high-quality labeled data and plentiful low-quality noisy labeled data. Particularly, such a synergistic framework is capable of simultaneously and robustly exploiting (i) the additional dark knowledge inside the images of low-quality labeled set via perturbation-based unsupervised consistency, and (ii) the productive information of their low-quality noisy labels via explicit label refinement. Comprehensive experiments on left atrium segmentation with simulated noisy labels and hepatic and retinal vessel segmentation with real-world noisy labels demonstrate the superior segmentation performance of our approach as well as its effectiveness on label denoising.
Zhe Xu 0012, Donghuan Lu, Jie Luo 0003, Yixin Wang 0003, Jiangpeng Yan, Kai Ma 0002, Yefeng Zheng 0001, Raymond Kai-Yu Tong
IEEE Trans. Medical Imaging4
2021 ACN: Adversarial Co-training Network for Brain Tumor Segmentation with Missing Modalities
Yixin Wang 0003, Yang Zhang 0002, Yang Liu 0250, Zihao Lin 0003, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
MICCAI (7)1
2021 Noisy Labels are Treasure: Mean-Teacher-Assisted Confident Learning for Hepatic Vessel Segmentation
Zhe Xu 0012, Donghuan Lu, Yixin Wang 0003, Jie Luo 0003, Jayender Jagadeesan, Kai Ma 0002, Yefeng Zheng 0001, Xiu Li 0001
MICCAI (1)3
2021 The state of the art in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: Results of the KiTS19 challenge
Nicholas Heller, Fabian Isensee, Klaus H. Maier-Hein, Xiaoshuai Hou, Chunmei Xie, Fengyi Li, Yang Nan 0002, Guangrui Mu, Miofei Han, Guang Yao, Yaozong Gao, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiawei Yang 0002, Guangwei Xiong, Jiang Tian, Christopher J. Weight
Medical Image Anal.14
2020 Double-Uncertainty Weighted Method for Semi-supervised Learning
Yixin Wang 0003, Yao Zhang 0010, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002
MICCAI (1)1