VLDB 2026 Research / reviewers in the wild / expert
Yuting Bai
dblp:226/9551 · also Yu-Ting Bai
· DBLP profile ↗
23ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-8047-1010ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Level Domain Adaptation and Contrastive Domain Isolation with Bilinear Fusion for Patient Drug Response PredictionabstractAccurate prediction of patient drug response is critical for precision cancer medicine but remains constrained by limited clinical data. While in vitro cell line data offer a scalable alternative, effective cross-domain transfer remains challenging. Many existing methods tend to overlook heterogeneous domain shifts across biological contexts, underrepresent the intrinsic differences between cell lines and patient tissues, and insufficiently capture high-order gene-drug interactions. To address these challenges, we propose MACB-DRP, a hierarchical transfer learning framework comprising three complementary stages that progressively coordinate adaptation across tissue, drug, and sample levels while enabling representation separation. The framework begins with tissue-aware domain adaptation, leveraging cancer-type classification and unsupervised alignment to preserve biologically meaningful structure across domains. It then incorporates drug-conditioned adversarial transfer for distribution alignment, coupled with bilinear fusion to model nonlinear and high-order gene-drug interactions. Finally, contrastive anchoring with feature-matched pairs enables fine-grained sample-level alignment, while feature-mismatched negatives preserve irreducible biological disparities. Experimental evaluation demonstrates that MACB-DRP achieves comprehensive predictive performance for patient drug responses, with robust results across multiple cancer types and nine drugs, and further reveals hierarchical structure across drugs and tissues in the visualization. These findings highlight the potential of biologically guided domain adaptation for improving translational pharmacogenomics. Yuting Bai, Hanwen Lv, Wanwan Shi, Zhiyi Zou, Jiawei Luo 0001 |
AAAI | 1 |
| 2026 | RageSense: Leveraging Acoustic Sensing and LLM-Based Intervention for Emotion Regulation in Mobile GamingabstractRageSense introduces a novel system for detecting and regulating player frustration during mobile gaming. Instead of relying on coarse emotion labels, RageSense estimates users’ valence and arousal levels in real time using near-ultrasonic acoustic sensing. By analyzing facial muscle movements via built-in smartphone speakers and microphones, our approach enables emotion sensing without requiring cameras or wearables, constituting a more unobtrusive, environment-resilient, and privacy-friendly approach than traditional emotion recognition. To transform detection into action, we integrate a large language model (LLM) that generates empathetic, context-aware interventions based on gameplay screenshots, behavioral signals, and emotional trajectories. These interventions are delivered in real time, tailored to the user’s emotional state, and designed to mitigate rage while enhancing player well-being. In a 53-participant field study, our system improved emotional state immediately after triggers and was preferred over random or template-based messages. To our knowledge, this is the first demonstration of near-ultrasonic, on-phone valence-arousal regression during mobile gameplay that directly drives real-time, context-aware interventions. Ruihao Zheng, Junbin Ren, Kaiyi Guo, Qian Zhang 0012, Dong She, Yuting Bai, Zhanpeng Jin, Yang Gao 0025 |
CHI | 7 |
| 2026 | Adapt, project and fuse: Parameter-efficient framework for vision large language models
Yuting Bai, Tonghua Su, Zixing Bai |
Neurocomputing | 1 |
| 2026 | A random forest-guided hybrid evolutionary framework for stochastic multiobjective knapsack problems
Yuting Bai, Yanhui Tang |
Inf. Sci. | 1 |
| 2026 | EPILOGUE: Multi-View Graph Contrastive Learning for Gene Function PredictionabstractThe integration of biological networks provides crucial support for accurate gene function prediction, a task that aims to assign genes to corresponding functional categories through computational methods. However, existing approaches struggle with multi-source heterogeneous networks due to their limited ability to capture complex nonlinear dependencies. Contrastive learning, which captures data distributions by measuring similarities and dissimilarities between samples, can generate semantically rich feature representations, offering a new approach to address the aforementioned issues. In this work, we propose EPILOGUE, a multi-view graph contrastive learning framework for gene function prediction. By integrating graph neural networks with contrastive learning, EPILOGUE enables the extraction of high-quality, discriminative gene representations for accurate functional annotation. Additionally, protein sequences are used as node features, offering biological information beyond network topology and supporting the learning of comprehensive semantic representations. Experiments on yeast and human datasets from the STRING database demonstrate that EPILOGUE outperforms nine state-of-the-art methods across six evaluation metrics, validating its effectiveness in learning semantically rich representations for gene function annotation. Yue Zhang 0045, Yuting Bai, Endai Guo, Kening Zhao, Weitian Huang, Hongmin Cai |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2026 | Image-Enhanced Multi-Modal Contrastive Transformer for Subcellular Spatial TranscriptomicsabstractRecent advances in spatial molecular imaging technologies have enabled gene expression profiling alongside high-resolution imaging, providing unprecedented opportunities to resolve molecular heterogeneity at subcellular resolution. However, these technologies fail to fully capture cellular characteristics due to the limited number of genes they can detect, which hinder downstream analysis. Spatial imaging data provide high-resolution and fine-grained morphology information, developing computational methods that effectively integrate image features with transcriptomic profiles is crucial for enabling comprehensive subcellular data analysis. In this study, we present SIMMT, an image-enhanced multi-modal contrastive transformer framework for identifying spatial domains and enhancing subcellular data. In the framework, we design a dual transformer architecture to learn multi-modal representations for cells by modeling transcriptomics and morphological images respectively. To fully capture modality interactions within spatial contexts, we introduce a contrastive learning module that enhances cell representation by aligning tissue morphology and gene expression at the cell level. We tested SIMMT on subcellular spatial transcriptomics datasets from human lung cancer tissue, mouse brain tissue, human colorectal cancer tissue, and human ovarian cancer tissue. The results demonstrated that SIMMT consistently outperformed state-of-the-art methods in spatial clustering and gene expression pattern analysis. Our method also effectively demonstrated its ability to identify tumor spatial heterogeneity and uncover potential gene biomarkers in the human bronchiolar adenoma (BA) dataset. Wanwan Shi, Ying Liu 0027, Qiu Xiao, Yuting Bai, Xinling Zeng, Chee Keong Kwoh 0001, Jiawei Luo 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Fine-Grained Cross-Attention Between Drug Structures and Genes for Perturbation PredictionabstractDrug perturbation prediction aims to accelerate drug target discovery by simulating transcriptional responses induced by chemical interventions. Although deep learning methods have made some progress in this field, they often adopt coarse-grained linear integration strategies (e.g., summation or concatenation) to integrate a single, holistic representation of the gene expression profile with a drug representation, which makes it difficult to capture the subtle perturbations. Furthermore, these methods are predominantly composed of Multi-Layer Perceptrons (MLPs), which cannot dynamically focus on key gene features and tend to overfit on the majority of genes that remain unchanged. To address these challenges, we introduce DGCAP, a novel framework specifically designed for the finegrained modeling relationships between drug structures and gene perturbation responses. Crucially, a cross-attention mechanism dynamically identifies and weights the associations between specific drug structures and individual genes, which enables DGCAP to effectively capture subtle perturbation patterns and mitigate overfitting on unchanged genes. Extensive evaluations across diverse datasets indicate that DGCAP outperforms state-of-the-art methods. Hanwen Lv, Yuting Bai, Jiawei Luo 0001 |
BIBM | 2 |
| 2025 | High-Frequency-Aware Graph Integration for Subcellular Spatial TranscriptomicsabstractRecent advances in spatial transcriptomics have enabled subcellular-resolution profiling of gene expression, offering unprecedented opportunities to investigate intracellular architecture and local microenvironmental interactions. Graph neural networks (GNNs) have shown great promise in modeling spatial transcriptomics data. However, existing GNN-based methods primarily focus on low-frequency signals, overlooking high-frequency signals critical for resolving transcriptional differences across subcellular compartments and cell boundaries. This limits their ability to characterize fine-grained structural and functional heterogeneity within tissues, hindering accurate spatial domain identification. In this study, we propose HiFi-ST, a High-Frequency-Aware Graph Integration framework for subcellular spatial transcriptomics. HiFi-ST employs a high-pass filter to extract high-frequency transcriptional differences, which are then integrated with spatial contexts through a transformer-based architecture. A contrastive learning module is designed to enhance cell representation by aligning spatial organization with transcriptional heterogeneity. Comprehensive experiments on subcellular datasets demonstrated that HiFi-ST consistently outperformed six state-of-the-art methods in spatial clustering, gene expression enhancement, and niche identification. Wanwan Shi, Yahui Long, Ying Liu 0027, Qiu Xiao, Yuting Bai, Xiaoyi Peng, Xiangtao Chen, Jiawei Luo 0001 |
BIBM | 6 |
| 2025 | Improving Multimodal Large Language Models through Combining Resampler and MLP ProjectionsabstractCurrent multimodal large language models (LLMs) achieve impressive performance through connecting visual encoders and LLMs by resampler or MLP projections. Although the MLP projections are effective and widely used in recent works when compared to resampler projections, it has the flaw of the inability to change the number of tokens, which limits the application scenarios. In this paper, we propose a new projection structure that inherits the efficiency of traditional MLP projections and has the ability to flexibly change the number of tokens. Furthermore, we propose a new multimodal LLM with the proposed projection applied. To validate the performance, we evaluate the proposed model on the first large-scale multimodal science question dataset ScienceQA, and the first multimodal LLM evaluation benchmark MME. Our model achieves a new state-of-the-art (SoTA) result of 94.27% on ScienceQA, which is higher than the previous SoTA result of 92.53% achieved by the LLaVA+GPT-4 judge model. On the MME benchmark, our model achieves better performance when compared to representative chatbot models BLIP-2 and InstructBLIP, with significantly lower training costs. Zixing Bai, Yuting Bai |
ICASSP | 2 |
| 2025 | Exploring the Role of CLIP Global Visual Features in Multimodal Large Language ModelsabstractThe next recognized development direction of large language models (LLMs) is to integrate and enhance multimodal capability. Although current multimodal large language models (MLLMs) have achieved impressive performance by combining the pre-trained visual encoder CLIP and LLM, these works mainly focus on using the CLIP patch visual features. In practice, we find that CLIP global visual features are more efficient than patch visual features in some scenarios, especially in multimodal reasoning tasks. Therefore, we explore the role of CLIP global visual features and propose a new MLLM with the usage of full global visual features in this paper. Our model adopts the parameter-efficient transfer learning (PETL) method Adapter to fine-tune the pre-trained models and a simple MLP-based network to connect the visual encoder and LLM. To validate the performance, we evaluate our model on the first large-scale multimodal science question dataset, ScienceQA. Our model achieves a new state-of-the-art (SoTA) result of 93.96% on ScienceQA, which is higher than the previous SoTA result of 92.53%. Zixing Bai, Yuting Bai |
ICASSP | 2 |
| 2025 | Wrist2Finger: Sensing Fingertip Force for Force-Aware Hand Interaction with a Ring-Watch WearableabstractHand pose tracking is essential for advancing applications in human-computer interaction. Current approaches, such as vision-based systems and wearable devices, face limitations in portability, usability, and practicality. We present a novel wearable system that reconstructs 3D hand pose and estimates per-finger forces using a minimal ring-watch sensor setup. A ring worn on the finger integrates an inertial measurement unit (IMU) to capture finger motion, while a smartwatch-based single-channel electromyography (EMG) sensor on the wrist detects muscle activations. By leveraging the complementary strengths of motion sensing and muscle signals, our approach achieves accurate hand pose tracking and grip force estimation in a compact wearable form factor. We develop a dual-branch transformer network that fuses IMU and EMG data with cross-modal attention to predict finger joint positions and forces simultaneously. A custom loss function imposes kinematic constraints for smooth force variation and realistic force saturation. Evaluation with 20 participants performing daily object interaction gestures demonstrates an average Mean Per Joint Position Error (MPJPE) of 0.57 cm and a fingertip force estimation (RMSE: 0.213, r=0.76). We showcase our system in a real-time Unity application, enabling virtual hand interactions that respond to user-applied forces. This minimal, force-aware tracking system has broad implications for VR/AR, assistive prosthetics, and ergonomic monitoring. Yingjing Xiao, Junbin Ren, Yuting Bai, Zhanpeng Jin, Yang Gao 0025 |
UIST | 4 |
| 2025 | STCGAN: a novel cycle-consistent generative adversarial network for spatial transcriptomics cellular deconvolutionabstractMOTIVATION: Spatial transcriptomics (ST) technologies have revolutionized our ability to map gene expression patterns within native tissue context, providing unprecedented insights into tissue architecture and cellular heterogeneity. However, accurately deconvolving cell-type compositions from ST spots remains challenging due to the sparse and averaged nature of ST data, which is essential for accurately depicting tissue architecture. While numerous computational methods have been developed for cell-type deconvolution and spatial distribution reconstruction, most fail to capture tissue complexity at the single-cell level, thereby limiting their applicability in practical scenarios. RESULTS: To this end, we propose a novel cycle-consistent generative adversarial network named STCGAN for cellular deconvolution in spatial transcriptomic. STCGAN first employs a cycle-consistent generative adversarial network (CGAN) to pre-train on ST data, ensuring that both the mapping from ST data to latent space and its reverse mapping are consistent, capturing complex spatial gene expression patterns and learning robust latent representations. Based on the learned representation, STCGAN then optimizes a trainable cell-to-spot mapping matrix to integrate scRNA-seq data with ST data, accurately estimating cellular composition within each capture spot and effectively reconstructing the spatial distribution of cells across the tissue. To further enhance deconvolution accuracy, we incorporate spatial-aware regularization that ensures accurate cellular distribution reconstruction within the spatial context. Benchmarking against seven state-of-the-art methods on five simulated and real datasets from various tissues, STCGAN consistently delivers superior cell-type deconvolution performance. AVAILABILITY: The code of STCGAN can be downloaded from https://github.com/cs-wangbo/STCGAN and all the mentioned datasets are available on Zenodo at https://zenodo.org/doi/10.5281/zenodo.10799113. Yahui Long, Yuting Bai, Jiawei Luo 0001, Chee Keong Kwoh 0001 |
Briefings Bioinform. | 3 |
| 2025 | Fusion Network Model Based on Broad Learning System for Multidimensional Time-Series ForecastingabstractMultidimensional time‐series prediction is significant in various fields, such as human production and life, weather forecasting, and artificial intelligence. However, a single model can only focus on specific features of time‐series data, making it unable to consider both linear and nonlinear components simultaneously. In this study, we propose a fusion network that combines the advantages of deep and broad networks for multidimensional time‐series prediction tasks. The complex multidimensional time‐series data are divided into nonlinear and time‐series data. Restricted Boltzmann machine and mapping functions are used for feature learning and generating mapping nodes at the mapping layer. The echo state network and gate recurrent unit are applied in the enhancement layer. The proposed model has been validated on PM2.5 and wind turbine power datasets, proving superior performance in multistep prediction tasks compared to the baseline models. Yuting Bai, Xinyi Xue, Xue-bo Jin 0001, Zhiyao Zhao |
Int. J. Intell. Syst. | 1 |
| 2025 | scTrans: Sparse attention powers fast and accurate cell type annotation in single-cell RNA-seq dataabstractCell type annotation is crucial in single-cell RNA sequencing data analysis because it enables significant biological discoveries and deepens our understanding of tissue biology. Given the high-dimensional and highly sparse nature of single-cell RNA sequencing data, most existing annotation tools focus on highly variable genes to reduce dimensionality and computational load. However, this approach inevitably results in information loss, potentially weakening the model's generalization performance and adaptability to novel datasets. To mitigate this issue, we developed scTrans, a single cell Transformer-based model, which employs sparse attention to utilize all non-zero genes, thereby effectively reducing the input data dimensionality while minimizing information loss. We validated the speed and accuracy of scTrans by performing cell type annotation on 31 different tissues within the Mouse Cell Atlas. Remarkably, even with datasets nearing a million cells, scTrans efficiently perform cell type annotation in limited computational resources. Furthermore, scTrans demonstrates strong generalization capabilities, accurately annotating cells in novel datasets and generating high-quality latent representations, which are essential for precise clustering and trajectory analysis. Zhiyi Zou, Ying Liu 0027, Yuting Bai, Jiawei Luo 0001, Zhaolei Zhang |
PLoS Comput. Biol. | 3 |
| 2025 | Cross-modal feature fusion via mutual assistance: a novel network for enhanced object detection
Xue-bo Jin 0001, Hui-Jun Ma, Tingli Su 0001, Jian-Lei Kong, Yuting Bai |
Vis. Comput. | 6 |
| 2024 | SpaGIC: graph-informed clustering in spatial transcriptomics via self-supervised contrastive learningabstractSpatial transcriptomics technologies enable the generation of gene expression profiles while preserving spatial context, providing the potential for in-depth understanding of spatial-specific tissue heterogeneity. Leveraging gene and spatial data effectively is fundamental to accurately identifying spatial domains in spatial transcriptomics analysis. However, many existing methods have not yet fully exploited the local neighborhood details within spatial information. To address this issue, we introduce SpaGIC, a novel graph-based deep learning framework integrating graph convolutional networks and self-supervised contrastive learning techniques. SpaGIC learns meaningful latent embeddings of spots by maximizing both edge-wise and local neighborhood-wise mutual information of graph structures, as well as minimizing the embedding distance between spatially adjacent spots. We evaluated SpaGIC on seven spatial transcriptomics datasets across various technology platforms. The experimental results demonstrated that SpaGIC consistently outperformed existing state-of-the-art methods in several tasks, such as spatial domain identification, data denoising, visualization, and trajectory inference. Additionally, SpaGIC is capable of performing joint analyses of multiple slices, further underscoring its versatility and effectiveness in spatial transcriptomics research. Wei Liu 0296, Yuting Bai, Jiawei Luo 0001 |
Briefings Bioinform. | 3 |
| 2024 | A novel broad learning system integrated with restricted Boltzmann machine and echo state network for time series forecasting
Yuting Bai, Xue-bo Jin 0001, Zhiyao Zhao, Tingli Su 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | AMNet: a new RGB-D instance segmentation network based on attention and multi-modality
Lihua Hu, Yuting Bai, Xiaoling Yao, Sulan Zhang |
Vis. Comput. | 3 |
| 2023 | Optimal Deployment for Hybrid Sensor Networks Based on Efficient Node ConfigurationabstractHybrid sensor networks, which contain mobile nodes and stationary nodes, are being used more and more widely. The second deployment of mobile nodes is a key problem to be solved, and the deployment performance of the network directly affects the monitoring effect of the network. Optimizing the configuration ratio of the two nodes can effectively reduce the network cost. In this paper, under the premise of knowing the coverage of the required monitoring area, the impact of sensor devices on node configuration is studied through parameter analysis, and the number and types of sensors that should be deployed in the hybrid sensor network are deduced, which can be conveniently and accurately used to design the actual hybrid sensor network. At the same time, for the secondary deployment of mobile nodes, this paper proposes a new mobile coverage method BS‐CCP (box search and concentric circle positioning) to improve the coverage of the hybrid sensor network and maximize the coverage of the target area with the specified sensor types and numbers. Compared with existing work, the method in this paper reduces the number of iterations and reduces the number of required nodes. Comparing BS‐CCP with the existing network mobile coverage algorithm, the experimental results show that the coverage obtained by this method is larger and more efficient. Qian Sun 0012, Xiaoyi Wang 0001, Zhiyao Zhao, Jiping Xu, Li Wang 0068, Huiyan Zhang 0002, Yuting Bai |
Int. J. Intell. Syst. | 9 |
| 2023 | AdaHOSVD: an adaptive higher-order singular value decomposition method for point cloud denoising
Lihua Hu, Wenhao Liang, Yuting Bai, Jifu Zhang |
Pattern Anal. Appl. | 3 |
| 2022 | Water quality evolution mechanism modeling and health risk assessment based on stochastic hybrid dynamic systems
Zhiyao Zhao, Yuqin Zhou, Xiaoyi Wang 0001, Yuting Bai |
Expert Syst. Appl. | 5 |
| 2022 | Continuous Positioning with Recurrent Auto-Regressive Neural Network for Unmanned Surface Vehicles in GPS Outages
Yuting Bai, Zhiyao Zhao, Xiaoyi Wang 0001, Xue-bo Jin 0001, Baihai Zhang |
Neural Process. Lett. | 1 |
| 2021 | Parallel deep prediction with covariance intersection fusion on non-stationary time series
Zhigang Shi, Yuting Bai, Xue-bo Jin 0001, Xiaoyi Wang 0001, Tingli Su 0001, Jian-Lei Kong |
Knowl. Based Syst. | 2 |