VLDB 2026 Research / reviewers in the wild / expert
Zhenchao Tang
dblp:184/5916
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConsistentID: Portrait Generation With Multimodal Fine-Grained Identity PreservingabstractDiffusion-based technologies have made significant strides, particularly in personalized and customized facial generation. However, existing methods struggle to achieve high-fidelity and detailed identity (ID) consistency. This is mainly due to two challenges: insufficient fine-grained control over specific facial areas and the absence of a comprehensive strategy for ID preservation that accounts for both intricate facial details and the overall facial structure. To address these limitations, we introduce ConsistentID, an innovative method crafted for diverse identity-preserving portrait generation under fine-grained multimodal facial prompts, utilizing only a single reference image. ConsistentID comprises two core components: a multimodal facial prompt generator and an ID-preservation network. The facial prompt generator combines localized facial features, facial feature descriptions, and overall facial descriptions to enhance the precision of facial detail reconstruction. The ID-preservation network, optimized with a facial attention localization strategy, ensures consistent identity preservation across facial regions. Together, these components leverage fine-grained multimodal identity information to improve identity preservation accuracy significantly. To drive ConsistentID's training, we propose a fine-grained portrait dataset, FGID, with over 500,000 facial images, offering greater diversity and comprehensiveness than existing public facial datasets. Experimental results substantiate that our ConsistentID achieves exceptional precision and diversity in personalized facial generation, surpassing existing methods in the MyStyle dataset. In addition, although ConsistentID introduces more multimodal ID information, it still maintains rapid inference speed during the generation process. Jiehui Huang, Wenhui Song, Zheng Chong, Zhenchao Tang, Yuhao Cheng, Long Chen 0005, Yiqiang Yan, Shengcai Liao, Xiaodan Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Accurate Protein-Protein Interaction Prediction: Based on Multiview Heterogeneous Graph Autoencoders and Random MaskingabstractProtein-protein interaction (PPI) and their interaction sites [PPI site (PPIS)] hold immense potential for elucidating cellular mechanisms and advancing targeted drug development. While deep learning has driven progress in PPI research by capturing protein features, it remains limited by its overreliance on sequence information and inability to effectively integrate protein internal structural features. To address these challenges, we propose MEGAE, a novel model capable of achieving high-precision prediction of PPI and PPIS. MEGAE reconstructs amino acid microenvironments through a vector quantization autoencoder, integrating physicochemical properties, structural details, and sequence data to provide a comprehensive representation of proteins. We innovatively introduce a multiview random masking training strategy, introducing controlled randomness during the reconstruction process to enhance the robustness of microenvironment embeddings. The model combines these fused embeddings with protein graphs and protein interaction networks, leveraging graph neural networks (GNNs) to capture multilevel relationships from local amino acid interactions to global signal network connections-thereby achieving precise predictions. Experimental results demonstrate that MEGAE outperforms state-of-the-art sequence- and structure-based methods across multiple datasets, exhibiting higher accuracy in predicting interaction types and interaction sites. This advancement underscores the potential of microenvironment-aware modeling in uncovering complex protein interactions. Shouzhi Chen, Zhenchao Tang, Linlin You, Calvin Yu-Chian Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | DUSTED: Dual-Attention Enhanced Spatial Transcriptomics DenoiserabstractSpatially Resolved Transcriptomics (SRT) has become an indispensable tool in various fields, including tumor microenvironment identification, neurobiology, and the study of complex tissue architecture. However, the accuracy of these insights is often compromised by noise in spatial transcriptomics data due to technical limitations. While recent advancements in denoising methods have shown some promise, they frequently fall short by neglecting spatial features, overlooking the variability in noise levels among genes, and relying heavily on external histological images for supplementary information. In our study, we propose DUSTED, a Dual-Attention Enhanced Spatial Transcriptomics Denoiser, designed to address these challenges. Built on a graph autoencoder framework, DUSTED utilizes gene channel attention and graph attention mechanisms to simultaneously consider spatial features and noise variability in gene expression data. Additionally, it integrates the negative binomial distribution with or without zero-inflation, ensuring a more accurate fit for gene expression distributions. Benchmark tests using simulated datasets demonstrate that DUSTED outperforms existing methods. Furthermore, in real-world applications with the HOCWTA and DLPFC datasets, DUSTED excels in enhancing the correlation between gene and protein expression, recovering spatial gene expression patterns, and improving clustering results. These improvements underscore its potential impact on advancing our understanding of tumor microenvironments, neural tissue organization, and other biologically significant areas. Zhenchao Tang |
AAAI | 3 |
| 2025 | U-N2C: A Dual Memory-Guided Disentanglement Framework for Unsupervised System Matrix Denoising in Magnetic Particle ImagingabstractRecently, Magnetic Particle Imaging, an emerging functional imaging modality, has exhibited outstanding spatial-temporal resolution and sensitivity. The general reconstruction pipeline of Magnetic Particle Imaging involves calibrating a System Matrix and then solving an ill-posed inverse problem combined with the measured particle signals. However, the introduction of noise during the System Matrix calibration procedure is inevitable, which degrades the detailed information in the reconstructed images. Therefore, frequency selection methods based on signal-to-noise ratio are commonly adopted. However, these methods lead to a decrease in the available high-frequency components, which damages the spatial resolution. To address this problem, we propose an unsupervised memory-guided denoising framework with unpaired noisy-clean System Matrix components, called U-N2C. Specifically, we design a Pattern Memory Block to memorize System Matrix patterns, directed by a position-aware frequency index embedding. Meanwhile, we devise a Noise Memory Block to implicitly approximate noise distributions. With the guidance of our dual memory blocks, we can disentangle the noise and content of the System Matrix in the latent space. Furthermore, benefiting from the ability to model complex noise, our method can generate pseudo but high-quality noisy-clean pairs and further enhance our denoising capability. Experiments on both synthetic and real noise demonstrate that our U-N2C achieves cutting-edge performance compared to other methods. Moreover, we conduct extensive qualitative and quantitative ablation studies to verify the effectiveness of our method. Our code has been available at U-N2C. Wenxuan Zou, Gen Shi, Siao Lei, Guanghui Li 0006, Guangxing Zhou, Yang Jing, Zhenchao Tang, Jie Tian 0001 |
IEEE Trans. Image Process. | 8 |
| 2024 | Comprehensive View Embedding Learning for Single-Cell Multimodal IntegrationabstractMotivation: Advances in single-cell measurement techniques provide rich multimodal data, which helps us to explore the life state of cells more deeply. However, multimodal integration, or, learning joint embeddings from multimodal data remains a current challenge. The difficulty in integrating unpaired single-cell multimodal data is that different modalities have different feature spaces, which easily leads to information loss in joint embedding. And few existing methods have fully exploited and fused the information in single-cell multimodal data. Result: In this study, we propose CoVEL, a deep learning method for unsupervised integration of single-cell multimodal data. CoVEL learns single-cell representations from a comprehensive view, including regulatory relationships between modalities, fine-grained representations of cells, and relationships between different cells. The comprehensive view embedding enables CoVEL to remove the gap between modalities while protecting biological heterogeneity. Experimental results on multiple public datasets show that CoVEL is accurate and robust to single-cell multimodal integration. Data availability: https://github.com/shapsider/scintegration. Zhenchao Tang, Jiehui Huang, Guanxing Chen, Calvin Yu-Chian Chen |
AAAI | 1 |
| 2024 | Progressive network based on detail scaling and texture extraction: A more general framework for image deraining
Jiehui Huang, Zhenchao Tang, Xuedong He, Defeng Zhou, Calvin Yu-Chian Chen |
Neurocomputing | 2 |
| 2024 | A knowledge distillation-guided equivariant graph neural network for improving protein interaction site prediction performance
Shouzhi Chen, Zhenchao Tang, Linlin You, Calvin Yu-Chian Chen |
Knowl. Based Syst. | 2 |
| 2024 | TMBL: Transformer-based multimodal binding learning model for multimodal sentiment analysis
Jiehui Huang, Zhenchao Tang, Calvin Yu-Chian Chen |
Knowl. Based Syst. | 3 |
| 2024 | DSIL-DDI: A Domain-Invariant Substructure Interaction Learning for Generalizable Drug-Drug Interaction PredictionabstractDrug-drug interactions (DDIs) trigger unexpected pharmacological effects in vivo, often with unknown causal mechanisms. Deep learning methods have been developed to better understand DDI. However, learning domain-invariant representations for DDI remains a challenge. Generalizable DDI predictions are closer to reality than source domain predictions. For existing methods, it is difficult to achieve out-of-distribution (OOD) predictions. In this article, focusing on substructure interaction, we propose DSIL-DDI, a pluggable substructure interaction module that can learn domain-invariant representations of DDIs from source domain. We evaluate DSIL-DDI on three scenarios: the transductive setting (all drugs in test set appear in training set), the inductive setting (test set contains new drugs that were not present in training set), and OOD generalization setting (training set and test set belong to two different datasets). The results demonstrate that DSIL-DDI improve the generalization and interpretability of DDI prediction modeling and provides valuable insights for OOD DDI predictions. DSIL-DDI can help doctors ensuring the safety of drug administration and reducing the harm caused by drug abuse. Zhenchao Tang, Guanxing Chen, Hualin Yang, Weihe Zhong, Calvin Yu-Chian Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Weakly Supervised Posture Mining for Fine-Grained ClassificationabstractBecause the subtle differences between the different sub-categories of common visual categories such as bird species, fine-grained classification has been seen as a challenging task for many years. Most previous works focus towards the features in the single discriminative region isolatedly, while neglect the connection between the different discriminative regions in the whole image. However, the relationship between different discriminative regions contains rich posture information and by adding the posture information, model can learn the behavior of the object which attribute to improve the classification performance. In this paper, we propose a novel fine-grained framework named PMRC (posture mining and reverse cross-entropy), which is able to combine with different backbones to good effect. In PMRC, we use the Deep Navigator to generate the discriminative regions from the images, and then use them to construct the graph. We aggregate the graph by message passing and get the classification results. Specifically, in order to force PMRC to learn how to mine the posture information, we design a novel training paradigm, which makes the Deep Navigator and message passing communicate and train together. In addition, we propose the reverse cross-entropy (RCE) and demomenstate that compared to the cross-entropy (CE), RCE can not only promote the accurracy of our model but also generalize to promote the accuracy of other kinds of fine-grained classification models. Experimental results on benchmark datasets confirm that PMRC can achieve state-of-the-art. Zhenchao Tang, Hualin Yang, Calvin Yu-Chian Chen |
CVPR | 1 |
| 2016 | An Effective Non-rigid Image Registration Method Based on Active Demons AlgorithmabstractIn order to solve the problem the homogeneous coefficient of the classic active demons algorithm can not take into account large deformation and small deformation at the same time, this paper presents a non-rigid registration algorithm based on active demons algorithm. The proposed algorithm introduces a new parameter called balance coefficient to the active demons algorithm, which will adjust the driving force combined with homogeneous coefficient. Not only the large deformation and the small deformation can be taken into account at the same time, but also the mutual restraint problem of the convergence speed and the registration accuracy can be eased in a certain extent. In order to further improve the registration accuracy and the convergence speed, and avoid falling into local extreme value, a coarse-to-fine multi-resolution strategy is introduced into the registration process. Experiments on checkboard test images, natural images and medical images demonstrate that the proposed method is faster and more accurate, and the registration accuracy is close to the latest TV-L1 optical flow image registration algorithm, which solves the problems of the active demons algorithm. Zhenchao Tang, Dayu Jia, Enqing Dong |
CBMS | 1 |