VLDB 2026 Research / reviewers in the wild / expert
Zelin Zang
dblp:226/7615
· DBLP profile ↗
31ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0003-2831-5437ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Departures: Distributional Transport for Single-Cell Perturbation Prediction with Neural Schrödinger BridgesabstractPredicting single-cell perturbation outcomes directly advances gene function analysis and facilitates drug candidate selection, making it a key driver of both basic and translational biomedical research. However, a major bottleneck in this task is the unpaired nature of single-cell data, as the same cell cannot be observed both before and after perturbation due to the destructive nature of sequencing. Although some neural generative transport models attempt to tackle unpaired single-cell perturbation data, they either lack explicit conditioning or depend on prior spaces for indirect distribution alignment, limiting precise perturbation modeling. In this work, we approximate Schrödinger Bridge (SB), which defines stochastic dynamic mappings recovering the entropy-regularized optimal transport (OT), to directly align the distributions of control and perturbed single-cell populations across different perturbation conditions. Unlike prior SB approximations that rely on bidirectional modeling to infer optimal source-target sample coupling, we leverage Minibatch-OT based pairing to avoid such bidirectional inference and the associated ill-posedness of defining the reverse process. This pairing directly guides bridge learning, yielding a scalable approximation to the SB. We approximate two SB models, one modeling discrete gene activation states and the other continuous expression distributions. Joint training enables accurate perturbation modeling and captures single-cell heterogeneity. Experiments on public genetic and drug perturbation datasets show that our model effectively captures heterogeneous single-cell responses and achieves state-of-the-art performance. Changxi Chi, Yufei Huang 0002, Jun Xia 0001, Jiangbin Zheng 0002, Yunfan Liu 0002, Zelin Zang, Stan Z. Li |
AAAI | 6 |
| 2026 | Learning Cell-Aware Hierarchical Multi-Modal Representations for Robust Molecular ModelingabstractUnderstanding how chemical perturbations propagate through biological systems is essential for robust molecular property prediction. While most existing methods focus on chemical structures alone, recent advances highlight the crucial role of cellular responses such as morphology and gene expression in shaping drug effects. However, current cell-aware approaches face two key limitations: (1) modality incompleteness in external biological data, and (2) insufficient modeling of hierarchical dependencies across molecular, cellular, and genomic levels. We propose CHMR (Cell-aware Hierarchical Multi-Modal Representations), a robust framework that jointly models local-global dependencies between molecules and cellular responses and captures latent biological hierarchies via a novel tree-structured vector quantization module. Evaluated on public benchmarks spanning 696 tasks, CHMR outperforms state-of-the-art baselines, yielding average improvements of 3.6% on classification and 17.2% on regression tasks. These results demonstrate the advantage of hierarchy-aware, multi-modal learning for reliable and biologically grounded molecular representations, offering a generalizable framework for integrative biomedical modeling. Mengran Li 0001, Zelin Zang, Wenbin Xing, Junzhou Chen 0001, Jiebo Luo 0001, Stan Z. Li |
AAAI | 2 |
| 2026 | MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language ModelsabstractAnswering complex medical questions requires not only domain expertise and patient-specific information, but also structured and multi-perspective reasoning. Existing multi-agent approaches often rely on fixed roles or shallow interaction prompts, limiting their ability to detect and resolve fine-grained logical inconsistencies. To address this, we propose MedLA, a logic-driven multi-agent framework built on large language models. Each agent organizes its reasoning process into an explicit logical tree based on syllogistic triads (major premise, minor premise, and conclusion), enabling transparent inference and premise-level alignment. Agents engage in a multi-round, graph-guided discussion to compare and iteratively refine their logic trees, achieving consensus through error correction and contradiction resolution. We demonstrate that MedLA consistently outperforms both static role-based systems and single-agent baselines on challenging benchmarks such as MedDDx and standard medical QA tasks. Furthermore, MedLA scales effectively across both open-source and commercial LLM backbones, achieving state-of-the-art performance and offering a generalizable paradigm for trustworthy medical reasoning. Fan Zhang 0010, Jinlin Wu, Guohui Fan, Zelin Zang |
AAAI | 8 |
| 2026 | POSITIVE4Rec: Incorporating the Recency Effect with Positional Inductive Bias for Sequential RecommendationabstractState-of-the-art attention-based models in Sequential Recommendation (SR) often struggle to accurately model the recency effect, where recent interactions disproportionately influence future behavior. Empirically, conventional learnable positional embeddings exhibit erratic fluctuations, failing to guarantee necessary influence decay over time. To address this, we propose POSITIVE4Rec, a model-agnostic framework injecting Positional Inductive Bias (PIB) directly into self-attention. A novel attention reweighting module imposes an adaptive monotonic trend on the attention distribution. Acting as a soft inductive bias rather than a rigid constraint, it prioritizes temporal proximity while retaining flexibility. This plug-and-play mechanism integrates seamlessly into diverse SR architectures. Extensive benchmark experiments show POSITIVE4Rec revitalizes state-of-the-art models with significant performance gains. Code is available at https://anonymous.4open.science/r/POSITIVE4Rec. Po-Chih Lin, Yixuan Dong, Fang-Yi Su, Haijie Yang, Zelin Zang, Hongliang Zhang 0002, Bingo Wing-Kuen Ling, Fuji Yang |
ICMR | 5 |
| 2026 | CellScout: Visual Analytics for Mining Biomarkers in Cell State DiscoveryabstractCell state discovery is crucial for understanding biological systems and enhancing medical outcomes. A key aspect of this process is identifying distinct biomarkers that define specific cell states. However, difficulties arise from the co-discovery process of cell states and biomarkers: biologists often use dimensionality reduction to visualize cells in a two-dimensional space. Then they usually interpret visually clustered cells as distinct states, from which they seek to identify unique biomarkers. However, this assumption is often this assumption often fails to hold due to internal inconsistencies in a cluster, making the process trial-and-error and highly uncertain. Therefore, biologists urgently need effective tools to help uncover the hidden association relationships between different cell populations and their potential biomarkers. To address this problem, we first designed a machine-learning algorithm based on the Mixture-of-Experts (MoE) technique to identify meaningful associations between cell populations and biomarkers. We further developed a visual analytics system-CellScout-in collaboration with biologists, to help them explore and refine these association relationships to advance cell state discovery. We validated our system through expert interviews, from which we further selected a representative case to demonstrate its effectiveness in discovering new cell states. Rui Sheng, Zelin Zang, Jiachen Wang 0001, Zixin Chen, Shaolun Ruan, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | DaCapo: Score Distillation as Stacked Bridge for Fast and High-quality 3D EditingabstractScore Distillation Sampling (SDS) has been successfully extended to text-driven 3D scene editing with 2D pretrained diffusion models. However, SDS-based editing methods suffer from lengthy optimization processes with slow inference and low quality. We attribute the issue of lengthy optimization to the stochastic optimization scheme used in SDS-based editing, where many steps may conflict with each other (e.g., the inherent trade-off between editing and preservation). To reduce this internal conflict and speed up the editing process, we propose to separate editing and preservation in time with a diffusion time schedule and frame the 3D editing optimization process as a diffusion bridge sampling process. Motivated by the analysis above, we introduce DaCapo, a fast diffusion sampling-like 3D editing method that incorporates a novel stacked bridge framework, which estimates a direct diffusion bridge between source and target distribution with only a pretrained 2D diffusion model. Specifically, It models the editing process as a combination of inversion and generation, where both processes happen simultaneously as a stack of Diffusion Bridges. DaCapo shows a 15× speed-up with comparable results to the state-of-the-art SDS-based method. It completes the process in just 2,500 steps on a single GPU and accommodates a variety of 3D representation methods. Yufei Huang 0002, Bangyan Liao, Lirong Wu, Siyuan Li 0002, Cheng Tan 0012, Zicheng Liu 0006, Yunfan Liu 0002, Zelin Zang, Chang Yu 0001, Zhen Lei 0001 |
CVPR | 10 |
| 2025 | USD: Unsupervised Soft Contrastive Learning for Fault Detection in Multivariate Time SeriesabstractUnsupervised fault detection in multivariate time series is critical for maintaining the integrity and efficiency of complex systems, with current methodologies largely focusing on statistical and machine learning techniques. However, these approaches often rest on the assumption that data distributions conform to Gaussian models, overlooking the diversity of patterns that can manifest in both normal and abnormal states, thereby diminishing discriminative performance. Our innovation addresses this limitation by introducing a combination of data augmentation and soft contrastive learning, specifically designed to capture the multifaceted nature of state behaviors more accurately. The data augmentation process enriches the dataset with varied representations of normal states, while soft contrastive learning fine-tunes the model’s sensitivity to the subtle differences between normal and abnormal patterns, enabling it to recognize a broader spectrum of anomalies. This dual strategy significantly boosts the model’s ability to distinguish between normal and abnormal states, leading to a marked improvement in fault detection performance across multiple datasets and settings, thereby setting a new benchmark for unsupervised fault detection in complex systems. Xiuxiu Qiu, Yiming Shi, Zelin Zang |
ICASSP | 4 |
| 2025 | Reconstructing 3D Hand-Instrument Interaction from a Single 2D Image in Medical Scenes
Xiangyu Zhu 0001, Jinlin Wu, Ming Feng, Zelin Zang, Hongbin Liu 0001, Zhen Lei 0001 |
MICCAI (10) | 5 |
| 2025 | FGeneBERT: function-driven pre-trained gene language model for metagenomicsabstractMetagenomic data, comprising mixed multi-species genomes, are prevalent in diverse environments like oceans and soils, significantly impacting human health and ecological functions. However, current research relies on K-mer, which limits the capture of structurally and functionally relevant gene contexts. Moreover, these approaches struggle with encoding biologically meaningful genes and fail to address the one-to-many and many-to-one relationships inherent in metagenomic data. To overcome these challenges, we introduce FGeneBERT, a novel metagenomic pre-trained model that employs a protein-based gene representation as a context-aware and structure-relevant tokenizer. FGeneBERT incorporates masked gene modeling to enhance the understanding of inter-gene contextual relationships and triplet enhanced metagenomic contrastive learning to elucidate gene sequence-function relationships. Pre-trained on over 100 million metagenomic sequences, FGeneBERT demonstrates superior performance on metagenomic datasets at four levels, spanning gene, functional, bacterial, and environmental levels and ranging from 1 to 213 k input sequences. Case studies of ATP synthase and gene operons highlight FGeneBERT's capability for functional recognition and its biological relevance in metagenomic research. Chenrui Duan, Zelin Zang, Yongjie Xu 0001, Hang He, Siyuan Li 0002, Zhen Lei 0001, Ju-Sheng Zheng, Stan Z. Li |
Briefings Bioinform. | 2 |
| 2025 | Complex hierarchical structures analysis in single-cell data with Poincaré deep manifold transformationabstractSingle-cell RNA sequencing (scRNA-seq) offers remarkable insights into cellular development and differentiation by capturing the gene expression profiles of individual cells. The role of dimensionality reduction and visualization in the interpretation of scRNA-seq data has gained widely acceptance. However, current methods face several challenges, including incomplete structure-preserving strategies and high distortion in embeddings, which fail to effectively model complex cell trajectories with multiple branches. To address these issues, we propose the Poincaré deep manifold transformation (PoincaréDMT) method, which maps high-dimensional scRNA-seq data to a hyperbolic Poincaré disk. This approach preserves global structure from a graph Laplacian matrix while achieving local structure correction through a structure module combined with data augmentation. Additionally, PoincaréDMT alleviates batch effects by integrating a batch graph that accounts for batch labels into the low-dimensional embeddings during network training. Furthermore, PoincaréDMT introduces the Shapley additive explanations method based on trained model to identify the important marker genes in specific clusters and cell differentiation process. Therefore, PoincaréDMT provides a unified framework for multiple key tasks essential for scRNA-seq analysis, including trajectory inference, pseudotime inference, batch correction, and marker gene selection. We validate PoincaréDMT through extensive evaluations on both simulated and real scRNA-seq datasets, demonstrating its superior performance in preserving global and local data structures compared to existing methods. Yongjie Xu 0001, Zelin Zang, Bozhen Hu, Cheng Tan 0012, Jun Xia 0001, Stan Z. Li |
Briefings Bioinform. | 2 |
| 2025 | MuST: multiple-modality structure transformation for single-cell spatial transcriptomicsabstractSpatial transcriptomics (ST) technologies have revolutionized the study of gene expression patterns in tissues by providing multimodal data, including transcriptomic (Tra.), spatial, and morphological modalities, thereby offering new opportunities to understand tissue biology beyond traditional Tra. However, we identify the modality bias phenomenon in ST data species, i.e. the inconsistent contribution of different modalities to the labels leads to a tendency for the analysis methods to retain the information of the dominant modality. How to mitigate the adverse effects of modality bias to satisfy various downstream tasks remains a fundamental challenge. This paper introduces Multiple-modality Structure Transformation, named MuST, a novel methodology to tackle the challenge. MuST integrates the multi-modality information contained in the ST data effectively into a uniform latent space to provide a foundation for all the downstream tasks. It learns intrinsic local structures by topology discovery strategy and topology fusion loss function to solve the inconsistencies among different modalities. Thus, these topology-based and deep learning techniques provide a solid foundation for a variety of analytical tasks while coordinating different modalities. The effectiveness of MuST is assessed by performance metrics and biological significance. The results show that it outperforms existing state-of-the-art methods with clear advantages in the precision of identifying and preserving structures of tissues and biomarkers. MuST offers a versatile toolkit for the intricate analysis of complex biological systems. Zelin Zang, Yongjie Xu 0001, Chenrui Duan, Zhen Lei 0001, Stan Z. Li |
Briefings Bioinform. | 1 |
| 2025 | GenURL: A General Framework for Unsupervised Representation LearningabstractUnsupervised representation learning (URL) that learns compact embeddings of high-dimensional data without supervision has achieved remarkable progress recently. However, the development of URLs for different requirements is independent, which limits the generalization of the algorithms, especially prohibitive as the number of tasks grows. For example, dimension reduction (DR) methods, t-SNE and UMAP, optimize pairwise data relationships by preserving the global geometric structure, while self-supervised learning, SimCLR and BYOL, focuses on mining the local statistics of instances under specific augmentations. To address this dilemma, we summarize and propose a unified similarity-based URL framework, GenURL, which can adapt to various URL tasks smoothly. In this article, we regard URL tasks as different implicit constraints on the data geometric structure that help to seek optimal low-dimensional representations that boil down to data structural modeling (DSM) and low-dimensional transformation (LDT). Specifically, DSM provides a structure-based submodule to describe the global structures, and LDT learns compact low-dimensional embeddings with given pretext tasks. Moreover, an objective function, general Kullback-Leibler (GKL) divergence, is proposed to connect DSM and LDT naturally. Comprehensive experiments demonstrate that GenURL achieves consistent state-of-the-art performance in self-supervised visual learning, unsupervised knowledge distillation (KD), graph embeddings (GEs), and DR. Siyuan Li 0002, Zicheng Liu 0006, Zelin Zang, Di Wu 0057, Zhiyuan Chen 0008, Stan Z. Li |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Deep Multimanifold Transformation-Based Multivariate Time Series Fault DetectionabstractUnsupervised fault detection in multivariate time series (MTS) plays a vital role in ensuring the stable operation of complex systems. Traditional methods often assume that normal data follow a single Gaussian distribution and identify anomalies as deviations from this distribution. However, this simplified assumption fails to capture the diversity and structural complexity of real-world time series, which can lead to misjudgments and reduced detection performance in practical applications. To address this issue, we propose a new method that combines a neighborhood-driven data augmentation strategy with a multimanifold representation learning framework. By incorporating information from local neighborhoods, the augmentation module can simulate contextual variations of normal data, enhancing the model's adaptability to distributional changes. In addition, we design a structure-aware feature learning approach that encourages natural clustering of similar patterns in the feature space while maintaining sufficient distinction between different operational states. Extensive experiments on several public benchmark datasets demonstrate that our method achieves superior performance in terms of both accuracy and robustness, showing strong potential for generalization and real-world deployment. Xiuxiu Qiu, Yiming Shi, Zelin Zang, Zhen Lei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | MuST: Maximizing the Latent Capacity of Spatial Transcriptomics Data with Multi-modality Structure TransformationabstractSpatial transcriptomics (ST) technologies have revolutionized the study of gene expression patterns in tissues by providing multimodality data in transcriptomic, spatial, and morphological, offering opportunities for understanding tissue biology beyond transcriptomics. However, we identify the modality bias phenomenon in ST data species, i.e., the inconsistent contribution of different modalities to the labels leads to a tendency for the analysis methods to retain the information of the dominant modality. How to mitigate the adverse effects of modality bias to satisfy various downstream tasks remains a fundamental challenge. This paper introduces Multiple-modality Structure Transformation, named MuST, a novel methodology to tackle the challenge. MuST integrates the multi-modality information contained in the ST data effectively into a uniform latent space to provide a foundation for all the downstream tasks. It learns intrinsic local structures by topology discovery strategy and topology fusion loss function to solve the inconsistencies among different modalities. Thus, these topology-based and deep learning techniques provide a solid foundation for a variety of analytical tasks while coordinating different modalities. The effectiveness of MuST is assessed by performance metrics and biological significance. The results show that it outperforms existing state-of-the-art methods with clear advantages in the precision of identifying and preserving structures of tissues and biomarkers. MuST offers a versatile toolkit for the intricate analysis of complex biological systems. The code is available at https://github.com/zangzelin/code_Must. Zelin Zang, Yongjie Xu 0001, Chenrui Duan, Zhen Lei 0001, Stan Z. Li |
BIBM | 1 |
| 2024 | Deep Manifold Transformation for Protein Representation LearningabstractProtein representation learning is critical in various tasks in biology, such as drug design and protein structure or function prediction, which has primarily benefited from protein language models and graph neural networks. These models can capture intrinsic patterns from protein sequences and structures through masking and task-related losses. However, the learned protein representations are usually not well optimized, leading to performance degradation due to limited data, difficulty adapting to new tasks, etc. To address this, we propose a new deep manifold transformation approach for universal protein representation learning (DMTPRL). It employs manifold learning strategies to improve the quality and adaptability of the learned embeddings. Specifically, we apply a novel manifold learning loss during training based on the graph inter-node similarity. Our proposed DMTPRL method outperforms state-of-the-art baselines on diverse downstream tasks across popular datasets. This validates our approach for learning universal and robust protein representations. We promise to release the code after acceptance. Bozhen Hu, Zelin Zang, Cheng Tan 0012, Stan Z. Li |
ICASSP | 2 |
| 2024 | DiffAug: Enhance Unsupervised Contrastive Learning with Domain-Knowledge-Free Diffusion-based Data AugmentationabstractUnsupervised Contrastive learning has gained prominence in fields such as vision, and biology, leveraging predefined positive/negative samples for representation learning. Data augmentation, categorized into hand-designed and model-based methods, has been identified as a crucial component for enhancing contrastive learning. However, hand-designed methods require human expertise in domain-specific data while sometimes distorting the meaning of the data. In contrast, generative model-based approaches usually require supervised or large-scale external data, which has become a bottleneck constraining model training in many domains. To address the problems presented above, this paper proposes DiffAug, a novel unsupervised contrastive learning technique with diffusion mode-based positive data generation. DiffAug consists of a semantic encoder and a conditional diffusion model; the conditional diffusion model generates new positive samples conditioned on the semantic encoding to serve the training of unsupervised contrast learning. With the help of iterative training of the semantic encoder and diffusion model, DiffAug improves the representation ability in an uninterrupted and unsupervised manner. Experimental evaluations show that DiffAug outperforms hand-designed and SOTA model-based augmentation methods on DNA sequence, visual, and bio-feature datasets. The code for review is released at DiffAug CODE. Zelin Zang, Hao Luo 0004, Kai Wang 0036, Fan Wang 0019, Stan Z. Li, Yang You 0001 |
ICML | 1 |
| 2024 | PhyloGen: Language Model-Enhanced Phylogenetic Inference via Graph Structure GenerationabstractPhylogenetic trees elucidate evolutionary relationships among species, but phylogenetic inference remains challenging due to the complexity of combining continuous (branch lengths) and discrete parameters (tree topology).
Traditional Markov Chain Monte Carlo methods face slow convergence and computational burdens. Existing Variational Inference methods, which require pre-generated topologies and typically treat tree structures and branch lengths independently, may overlook critical sequence features, limiting their accuracy and flexibility.
We propose PhyloGen, a novel method leveraging a pre-trained genomic language model to generate and optimize phylogenetic trees without dependence on evolutionary models or aligned sequence constraints. PhyloGen views phylogenetic inference as a conditionally constrained tree structure generation problem, jointly optimizing tree topology and branch lengths through three core modules: (i) Feature Extraction, (ii) PhyloTree Construction, and (iii) PhyloTree Structure Modeling.
Meanwhile, we introduce a Scoring Function to guide the model towards a more stable gradient descent.
We demonstrate the effectiveness and robustness of PhyloGen on eight real-world benchmark datasets. Visualization results confirm PhyloGen provides deeper insights into phylogenetic relationships. Chenrui Duan, Zelin Zang, Siyuan Li 0002, Yongjie Xu 0001, Stan Z. Li |
NeurIPS | 2 |
| 2024 | Deep Graph Neural Networks via Posteriori-Sampling-based Node-Adaptative Residual ModuleabstractGraph Neural Networks (GNNs), a type of neural network that can learn from graph-structured data through neighborhood information aggregation, have shown superior performance in various downstream tasks. However, as the number of layers increases, node representations becomes indistinguishable, which is known as over-smoothing. To address this issue, many residual methods have emerged. In this paper, we focus on the over-smoothing issue and related residual methods. Firstly, we revisit over-smoothing from the perspective of overlapping neighborhood subgraphs, and based on this, we explain how residual methods can alleviate over-smoothing by integrating multiple orders neighborhood subgraphs to avoid the indistinguishability of the single high-order neighborhood subgraphs. Additionally, we reveal the drawbacks of previous residual methods, such as the lack of node adaptability and severe loss of high-order neighborhood subgraph information, and propose a \textbf{Posterior-Sampling-based, Node-Adaptive Residual module (PSNR)}. We theoretically demonstrate that PSNR can alleviate the drawbacks of previous residual methods. Furthermore, extensive experiments verify the superiority of the PSNR module in fully observed node classification and missing feature scenarios. Our code
is available at \href{https://github.com/jingbo02/PSNR-GNN}{https://github.com/jingbo02/PSNR-GNN}. Ruqiong Zhang, Jun Xia 0001, Zhizhi Yu, Zelin Zang, Di Jin 0001, Carl Yang 0001, Stan Z. Li |
NeurIPS | 6 |
| 2024 | DMT-EV: An Explainable Deep Network for Dimension ReductionabstractDimension reduction (DR) is commonly utilized to capture the intrinsic structure and transform high-dimensional data into low-dimensional space while retaining meaningful properties of the original data. It is used in various applications, such as image recognition, single-cell sequencing analysis, and biomarker discovery. However, contemporary parametric-free and parametric DR techniques suffer from several significant shortcomings, such as the inability to preserve global and local features and the poor generalisation performance. On the other hand, regarding explainability, it is crucial to comprehend the embedding process, especially the contribution of each part to the embedding process, while understanding how each feature affects the embedding results that identify critical components and help diagnose the embedding process. To address these problems, we have developed a deep neural network method called DMT-EV, which provides not only excellent performance in structural maintainability but also explainability to the DR therein. DMT-EV starts with data augmentation and a manifold-based loss function to improve embedding performance. The explanation is based on saliency maps and aims to examine the trained DMT-EV parameters and contributions of components during the embedding process. The proposed techniques are integrated with a visual interface to help the user to adjust DMT-EV to achieve better DR performance and explainability. The interactive visual interface makes it easier to illustrate the data features, compare different DR techniques, and investigate DR. An in-depth experimental comparison shows that DMT-EV consistently outperforms the state-of-the-art methods in both performance measures and explainability. Zelin Zang, Shenghui Cheng, Hanchen Xia, Yaoting Sun, Yongjie Xu 0001, Baigui Sun, Stan Z. Li |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Deep Manifold Graph Auto-Encoder For Attributed Graph EmbeddingabstractRepresenting graph data in a low-dimensional space for subsequent tasks is the purpose of attributed graph embedding. Most existing neural network approaches learn latent representations by minimizing reconstruction errors. Rare work considers the data distribution and the topological structure of latent codes simultaneously, which often results in inferior embeddings in real-world graph data. This paper proposes a novel Deep Manifold (Variational) Graph Auto-Encoder (DMVGAE/DMGAE) method for attributed graph data to improve the stability and quality of learned representations to tackle the crowding problem. The node-to-node geodesic similarity is preserved between the original and latent space under a pre-defined distribution. The proposed method surpasses state-of-the-art baseline algorithms by a significant margin on different downstream tasks across popular datasets, which validates our solutions. We promise to release the code after acceptance. Bozhen Hu, Zelin Zang, Jun Xia 0001, Lirong Wu, Cheng Tan 0012, Stan Z. Li |
ICASSP | 2 |
| 2023 | Boosting Novel Category Discovery Over Domains with Soft Contrastive Learning and All in One ClassifierabstractUnsupervised domain adaptation (UDA) has proven to be highly effective in transferring knowledge from a label-rich source domain to a label-scarce target domain. However, the presence of additional novel categories in the target domain has led to the development of open-set domain adaptation (ODA) and universal domain adaptation (UNDA). Existing ODA and UNDA methods treat all novel categories as a single, unified unknown class and attempt to detect it during training. However, we found that domain variance can lead to more significant view-noise in unsupervised data augmentation, which affects the effectiveness of contrastive learning (CL) and causes the model to be overconfident in novel category discovery. To address these issues, a framework named Soft-contrastive All-in-one Network (SAN) is proposed for ODA and UNDA tasks. SAN includes a novel data-augmentation-based soft contrastive learning (SCL) loss to fine-tune the backbone for feature transfer and a more human-intuitive classifier to improve new class discovery capability. The SCL loss weakens the adverse effects of the data augmentation view-noise problem which is amplified in domain transfer tasks. The All-in-One (AIO) classifier overcomes the overconfidence problem of current mainstream closed-set and open-set classifiers. Visualization and ablation experiments demonstrate the effectiveness of the proposed innovations. Furthermore, extensive experiment results on ODA and UNDA show that SAN outperforms existing state-of-the-art methods. Zelin Zang, Senqiao Yang, Fei Wang 0032, Baigui Sun, Xuansong Xie, Stan Z. Li |
ICCV | 1 |
| 2023 | Architecture-Agnostic Masked Image Modeling - From ViT back to CNNabstractMasked image modeling, an emerging self-supervised pre-training method, has shown impressive success across numerous downstream vision tasks with Vision transformers. Its underlying idea is simple: a portion of the input image is masked out and then reconstructed via a pre-text task. However, the working principle behind MIM is not well explained, and previous studies insist that MIM primarily works for the Transformer family but is incompatible with CNNs. In this work, we observe that MIM essentially teaches the model to learn better middle-order interactions among patches for more generalized feature extraction. We then propose an Architecture-Agnostic Masked Image Modeling framework (A$^2$MIM), which is compatible with both Transformers and CNNs in a unified way. Extensive experiments on popular benchmarks show that A$^2$MIM learns better representations without explicit design and endows the backbone model with the stronger capability to transfer to various downstream tasks. Siyuan Li 0002, Di Wu 0057, Fang Wu 0002, Zelin Zang, Stan Z. Li |
ICML | 4 |
| 2023 | UDRN: Unified Dimensional Reduction Neural Network for feature selection and feature projection
Zelin Zang, Yongjie Xu 0001, Linyan Lu, Yulan Geng, Senqiao Yang, Stan Z. Li |
Neural Networks | 1 |
| 2022 | Exploring Localization for Self-supervised Fine-grained Contrastive Learning
Di Wu 0057, Siyuan Li 0002, Zelin Zang, Stan Z. Li |
BMVC | 3 |
| 2022 | DLME: Deep Local-Flatness Manifold Embedding
Zelin Zang, Siyuan Li 0002, Di Wu 0057, Kai Wang 0036, Baigui Sun, Hao Li 0030, Stan Z. Li |
ECCV (21) | 1 |
| 2022 | Generalized Clustering and Multi-Manifold Learning with Geometric Structure PreservationabstractThough manifold-based clustering has become a popular research topic, we observe that one important factor has been omitted by these works, namely that the defined clustering loss may corrupt the local and global structure of the latent space. In this paper, we propose a novel Generalized Clustering and Multi-manifold Learning (GCML) framework with geometric structure preservation for generalized data, i.e., not limited to 2-D image data and has a wide range of applications in speech, text, and biology domains. In the proposed framework, manifold clustering is done in the latent space guided by a clustering loss. To overcome the problem that the clustering-oriented loss may deteriorate the geometric structure of the latent space, an isometric loss is proposed for preserving intra-manifold structure locally and a ranking loss for inter-manifold structure globally. Extensive experimental results have shown that GCML exhibits superior performance to counterparts in terms of qualitative visualizations and quantitative metrics, which demonstrates the effectiveness of preserving geometric structure. Code has been made available at: https://github.com/LirongWu/GCML. Lirong Wu, Zicheng Liu 0006, Jun Xia 0001, Zelin Zang, Siyuan Li 0002, Stan Z. Li |
WACV | 4 |
| 2022 | Surrogate Representation Learning with Isometric Mapping for Gray-box Graph Adversarial AttacksabstractGray-box graph attacks aim to disrupt the victim model's performance by using inconspicuous attacks with limited knowledge of the victim model. The details of the victim model and the labels of the test nodes are invisible to the attacker. The attacker constructs an imaginary surrogate model trained under supervision to obtain the gradient on the node attributes or graph structure. However, there is a lack of discussion on the training of surrogate models and the reliability of provided gradient information. The general node classification models lose the topology of the nodes on the graph, which is, in fact, an exploitable prior for the attacker. This paper investigates the effect of surrogate representation learning on the transferability of gray-box graph adversarial attacks. We propose Surrogate Representation Learning with Isometric Mapping (SRLIM) to reserve the topology in the surrogate embedding. By isometric mapping, our proposed SRLIM can constrain the topological structure of nodes from the input layer to the embedding space, that is, to maintain the similarity of nodes in the propagation process. Experiments prove the effectiveness of our approach through the improvement in the performance of the adversarial attacks generated by the gradient-based attacker in untargeted poisoning gray-box scenarios. Zelin Zang, Stan Z. Li |
WSDM | 3 |
| 2022 | Deep manifold embedding of attributed graphs
Zelin Zang, Siyuan Li 0002, Di Wu 0057, Jianzhu Guo, Yongjie Xu 0001, Stan Z. Li |
Neurocomputing | 1 |
| 2021 | Invertible Manifold Learning for Dimension Reduction
Siyuan Li 0002, Zelin Zang, Lirong Wu, Jun Xia 0001, Stan Z. Li |
ECML/PKDD (3) | 3 |
| 2020 | Two-dimensional discrete feature based spatial attention CapsNet For sEMG signal recognition
Guoqi Chen, Wanliang Wang, Zheng Wang 0048, Honghai Liu 0001, Zelin Zang, Weikun Li |
Appl. Intell. | 5 |
| 2018 | Fish Swarm Based Man-Machine Cooperative Photographing Location Positioning Algorithm
Zelin Zang, Wanliang Wang, Linyan Lu, Yanwei Zhao |
CDVE | 1 |