Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sicheng Yang 0001

dblp:176/6714-1 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0008-5201-4494ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 67% Representation and self-supervised learning · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
image tokenization
1.012026
VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026
Machine learning › Generative modeling
variational autoencoder
1.012026
VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026
Machine learning › Representation and self-supervised learning
vector quantization
1.012026
VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026

Methods — techniques the papers use, named apart from their topics

vector quantization · 1.0variational autoencoder · 1.0distribution regularization · 1.0
YearPublicationVenuePosition
2026 VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling
abstract
Vector quantization (VQ) transforms continuous image features into discrete representations, providing compressed, tokenized inputs for generative models. However, VQ-based frameworks suffer from several issues, such as non-smooth latent spaces, weak alignment between representations before and after quantization, and poor coherence between the continuous and discrete domains. These issues lead to unstable codeword learning and underutilized codebooks, ultimately degrading the performance of both reconstruction and downstream generation tasks. To this end, we propose VAEVQ, which comprises three key components: (1) Variational Latent Quantization (VLQ), replacing the AE with a VAE for quantization to leverage its structured and smooth latent space, thereby facilitating more effective codeword activation; (2) Representation Coherence Strategy (RCS), adaptively modulating the alignment strength between pre- and post-quantization features to enhance consistency and prevent overfitting to noise; and (3) Distribution Consistency Regularization (DCR), aligning the entire codebook distribution with the continuous latent distribution to improve utilization. Extensive experiments on two benchmark datasets demonstrate that VAEVQ outperforms state-of-the-art methods.
Sicheng Yang 0001, Xing Hu 0010, Qiang Wu 0012
AAAI1
2026 LCM-Net: LLM-Driven Cross-Modality MoE Feature Fusion Network for Cancer Survival Analysis
abstract
Cancer survival analysis aims to predict survival outcomes to evaluate the efficacy and prognosis of treatment. Although current approaches have designed diverse cross-modal learning methods to integrate genetic data and pathology images, they are frequently hindered by data redundancy. Pattern representation in high-dimensional genetic data remains a significant hurdle. Pathology data analysis is computationally intensive because of the giga-pixel resolution. Moreover, the heterogeneity of data types poses a barrier to extending multimodal fusion methods. To address the aforementioned issues, we propose a novel LLM-driven Cross-Modality MoE-feature Fusion Network (LCM-Net) with three innovative modules for boosting cancer survival prediction. Specifically, the Genomic Language Alignment (GLA) module integrates genomic features with learnable prompts. Utilizing large language models, it encodes genomic information into concise and semantically relevant representations. Then, we devise the Pathological Feature Refinement (PFR) module to serve as a plug-and-play component that filters out irrelevant regions in pathology images. Finally, we propose a Multimodal Expert Integration (MEI) module to effectively leverage the capabilities of different experts, integrating the processed features from both the genomic and pathological domains. Extensive experiments on five public datasets demonstrate that our approach outperforms state-of-the-art methods, and the ablation study confirms the effectiveness of the proposed modules. Our code is publicly available at https://github.com/script-Yang/LCM-Net.
Sicheng Yang 0001, Haipeng Zhou, Weiming Wang 0002, Shifu Chen, Guang Yang 0006, Huazhu Fu, Lei Zhu 0003
IEEE Trans. Medical Imaging1
2025 CoC: Chain-of-Cancer Based on Cross-Modal Autoregressive Traction for Survival Prediction
Haipeng Zhou, Sicheng Yang 0001, Harry Qin, Lei Zhu 0003
MICCAI (15)2
2025 VQ-Seg: Vector-Quantized Token Perturbation for Semi-Supervised Medical Image Segmentation
abstract
Consistency learning with feature perturbation is a widely used strategy in semi-supervised medical image segmentation. However, many existing perturbation methods rely on dropout, and thus require a careful manual tuning of the dropout rate, which is a sensitive hyperparameter and often difficult to optimize and may lead to suboptimal regularization. To overcome this limitation, we propose VQ-Seg, the first approach to employ vector quantization (VQ) to discretize the feature space and introduce a novel and controllable Quantized Perturbation Module (QPM) that replaces dropout. Our QPM perturbs discrete representations by shuffling the spatial locations of codebook indices, enabling effective and controllable regularization. To mitigate potential information loss caused by quantization, we design a dual-branch architecture where the post-quantization feature space is shared by both image reconstruction and segmentation tasks. Moreover, we introduce a Post-VQ Feature Adapter (PFA) to incorporate guidance from a foundation model (FM), supplementing the high-level semantic information lost during quantization. Furthermore, we collect a large-scale Lung Cancer (LC) dataset comprising 828 CT scans annotated for central-type lung carcinoma. Extensive experiments on the LC dataset and other public benchmarks demonstrate the effectiveness of our method, which outperforms state-of-the-art approaches. Codes will be released.
Sicheng Yang 0001, Zhaohu Xing, Lei Zhu 0003
NeurIPS1
2024 Cross-conditioned Diffusion Model for Medical Image to Image Translation
Zhaohu Xing, Sicheng Yang 0001, Sixiang Chen, Tian Ye 0001, Harry Qin, Lei Zhu 0003
MICCAI (7)2