Zhiyang He

dblp:24/10650 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Theory of computation · 3 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment via Explicit Injection
abstract
He Geng, Yangmin Huang, Lixian Lai, Qianyun Du, Hui Chu, Zhiyang He, Jiaxue Hu, Xiaodong Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
He Geng, Yangmin Huang, Lixian Lai, Qianyun Du, Hui Chu, Zhiyang He, Jiaxue Hu, Xiaodong Tao
ACL (1)6
2026 Distilling Magic States in the Bicycle Architecture
Shifan Xu, Patrick Rall, Zhiyang He, Yongshan Ding 0001
ISCA4
2025 KANTrust: A Multi-Omics Framework for Uncertainty-Aware Disease Subtyping
abstract
The integration of multi-omics data, including DNA methylation, mRNA expression, and miRNA profiles, is crucial for accurate disease subtyping and outcome prediction in complex disorders such as Alzheimer's disease and various cancers. However, the inherent heterogeneity and inconsistency among omics views present significant challenges for reliable data fusion. To address these issues, we propose KANTrust, a novel framework for trustworthy multi-omics classification that explicitly models both epistemic and aleatoric uncertainties. Our method combines a Kolmogorov-Arnold Network (KAN)enhanced robust representation module, a contrastive evidence consistency module, and an evidence-theoretic fusion module to achieve reliable multi-view integration. KANTrust adaptively highlights informative features within each omics modality, promotes semantic alignment across views, and quantifies uncertainty through a Dempster-Shafer framework. Experimental evaluations on four real-world biomedical datasets demonstrate that KANTrust consistently outperforms state-of-the-art methods in both binary and multi-class classification tasks. Code is available at https://github.com/wcj6/KANTrust.
Chunjiang Wang, Rui Yan 0009, Kun Zhang 0040, Zihang Jiang, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou
BIBM5
2025 MVP-CBM: Multi-layer Visual Preference-enhanced Concept Bottleneck Model for Explainable Medical Image Classification
abstract
The concept bottleneck model (CBM), as a technique improving interpretability via linking predictions to human-understandable concepts, makes high-risk and life-critical medical image classification credible. Typically, existing CBM methods associate the final layer of visual encoders with concepts to explain the model’s predictions. However, we empirically discover the phenomenon of concept preference variation, that is, the concepts are preferably associated with the features at different layers than those only at the final layer; yet a blind last-layer-based association neglects such a preference variation and thus weakens the accurate correspondences between features and concepts, impairing model interpretability. To address this issue, we propose a novel Multi-layer Visual Preference-enhanced Concept Bottleneck Model (MVP-CBM), which comprises two key novel modules: (1) intra-layer concept preference modeling, which captures the preferred association of different concepts with features at various visual layers, and (2) multi-layer concept sparse activation fusion, which sparsely aggregates concept activations from multiple layers to enhance performance. Thus, by explicitly modeling concept preferences, MVP-CBM can comprehensively leverage multi-layer visual information to provide a more nuanced and accurate explanation of model decisions. Extensive experiments on several public medical classification benchmarks demonstrate that MVP-CBM achieves state-of-the-art accuracy and interoperability, verifying its superiority. Code is available at https://github.com/wcj6/MVP-CBM.
Chunjiang Wang, Kun Zhang 0040, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou
IJCAI4
2025 Pre-trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
Fenghe Tang, Zhiyang He, Xiaodong Tao, Zihang Jiang, Shaohua Kevin Zhou
MICCAI (10)3
2025 SimCroP: Radiograph Representation Learning with Similarity-Driven Cross-Granularity Pre-training
Rongsheng Wang 0003, Fenghe Tang, Qingsong Yao, Rui Yan 0009, Zhen Huang 0007, Haoran Lai, Zhiyang He, Xiaodong Tao, Zihang Jiang, Shaohua Kevin Zhou
MICCAI (5)8
2025 ECAMP: Entity-centered Context-aware Medical Vision Language Pre-training
Rongsheng Wang 0003, Qingsong Yao, Zihang Jiang, Haoran Lai, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou
Medical Image Anal.5
2024 CARZero: Cross-Attention Alignment for Radiology Zero-Shot Classification
abstract
The advancement of Zero-Shot Learning in the medi-cal domain has been driven forward by using pretrained models on large-scale image-text pairs, focusing on image-text alignment. However, existing methods primarily rely on cosine similarity for alignment, which may not fully capture the complex relationship between medical images and reports. To address this gap, we introduce a novel approach called Cross-Attention Alignment for Radiology Zero-Shot Classification (CARZero). Our approach innovatively leverages cross-attention mechanisms to process image and report features, creating a Similarity Representation that more accurately reflects the intricate relationships in medical semantics. This representation is then linearly projected to form an image-text similarity matrix for cross-modality alignment. Additionally, recognizing the pivotal role of prompt selection in zero-shot learning, CARZero in-corporates a Large Language Model-based prompt alignment strategy. This strategy standardizes diverse diagnostic expressions into a unified format for both training and inference phases, overcoming the challenges of manual prompt design. Our approach is simple yet effective, demonstrating state-of-the-art performance in zero-shot classification on five official chest radiograph diagnostic test sets, including remarkable results on datasets with long-tail distributions of rare diseases. This achievement is attributed to our new image-text alignment strategy, which effectively addresses the complex relationship between medical images and reports. Code and models are available at https://github.com/laihaoran/CARZero.
Haoran Lai, Qingsong Yao, Zihang Jiang, Rongsheng Wang 0003, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou
CVPR5
2022 Breaking the nk barrier for minimum k-cut on simple graphs
abstract
In the minimum k-cut problem, we want to find the minimum number of edges whose deletion breaks the input graph into at least k connected components. The classic algorithm of Karger and Stein runs in Õ(n2k−2) time, and recent, exciting developments have improved the running time to O(nk). For general, weighted graphs, this is tight assuming popular hardness conjectures.
Zhiyang He, Jason Li 0006
STOC1
2021 Near-Linear-Time, Optimal Vertex Cut Sparsifiers in Directed Acyclic Graphs
abstract
Let $G$ be a graph and $S, T \subseteq V(G)$ be (possibly overlapping) sets of terminals, $|S|=|T|=k$. We are interested in computing a vertex sparsifier for terminal cuts in $G$, i.e., a graph $H$ on a smallest possible number of vertices, where $S \cup T \subseteq V(H)$ and such that for every $A \subseteq S$ and $B \subseteq T$ the size of a minimum $(A,B)$-vertex cut is the same in $G$ as in $H$. We assume that our graphs are unweighted and that terminals may be part of the min-cut. In previous work, Kratsch and Wahlström (FOCS 2012/JACM 2020) used connections to matroid theory to show that a vertex sparsifier $H$ with $O(k^3)$ vertices can be computed in randomized polynomial time, even for arbitrary digraphs $G$. However, since then, no improvements on the size $O(k^3)$ have been shown. In this paper, we draw inspiration from the renowned Bollobás's Two-Families Theorem in extremal combinatorics and introduce the use of total orderings into Kratsch and Wahlström's methods. This new perspective allows us to construct a sparsifier $H$ of $Θ(k^2)$ vertices for the case that $G$ is a DAG. We also show how to compute $H$ in time near-linear in the size of $G$, improving on the previous $O(n^{ω+1})$. Furthermore, $H$ recovers the closest min-cut in $G$ for every partition $(A,B)$, which was not previously known. Finally, we show that a sparsifier of size $Ω(k^2)$ is required, both for DAGs and for undirected edge cuts.
Zhiyang He, Jason Li 0006, Magnus Wahlström
ESA1
2019 Hypergraphs with Few Berge Paths of Fixed Length between Vertices
abstract
In this paper we study the maximum number of hyperedges which may be in an $r$-uniform hypergraph under the restriction that no pair of vertices has more than $t$ Berge paths of length $k$ between them. When $r=t=2$, this is the even-cycle problem asking for ${ex}(n, C_{2k})$. We extend results of Füredi and Simonovits and of Conlon, who studied the problem when $r=2$. In particular, we show that for fixed $k$ and $r$, there is a constant $t$ such that the maximum number of edges can be determined in order of magnitude.
Zhiyang He, Michael Tait
SIAM J. Discret. Math.1
2018 Medical Exam Question Answering with Large-scale Reading Comprehension
abstract
Reading and understanding text is one important component in computer aided diagnosis in clinical medicine, also being a major research problem in the field of NLP. In this work, we introduce a question-answering task called MedQA to study answering questions in clinical medicine using knowledge in a large-scale document collection. The aim of MedQA is to answer real-world questions with large-scale reading comprehension. We propose our solution SeaReader---a modular end-to-end reading comprehension model based on LSTM networks and dual-path attention architecture. The novel dual-path attention models information flow from two perspectives and has the ability to simultaneously read individual documents and integrate information across multiple documents. In experiments our SeaReader achieved a large increase in accuracy on MedQA over competing models. Additionally, we develop a series of novel techniques to demonstrate the interpretation of the question answering process in SeaReader.
Xiao Zhang 0001, Ji Wu 0002, Zhiyang He, Xien Liu
AAAI3
2017 Multi-label text classification based on the label correlation mixture model
abstract
In the current paper, we propose a probabilistic generative model, the label correlation mixture model (LCMM), to depict multi-labeled document data, which can be utilized for multi-label text classification. LCMM assumes two stochastic generative processes, which correspond to two submodels: 1) a label correlation model; and 2) a label mixture model. The former model formulates labels’ generative process, in which a label correlation network is created to depict the dependency between labels. Moreover, we present an efficient inference algorithm for calculating the generative probability of a multi-label class. Furthermore, in order to optimize the label correlation network, we propose a parameter-learning algorithm based on gradient descent. The second submodel in the LCMM depicts the generative process of words in a document with the given labels. Different traditional mixture models can be adopted in this generative process, such as the mixture of language models, or topic models. In the multi-label classification stage, we propose a two-step strategy to most efficiently utilize the LCMM based on the framework of Bayes decision theory. We conduct extensive multi-label classification experiments on three standard text data sets. The experimental results show significant performance improvements comparing to existing approaches. For example, the improvements on accuracy and macro F-score measures in the OHSUMED data set achieve 28.3% and 37.0%, respectively. These performance enhancements demonstrate the effectiveness of the proposed models and solutions.
Zhiyang He, Ji Wu 0002, Ping Lv
Intell. Data Anal.1
2016 Hidden Softmax Sequence Model for Dialogue Structure Analysis
abstract
We propose a new unsupervised learning model, hidden softmax sequence model (HSSM), based on Boltzmann machine for dialogue structure analysis.The model employs three types of units in the hidden layer to discovery dialogue latent structures: softmax units which represent latent states of utterances; binary units which represent latent topics specified by dialogues; and a binary unit that represents the global general topic shared across the whole dialogue corpus.In addition, the model contains extra connections between adjacent hidden softmax units to formulate the dependency between latent states.Two different kinds of real world dialogue corpora, Twitter-Post and AirTicketBooking, are utilized for extensive comparing experiments, and the results illustrate that the proposed model outperforms sate-ofthe-art popular approaches.
Zhiyang He, Xien Liu, Ping Lv, Ji Wu 0002
ACL (1)1
2016 Target-Based State and Tracking Algorithm for Spoken Dialogue System
Miao Li 0003, Zhiyang He, Ji Wu 0002
INTERSPEECH2
2014 Label correlation mixture model for multi-label text categorization
abstract
Multi-label text categorization is more difficult but practical than the conventional binary or multi-class text categorization. This paper propose a novel probabilistic generative model, label correlation mixture model (LCMM), to depict the multiple labeled documents, which can be used for multi-label text categorization. In LCMM, labels and topics have the one-to-one correspondences. LCMM consists of two parts: label correlation model and multi-label conditioned document model. The former one formulates the generating process of labels and the dependencies between the labels are taken into account. We also propose an efficient algorithm for calculating the probability of generating an arbitrary subset of labels. Multi-label conditioned document model can be regarded as a supervised label mixture model, in which the labels for a document are known. To evaluate LCMM, multi-label text categorization experiments on three standard text data sets are performed. The experimental results demonstrate the effectiveness of LCMM, comparing to other reported methods.
Zhiyang He, Ji Wu 0002, Ping Lv
SLT1
2011 An Active Learning Approach to Task Adaptation
Ji Wu 0002, Zhiyang He, Ping Lv
INTERSPEECH2