EDBT 2026 Demo / reviewers in the wild / expert
Zhiquan Tan
dblp:326/0177
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0001-6658-1934ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Representation and self-supervised learning · 40% Learning theory · 16% Trustworthy machine learning · 12% | |
| Theoretical computer science
1 paper |
Coding theory · 100% |
Topics — the 26 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
2.5 | 4 | 2024 | Matrix Information Theory for Self-Supervised Learning · ICML 2024 Provable Contrastive Continual Learning · ICML 2024 Contrastive Learning is Spectral Clustering on Similarity Graph · ICLR 2024 |
Machine learning › Learning theory
information-theoretic analysis |
1.5 | 2 | 2024 | Information Flow in Self-Supervised Learning · ICML 2024 Unveiling the Dynamics of Information Interplay in Supervised Learning · ICML 2024 |
Machine learning › Representation and self-supervised learning
representation analysis |
1.5 | 2 | 2024 | Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models · NeurIPS 2024 Unveiling the Dynamics of Information Interplay in Supervised Learning · ICML 2024 |
Computer vision › Segmentation and scene understanding › category discovery
generalized category discovery |
0.9 | 1 | 2025 | Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution Regularization · ICML 2025 |
Coding theory › error-correcting codes › erasure coding
burst erasure correction |
0.9 | 1 | 2025 | Multiplexed Streaming Codes for Messages With Different Decoding Delays in Channel With Burst and Random Erasures · IEEE Trans. Commun. 2025 |
Coding theory › error-correcting codes
erasure coding |
0.9 | 1 | 2025 | Multiplexed Streaming Codes for Messages With Different Decoding Delays in Channel With Burst and Random Erasures · IEEE Trans. Commun. 2025 |
Coding theory › error-correcting codes › erasure coding
streaming codes |
0.9 | 1 | 2025 | Multiplexed Streaming Codes for Messages With Different Decoding Delays in Channel With Burst and Random Erasures · IEEE Trans. Commun. 2025 |
Machine learning › Learning paradigms
continual learning |
0.8 | 1 | 2024 | Provable Contrastive Continual Learning · ICML 2024 |
Machine learning › Learning theory
generalization bounds |
0.8 | 1 | 2024 | Provable Contrastive Continual Learning · ICML 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.8 | 1 | 2024 | Provable Contrastive Continual Learning · ICML 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
maximum entropy coding |
0.8 | 1 | 2024 | Matrix Information Theory for Self-Supervised Learning · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
non-contrastive learning |
0.8 | 1 | 2024 | Matrix Information Theory for Self-Supervised Learning · ICML 2024 |
Machine learning › Transfer learning and domain adaptation
optimal transport alignment |
0.8 | 1 | 2024 | OTMatch: Improving Semi-Supervised Learning with Optimal Transport · ICML 2024 |
Machine learning › Learning theory › generalization bounds
performance guarantees |
0.8 | 1 | 2024 | Provable Contrastive Continual Learning · ICML 2024 |
Machine learning › Learning paradigms
semi-supervised learning |
0.8 | 1 | 2024 | OTMatch: Improving Semi-Supervised Learning with Optimal Transport · ICML 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
theoretical analysis of contrastive learning |
0.8 | 1 | 2024 | Contrastive Learning is Spectral Clustering on Similarity Graph · ICLR 2024 |
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution |
0.7 | 1 | 2023 | Trade-off Between Efficiency and Consistency for Removal-based Explanations · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Trade-off Between Efficiency and Consistency for Removal-based Explanations · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › interpretability › post-hoc explanation
removal-based explanation |
0.7 | 1 | 2023 | Trade-off Between Efficiency and Consistency for Removal-based Explanations · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › contrastive learning › negative-free contrastive learning
barlow twins |
0.2 | 1 | 2024 | Information Flow in Self-Supervised Learning · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder |
0.2 | 1 | 2024 | Information Flow in Self-Supervised Learning · ICML 2024 |
Machine learning › Deep learning architectures and training
neural collapse |
0.2 | 1 | 2024 | Unveiling the Dynamics of Information Interplay in Supervised Learning · ICML 2024 |
Machine learning › Deep learning architectures and training
training dynamics |
0.2 | 1 | 2024 | Unveiling the Dynamics of Information Interplay in Supervised Learning · ICML 2024 |
Computer vision › Vision and language › vision-language model
vision-language model alignment |
0.2 | 1 | 2024 | Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability › shapley value
SHAP |
0.2 | 1 | 2023 | Trade-off Between Efficiency and Consistency for Removal-based Explanations · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
matrix information theory · 1.5sliding window analysis · 0.9pseudo-labeling · 0.9merging approach · 0.9contrastive learning · 0.9mutual information ratio · 0.8matrix mutual information · 0.8matrix joint entropy · 0.8masked autoencoder · 0.8kernel function · 0.8entropy difference ratio · 0.8InfoNCE loss · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution RegularizationabstractGeneralized Category Discovery (GCD) aims to identify unlabeled samples by leveraging the base knowledge from labeled ones, where the unlabeled set consists of both base and novel classes.
Since clustering methods are time-consuming at inference, parametric-based approaches have become more popular.
However, recent parametric-based methods suffer from inferior base discrimination due to unreliable self-supervision.
To address this issue, we propose a Reciprocal Learning Framework (RLF) that introduces an auxiliary branch devoted to base classification.
During training, the main branch filters the pseudo-base samples to the auxiliary branch.
In response, the auxiliary branch provides more reliable soft labels for the main branch, leading to a virtuous cycle.
Furthermore, we introduce Class-wise Distribution Regularization (CDR) to mitigate the learning bias towards base classes.
CDR essentially increases the prediction confidence of the unlabeled data and boosts the novel class performance.
Combined with both components, our proposed method, RLCD, achieves superior performance in all classes with negligible extra computation.
Comprehensive experiments across seven GCD datasets validate its superiority.
Our codes are available at https://github.com/APORduo/RLCD. Zhiquan Tan, Linglan Zhao, Xiangzhong Fang, Weiran Huang 0001 |
ICML | 2 |
| 2025 | Multiplexed Streaming Codes for Messages With Different Decoding Delays in Channel With Burst and Random ErasuresabstractIn a real-time transmission scenario, messages are transmitted through a channel that is subject to packet loss. The destination must recover the messages within the required deadline. In this paper, we consider a setup where two different types of messages with distinct decoding deadlines are transmitted through a channel model that introduces either one burst erasure of length at most B, or N random erasures in any fixed-sized sliding window. The message with a short decoding deadline$T_{\mathrm {u}}$is referred to as an urgent message, while the other one with a decoding deadline$T_{\mathrm {v}}$($T_{\mathrm {v}} \gt T_{\mathrm {u}}$) is referred to as a less urgent message. We consider the scenario where$T_{\mathrm {v}} \gt T_{\mathrm {u}} + B$and propose a non-trivial achievable region$\mathcal {R}$for the aforementioned channel model. We propose a novel merging approach to encode two message streams of different urgency levels into a single flow and present explicit constructions for encoding, contributing to the establishment of the achievability of region$\mathcal {R}$. Our comprehensive analysis demonstrates that this region encompasses the rate pairs of existing encoding schemes and coincides with the capacity region in burst channel scenarios. Lastly, we investigate the property of the achievable region$\mathcal {R}$, proving that it is the largest one obtained from all the rate pairs under the merging method. Dingli Yuan, Zhiquan Tan |
IEEE Trans. Commun. | 2 |
| 2024 | Contrastive Learning is Spectral Clustering on Similarity GraphabstractContrastive learning is a powerful self-supervised learning method, but we have a limited theoretical understanding of how it works and why it works. In this paper, we prove that contrastive learning with the standard InfoNCE loss is equivalent to spectral clustering on the similarity graph. Using this equivalence as the building block, we extend our analysis to the CLIP model and rigorously characterize how similar multi-modal objects are embedded together. Motivated by our theoretical insights, we introduce the Kernel-InfoNCE loss, incorporating mixtures of kernel functions that outperform the standard Gaussian kernel on several vision datasets. Zhiquan Tan, Yifan Zhang 0029, Jingqin Yang, Yang Yuan 0010 |
ICLR | 1 |
| 2024 | Unveiling the Dynamics of Information Interplay in Supervised LearningabstractIn this paper, we use matrix information theory as an analytical tool to analyze the dynamics of the information interplay between data representations and classification head vectors in the supervised learning process. Specifically, inspired by the theory of Neural Collapse, we introduce matrix mutual information ratio (MIR) and matrix entropy difference ratio (HDR) to assess the interactions of data representation and class classification heads in supervised learning, and we determine the theoretical optimal values for MIR and HDR when Neural Collapse happens. Our experiments show that MIR and HDR can effectively explain many phenomena occurring in neural networks, for example, the standard supervised training dynamics, linear mode connectivity, and the performance of label smoothing and pruning. Additionally, we use MIR and HDR to gain insights into the dynamics of grokking, which is an intriguing phenomenon observed in supervised training, where the model demonstrates generalization capabilities long after it has learned to fit the training data. Furthermore, we introduce MIR and HDR as loss terms in supervised and semi-supervised learning to optimize the information interactions among samples and classification heads. The empirical results provide evidence of the method’s effectiveness, demonstrating that the utilization of MIR and HDR not only aids in comprehending the dynamics throughout the training process but can also enhances the training procedure itself. Kun Song 0004, Zhiquan Tan, Bochao Zou, Huimin Ma 0001, Weiran Huang 0001 |
ICML | 2 |
| 2024 | Information Flow in Self-Supervised LearningabstractIn this paper, we conduct a comprehensive analysis of two dual-branch (Siamese architecture) self-supervised learning approaches, namely Barlow Twins and spectral contrastive learning, through the lens of matrix mutual information. We prove that the loss functions of these methods implicitly optimize both matrix mutual information and matrix joint entropy. This insight prompts us to further explore the category of single-branch algorithms, specifically MAE and U-MAE, for which mutual information and joint entropy become the entropy. Building on this intuition, we introduce the Matrix Variational Masked Auto-Encoder (M-MAE), a novel method that leverages the matrix-based estimation of entropy as a regularizer and subsumes U-MAE as a special case. The empirical evaluations underscore the effectiveness of M-MAE compared with the state-of-the-art methods, including a 3.9% improvement in linear probing ViT-Base, and a 1% improvement in fine-tuning ViT-Large, both on ImageNet. Zhiquan Tan, Jingqin Yang, Weiran Huang 0001, Yang Yuan 0010, Yifan Zhang 0029 |
ICML | 1 |
| 2024 | OTMatch: Improving Semi-Supervised Learning with Optimal TransportabstractSemi-supervised learning has made remarkable strides by effectively utilizing a limited amount of labeled data while capitalizing on the abundant information present in unlabeled data. However, current algorithms often prioritize aligning image predictions with specific classes generated through self-training techniques, thereby neglecting the inherent relationships that exist within these classes. In this paper, we present a new approach called OTMatch, which leverages semantic relationships among classes by employing an optimal transport loss function to match distributions. We conduct experiments on many standard vision and language datasets. The empirical results show improvements in our method above baseline, this demonstrates the effectiveness and superiority of our approach in harnessing semantic relationships to enhance learning performance in a semi-supervised setting. Zhiquan Tan, Kaipeng Zheng, Weiran Huang 0001 |
ICML | 1 |
| 2024 | Provable Contrastive Continual LearningabstractContinual learning requires learning incremental tasks with dynamic data distributions. So far, it has been observed that employing a combination of contrastive loss and distillation loss for training in continual learning yields strong performance. To the best of our knowledge, however, this contrastive continual learning framework lacks convincing theoretical explanations. In this work, we fill this gap by establishing theoretical performance guarantees, which reveal how the performance of the model is bounded by training losses of previous tasks in the contrastive continual learning framework. Our theoretical explanations further support the idea that pre-training can benefit continual learning. Inspired by our theoretical analysis of these guarantees, we propose a novel contrastive continual learning algorithm called CILA, which uses adaptive distillation coefficients for different tasks. These distillation coefficients are easily computed by the ratio between average distillation losses and average contrastive losses from previous tasks. Our method shows great improvement on standard benchmarks and achieves new state-of-the-art performance. Yichen Wen, Zhiquan Tan, Kaipeng Zheng, Chuanlong Xie, Weiran Huang 0001 |
ICML | 2 |
| 2024 | Matrix Information Theory for Self-Supervised LearningabstractThe maximum entropy encoding framework provides a unified perspective for many non-contrastive learning methods like SimSiam, Barlow Twins, and MEC. Inspired by this framework, we introduce Matrix-SSL, a novel approach that leverages matrix information theory to interpret the maximum entropy encoding loss as matrix uniformity loss. Furthermore, Matrix-SSL enhances the maximum entropy encoding method by seamlessly incorporating matrix alignment loss, directly aligning covariance matrices in different branches. Experimental results reveal that Matrix-SSL outperforms state-of-the-art methods on the ImageNet dataset under linear evaluation settings and on MS-COCO for transfer learning tasks. Specifically, when performing transfer learning tasks on MS-COCO, our method outperforms previous SOTA methods such as MoCo v2 and BYOL up to 3.3% with only 400 epochs compared to 800 epochs pre-training. We also try to introduce representation learning into the language modeling regime by fine-tuning a 7B model using matrix cross-entropy loss, with a margin of 3.1% on the GSM8K dataset over the standard cross-entropy loss. Yifan Zhang 0029, Zhiquan Tan, Jingqin Yang, Weiran Huang 0001, Yang Yuan 0010 |
ICML | 2 |
| 2024 | Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language ModelsabstractLarge Language Models (LLMs) have transformed natural language processing and extended their powerful capabilities to multi-modal domains. As LLMs continue to advance, it is crucial to develop diverse and appropriate metrics for their evaluation. In this paper, we introduce a novel rank-based metric, Diff-eRank, grounded in information theory and geometry principles. Diff-eRank assesses LLMs by analyzing their hidden representations, providing a quantitative measure of how efficiently they eliminate redundant information during training. We demonstrate the applicability of Diff-eRank in both single-modal (e.g., language) and multi-modal settings. For language models, our results show that Diff-eRank increases with model size and correlates well with conventional metrics such as loss and accuracy. In the multi-modal context, we propose an alignment evaluation method based on the eRank, and verify that contemporary multi-modal LLMs exhibit strong alignment performance based on our method. Our code is publicly available at https://github.com/waltonfuture/Diff-eRank. Lai Wei 0005, Zhiquan Tan, Chenghai Li, Jindong Wang 0001, Weiran Huang 0001 |
NeurIPS | 2 |
| 2023 | Trade-off Between Efficiency and Consistency for Removal-based ExplanationsabstractIn the current landscape of explanation methodologies, most predominant approaches, such as SHAP and LIME, employ removal-based techniques to evaluate the impact of individual features by simulating various scenarios with specific features omitted. Nonetheless, these methods primarily emphasize efficiency in the original context, often resulting in general inconsistencies. In this paper, we demonstrate that such inconsistency is an inherent aspect of these approaches by establishing the Impossible Trinity Theorem, which posits that interpretability, efficiency, and consistency cannot hold simultaneously. Recognizing that the attainment of an ideal explanation remains elusive, we propose the utilization of interpretation error as a metric to gauge inefficiencies and inconsistencies. To this end, we present two novel algorithms founded on the standard polynomial basis, aimed at minimizing interpretation error. Our empirical findings indicate that the proposed methods achieve a substantial reduction in interpretation error, up to 31.8 times lower when compared to alternative techniques. Yifan Zhang 0029, Haowei He, Zhiquan Tan, Yang Yuan 0010 |
NeurIPS | 3 |
| 2023 | Coded Real Number Matrix Multiplication for On-Device Edge ComputingabstractThis letter addresses the challenge of multiplying large real-numbered matrices$\mathbf {A}$and$\mathbf {B}$in on-device edge computing environments with a master node and$N$worker nodes. To ensure accuracy and mitigate adversarial behavior, existing approaches rely on polynomial coding techniques. However, polynomial decoding on real numbers is numerically unstable, leading prior works to employ finite fields or complex number fields. Nevertheless, computing on finite fields may encounter computation overflows. In this letter, we propose a novel, numerically stable coded matrix multiplication scheme on real number fields. Our approach detects malicious workers and mitigates the impact of slower workers (stragglers), all while operating solely on real number fields. Zhiquan Tan, Dingli Yuan, Yifan Zhang 0029 |
IEEE Signal Process. Lett. | 1 |
| 2022 | Data Integrity Check in Distributed Storage SystemsabstractIn this paper, we propose a method of checking data integrity in distributed storage systems. Compare with conventionally used cyclic redundancy check (CRC) method, the proposed method may usually achieve a 99% reduction of bandwidth with the expense of losing only a little detection capacity. We also provide an analysis of performance and a suggestion of auxiliary matrix used in the proposed framework. Besides, a theoretical upper bound of CRC method’s detection capacity is given. Zhiquan Tan, Sian-Jheng Lin, Yunghsiang Sam Han, Bo Bai 0001, Gong Zhang 0001 |
ISIT | 1 |