Yuanchi Zhang

dblp:304/3086 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 42% Language models and text generation · 16% Deep learning architectures and training · 11%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.222026
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis · AAAI 2026
Continual Knowledge Distillation for Neural Machine Translation · ACL (1) 2023
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
expert pruning
1.012026
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis · AAAI 2026
Machine learning › Deep learning architectures and training
mixture of experts
1.012026
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
quantization
1.012026
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis · AAAI 2026
Robotics › Robot navigation and mapping
active perception
0.912025
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models · ACL (1) 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models · ACL (1) 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
Continual Knowledge Distillation for Neural Machine Translation · ACL (1) 2023
Natural language and speech › Machine translation
neural machine translation
0.712023
Continual Knowledge Distillation for Neural Machine Translation · ACL (1) 2023
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.212024
Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

structured pruning · 1.0mixed-precision quantization · 1.0micro-expert analysis · 1.0multimodal large language model evaluation · 0.9self-distillation · 0.8in-context learning · 0.8dialogue simulation · 0.8knowledge distillation · 0.7continual learning · 0.7
YearPublicationVenuePosition
2026 CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
abstract
Large Language Models (LLMs) with Mixture-of-Experts (MoE) architectures are distinguished by their strong performance scaling with increasing parameters across a wide range of tasks, yet they also suffer from substantial computational and storage overheads. Notably, the performance gains of MoE models do not scale proportionally with the growth in expert parameters. While prior works attempt to reduce parameters via expert-level pruning, merging, or decomposition, they still suffer from challenges in both performance and computational efficiency. In this paper, we address these challenges by introducing micro-expert as a finer-grained compression unit that spans across matrices. We first establish a more fundamental perspective, viewing MoE layers as mixtures of micro-experts, and present CAMERA, a lightweight and training-free framework for identifying micro-expert redundancy. Our analysis uncovers significant variance in micro-expert contributions during decoding. Based on this insight, we further propose CAMERA-P, a structured micro-expert pruning framework, and CAMERA-Q, a mixed-precision quantization idea designed for micro-experts. Extensive experiments on nine downstream tasks show that CAMERA-P consistently outperforms strong baselines under pruning ratios ranging from 20% to 60%. Furthermore, CAMERA-Q achieves superior results under aggressive 2-bit quantization, surpassing existing matrix- and channel-level ideas. Notably, our method enables complete micro-expert analysis of Qwen2-57B-A14B in less than 5 minutes on a single NVIDIA A100-40GB GPU.
Yuzhuang Xu, Xu Han 0007, Yuanchi Zhang, Shiyu Ji, Qingfu Zhu, Wanxiang Che
AAAI3
2025 ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
abstract
Ziyue Wang, Chi Chen, Fuwen Luo, Yurui Dong, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang, Peng Li, Yang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziyue Wang 0002, Chi Chen 0005, Fuwen Luo, Yurui Dong 0001, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang 0014, Peng Li 0030, Yang Liu 0005
ACL (1)5
2024 Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
abstract
Xiaolong Wang, Yile Wang, Yuanchi Zhang, Fuwen Luo, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiaolong Wang 0014, Yile Wang 0001, Yuanchi Zhang, Fuwen Luo, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)3
2024 Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
abstract
Yuanchi Zhang, Yile Wang, Zijun Liu, Shuo Wang, Xiaolong Wang, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yuanchi Zhang, Yile Wang 0001, Shuo Wang 0013, Xiaolong Wang 0014, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)1
2023 Continual Knowledge Distillation for Neural Machine Translation
abstract
While many parallel corpora are not publicly accessible for data copyright, data privacy and competitive differentiation reasons, trained translation models are increasingly available on open platforms.In this work, we propose a method called continual knowledge distillation to take advantage of existing translation models to improve one model of interest.The basic idea is to sequentially transfer knowledge from each trained model to the distilled model.Extensive experiments on Chinese-English and German-English datasets show that our method achieves significant and consistent improvements over strong baselines under both homogeneous and heterogeneous trained model settings and is robust to malicious models.
Yuanchi Zhang, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)1
2022 DirectQuote: A Dataset for Direct Quotation Extraction and Attribution in News Articles
abstract
Quotation extraction and attribution are challenging tasks, aiming at determining the spans containing quotations and attributing each quotation to the original speaker. Applying this task to news data is highly related to fact-checking, media monitoring and news tracking. Direct quotations are more traceable and informative, and therefore of great significance among different types of quotations. Therefore, this paper introduces DirectQuote, a corpus containing 19,760 paragraphs and 10,279 direct quotations manually annotated from online news media. To the best of our knowledge, this is the largest and most complete corpus that focuses on direct quotations in news texts. We ensure that each speaker in the annotation can be linked to a specific named entity on Wikidata, benefiting various downstream tasks. In addition, for the first time, we propose several sequence labeling models as baseline methods to extract and attribute quotations simultaneously in an end-to-end manner.
Yuanchi Zhang
LREC1