Chengxing Zhou

dblp:358/6095 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0006-8809-2422ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 28% Generative modeling · 28% Vision and language · 26%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention Into Convolutions · ICCV 2025
Machine learning › Generative modeling › diffusion model
efficient diffusion model
0.912025
Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention Into Convolutions · ICCV 2025
Machine learning › Efficient and distributed learning › distillation
self-attention distillation
0.912025
Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention Into Convolutions · ICCV 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention Into Convolutions · ICCV 2025
Computer vision › Vision and language
vision-language model
0.912025
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference · ACL (1) 2025
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.812024
ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks · ACM Multimedia 2024
Machine learning › Trustworthy machine learning
language model interpretability
0.312025
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

region activation analysis · 0.9pyramid convolution · 0.9parameter-efficient training · 0.9knowledge distillation · 0.9benchmark reformulation · 0.8
YearPublicationVenuePosition
2025 Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
abstract
Siyuan Wang, Dianyi Wang, Chengxing Zhou, Zejun Li, Zhihao Fan, Xuanjing Huang, Zhongyu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Siyuan Wang 0025, Dianyi Wang, Chengxing Zhou, Zhihao Fan, Xuanjing Huang 0001, Zhongyu Wei
ACL (1)3
2025 Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention Into Convolutions
abstract
Contemporary diffusion models built upon U-Net or Diffusion Transformer (DiT) architectures have revolutionized image generation through transformer-based attention mechanisms. The prevailing paradigm has commonly employed self-attention with quadratic computational complexity to handle global spatial relationships in complex images, thereby synthesizing high-fidelity images with coherent visual semantics.Contrary to conventional wisdom, our systematic layer-wise analysis reveals an interesting discrepancy: self-attention in pre-trained diffusion models predominantly exhibits localized attention patterns, closely resembling convolutional inductive biases. This suggests that global interactions in self-attention may be less critical than commonly assumed.Driven by this, we propose \(Δ\)ConvFusion to replace conventional self-attention modules with Pyramid Convolution Blocks (\(Δ\)ConvBlocks).By distilling attention patterns into localized convolutional operations while keeping other components frozen, \(Δ\)ConvFusion achieves performance comparable to transformer-based counterparts while reducing computational cost by 6929$\times$ and surpassing LinFusion by 5.42$\times$ in efficiency--all without compromising generative fidelity.
ZiYi Dong, Chengxing Zhou, Weijian Deng, Pengxu Wei, Xiangyang Ji, Liang Lin 0004
ICCV2
2024 ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks
abstract
Recent years have witnessed remarkable progress in the development of large vision-language models (LVLMs). Benefiting from the strong language backbones and efficient cross-modal alignment strategies, LVLMs exhibit surprising capabilities to perceive visual signals and perform visually grounded reasoning. However, the capabilities of LVLMs have not been comprehensively and quantitatively evaluated. Most existing multi-modal benchmarks require task-oriented input-output formats, posing great challenges to automatically assess the free-form text output of LVLMs. To effectively leverage the annotations available and reduce the manual efforts required for constructing new benchmarks, we propose to re-formulate existing benchmarks into unified LVLM-compatible formats. Through systematic data collection and reformulation, we present ReForm-Eval benchmark, offering substantial data for evaluating various capabilities of LVLMs. Through extensive experiments and analysis in ReForm-Eval, we demonstrate the comprehensiveness and reliability of ReForm-Eval in assessing various LVLMs. Our benchmark and evaluation framework is now available at https://github.com/FudanDISC/ReForm-Eval
Mengfei Du, Qingwen Liu 0002, Binhao Wu, Jiwen Zhang, Chengxing Zhou, Zhihao Fan, Jie Fu 0001, Jingjing Chen 0001, Zhongyu Wei, Xuanjing Huang 0001
ACM Multimedia7