EDBT 2026 Demo / reviewers in the wild / expert
Fangmin Chen
dblp:320/5863
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0002-9613-7898ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 50% Deep learning architectures and training · 50% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
1.0 | 1 | 2026 | S2O: Early Stopping for Sparse Attention via Online Permutation · ACL (1) 2026 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
1.0 | 1 | 2026 | S2O: Early Stopping for Sparse Attention via Online Permutation · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models · AAAI 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.9 | 1 | 2025 | ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models · AAAI 2025 |
Image and video processing › super-resolution › image super-resolution
single image super-resolution |
0.7 | 1 | 2023 | Unfolding Once is Enough: A Deployment-Friendly Transformer Unit for Super-Resolution · ACM Multimedia 2023 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.3 | 1 | 2026 | S2O: Early Stopping for Sparse Attention via Online Permutation · ACL (1) 2026 |
Hardware accelerators and domain-specific architectures
efficient inference |
0.2 | 1 | 2023 | Unfolding Once is Enough: A Deployment-Friendly Transformer Unit for Super-Resolution · ACM Multimedia 2023 |
Hardware accelerators and domain-specific architectures › dataflow optimization
operator fusion |
0.2 | 1 | 2023 | Unfolding Once is Enough: A Deployment-Friendly Transformer Unit for Super-Resolution · ACM Multimedia 2023 |
Methods — techniques the papers use, named apart from their topics
vision transformer · 1.3operator fusion · 1.3layer normalization substitution · 1.3online permutation · 1.0early stopping · 1.0distribution correction · 0.9bit balance strategy · 0.9binary tensor core · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S2O: Early Stopping for Sparse Attention via Online PermutationabstractYu Zhang, Songwei Liu, Chenqian Yan, Linsheng, Beichen Ning, Fangmin Chen, Xing Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Songwei Liu, Chenqian Yan, Beichen Ning, Fangmin Chen |
ACL (1) | 6 |
| 2025 | ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language ModelsabstractLarge Language Models (LLMs) have revolutionized natural language processing tasks. However, their practical application is constrained by substantial memory and computational demands. Post-training quantization (PTQ) is considered an effective method to accelerate LLM inference. Despite its growing popularity in LLM model compression, PTQ deployment faces two major challenges. First, low-bit quantization leads to performance degradation. Second, restricted by the limited integer computing unit type on GPUs, quantized matrix operations with different precisions cannot be effectively accelerated. To address these issues, we introduce a novel arbitrary-bit quantization algorithm and inference framework, ABQ-LLM. It achieves superior performance across various quantization settings and enables efficient arbitrary-precision quantized inference on the GPU. ABQ-LLM introduces several key innovations: (1) a distribution correction method for transformer blocks to mitigate distribution differences caused by full quantization of weights and activations, improving performance at low bit-widths. (2) the bit balance strategy to counteract performance degradation from asymmetric distribution issues at very low bit-widths (e.g., 2-bit). (3) an innovative quantization acceleration framework that reconstructs the quantization matrix multiplication of arbitrary precision combinations based on BTC (Binary TensorCore) equivalents, gets rid of the limitations of INT4/INT8 computing units. ABQ-LLM can convert each component bit width gain into actual acceleration gain, maximizing performance under mixed precision(e.g., W6A6, W2A8). Based on W2*A8 quantization configuration on LLaMA-7B model, it achieved a WikiText2 perplexity of 7.59 (2.17⬇ vs 9.76 in AffineQuant). Compared to SmoothQuant, we realized 1.6x acceleration improvement and 2.7x memory compression gain. Songwei Liu, Yusheng Xie, Miao Wei, Fangmin Chen, Xing Mei |
AAAI | 8 |
| 2023 | Unfolding Once is Enough: A Deployment-Friendly Transformer Unit for Super-ResolutionabstractRecent years have witnessed a few attempts of vision transformers for single image super-resolution (SISR). Since the high resolution of intermediate features in SISR models increases memory and computational requirements, efficient SISR transformers are more favored. Based on some popular transformer backbone, many methods have explored reasonable schemes to reduce the computational complexity of the self-attention module while achieving impressive performance. However, these methods only focus on the performance on the training platform (e.g., Pytorch/Tensorflow) without further optimization for the deployment platform (e.g., TensorRT). Therefore, they inevitably contain some redundant operators, posing challenges for subsequent deployment in real-world applications. In this paper, we propose a deployment-friendly transformer unit, namely UFONE (i.e., UnFolding ONce is Enough), to alleviate these problems. In each UFONE, we introduce an Inner-patch Transformer Layer (ITL) to efficiently reconstruct the local structural information from patches and a Spatial-Aware Layer (SAL) to exploit the long-range dependencies between patches. Based on UFONE, we propose a Deployment-friendly Inner-patch Transformer Network (DITN) for the SISR task, which can achieve favorable performance with low latency and memory usage on both training and deployment platforms. Furthermore, to further boost the deployment efficiency of the proposed DITN on TensorRT, we also provide an efficient substitution for layer normalization and propose a fusion optimization strategy for specific operators. Extensive experiments show that our models can achieve competitive results in terms of qualitative and quantitative performance with high deployment efficiency. Yong Liu 0031, Hang Dong 0001, Boyang Liang, Songwei Liu, Qingji Dong, Kai Chen 0023, Fangmin Chen, Lean Fu, Fei Wang 0008 |
ACM Multimedia | 7 |