EDBT 2026 Demo / reviewers in the wild / expert
Yiming Dong
dblp:319/5065
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 49% Optimization for machine learning · 14% Efficient and distributed learning · 14% | |
| Computer graphics and multimedia
2 papers |
Geometric modeling and processing · 72% Visual content generation and editing · 28% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
1.9 | 3 | 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025 Gauge Equivariant Transformer · NeurIPS 2021 Efficient Equivariant Network · NeurIPS 2021 |
Machine learning › Deep learning architectures and training
equivariant neural network |
1.0 | 2 | 2021 | Gauge Equivariant Transformer · NeurIPS 2021 Efficient Equivariant Network · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › equivariant neural network
equivariant self-attention |
1.0 | 2 | 2021 | Gauge Equivariant Transformer · NeurIPS 2021 Efficient Equivariant Network · NeurIPS 2021 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
1.0 | 1 | 2026 | From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning |
1.0 | 1 | 2026 | From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning · AAAI 2026 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
0.9 | 1 | 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025 |
Machine learning › Optimization for machine learning
learning rate schedule |
0.9 | 1 | 2025 | Stepsize anything: A unified learning rate schedule for budgeted-iteration training · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.9 | 1 | 2025 | On the O(√d/K1/4) Convergence Rate of AdamW Measured by ℓ1 Norm · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025 |
Visual content generation and editing › image colorization
line art colorization |
0.9 | 1 | 2025 | KISSColor: Kinetic and Intuitive Stroke Stretching for Vector Drawing Colorization · ACM Trans. Graph. 2025 |
Geometric modeling and processing › shape representation
vector graphics |
0.9 | 1 | 2025 | KISSColor: Kinetic and Intuitive Stroke Stretching for Vector Drawing Colorization · ACM Trans. Graph. 2025 |
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
adaptive gradient methods |
0.9 | 1 | 2025 | On the O(sqrt(d)/T^(1/4)) Convergence Rate of RMSProp and Its Momentum Extension Measured by l_1 Norm · J. Mach. Learn. Res. 2025 |
Mathematical optimization
continuous optimization |
0.9 | 1 | 2025 | On the O(sqrt(d)/T^(1/4)) Convergence Rate of RMSProp and Its Momentum Extension Measured by l_1 Norm · J. Mach. Learn. Res. 2025 |
Mathematical optimization
nonconvex optimization |
0.9 | 1 | 2025 | On the O(sqrt(d)/T^(1/4)) Convergence Rate of RMSProp and Its Momentum Extension Measured by l_1 Norm · J. Mach. Learn. Res. 2025 |
Machine learning › Deep learning architectures and training › equivariant neural network
group equivariant convolution |
0.5 | 1 | 2021 | Efficient Equivariant Network · NeurIPS 2021 |
Geometric modeling and processing
manifold learning |
0.5 | 1 | 2021 | Gauge Equivariant Transformer · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
token distribution analysis · 2.0diversity-control strategies · 2.0parallel transport · 1.0multi-head self-attention · 1.0winding-number fields · 0.9skip connections · 0.9multi-latent attention · 0.9momentum · 0.9mixed-integer programming · 0.9l1 norm convergence analysis · 0.9kinetic data structures · 0.9group-query attention · 0.9convergence analysis · 0.9condition number analysis · 0.9complexity analysis · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Macro to Micro: Probing Dataset Diversity in Language Model Fine-TuningabstractDataset diversity plays a pivotal role for the successful training of many machine learning models, particularly in the supervised fine-tuning (SFT) stage of large language model (LLM) development. Despite increasing recognition of its importance, systematic analyses of dataset diversity still remain underexplored. To address this gap, this work presents a systematic taxonomy of existing diversity-control strategies, which primarily focus on the instruction component, operating at either macroscopic (entire instruction semantics) or mesoscopic levels (instruction units), and furthermore introduces a novel analysis of microscopic diversity within the response component, specifically analyzing the statistical distribution of tokens in SFT training samples. In the experimental evaluation, we construct fixed-size datasets (e.g., 10,000 samples each) from a corpus of 117,000 open-source SFT samples, incorporating six distinct diversity-control strategies spanning macro-, meso-, and microscopic levels applied to both instructions and responses. We then fine-tune LLMs on these datasets to assess the six diversity-control strategies. Results reveal that while macroscopic and mesoscopic strategies lead to higher performance with increasing diversity, the microscopic strategy in responses exhibits not only a stronger correlation between model performance and the degree of diversity, but also superior performance with maximum diversity across all strategies. These findings offer actionable insights for constructing high-performance SFT datasets. Yiming Dong |
AAAI | 3 |
| 2025 | On the O(√d/K1/4) Convergence Rate of AdamW Measured by ℓ1 Norm
Huan Li 0007, Yiming Dong, Zhouchen Lin |
NeurIPS | 2 |
| 2025 | Stepsize anything: A unified learning rate schedule for budgeted-iteration trainingabstractThe expanding computational costs and limited resources underscore the critical need for budgeted-iteration training, which aims to achieve optimal learning within predetermined iteration budgets. While learning rate schedules fundamentally govern the performance of different networks and tasks, particularly in budgeted-iteration scenarios, their design remains largely heuristic, lacking theoretical foundations. In addition, the optimal learning rate schedule requires extensive trial-and-error selection, making the training process inefficient. In this work, we propose the Unified Budget-Aware (UBA) schedule, a theoretically grounded learning rate schedule that consistently outperforms commonly-used schedules among diverse architectures and tasks under different constrained training budgets. First, we bridge the gap by constructing a novel training budget-aware optimization framework, which explicitly accounts for the robustness to landscape curvature variations. From this framework, we derive the UBA schedule, controlled by a single hyper-parameter $\varphi$ that provides a trade-off between flexibility and simplicity, eliminating the need for per-network numerical optimization. Moreover, we establish a theoretical connection between $\varphi$ and the condition number, adding interpretation and justification to our approach. Besides, we prove the convergence for different values of $\varphi$. We offer practical guidelines for $\varphi$ selection via theoretical analysis and empirical results. Extensive experimental results show that UBA $\textit{consistently surpasses}$ the commonly-used schedules across diverse vision and language tasks, spanning network architectures (e.g., ResNet, OLMo) and scales, under different training-iteration budgets. Anda Tang, Yiming Dong, Yutao Zeng, Zhouchen Lin |
NeurIPS | 2 |
| 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value HeadsabstractTransformer models have driven breakthroughs across various language tasks by their strong capability to learn rich contextual representations. Scaling them to improve representation, however, often demands substantial memory and compute costs, such as the Key-Value (KV) cache used during auto-regressive decoding. Skip connections offer a promising way to improve representation without bloating resource usage, yet most prior works either improve expressivity while leaving KV costs unchanged, or reduce memory at the cost of weaker representation. In this work, we propose SkipV1Former, a Transformer variant that uses skip connections from the first layer's Value heads to strengthen model representation and reduce KV cache. Specifically, from the second block onward, each layer reuses half of its Value heads from the very first layer, while computing the other half as usual-cutting Value projections and V cache by nearly 50 \%. Theoretically, we show that routing uncompressed first-layer Values into deeper layers restores information lost to compression and accelerates the model’s implicit mesa-optimization-a key pattern of Transformer in auto-regressive tasks. Empirically, across different model scales, SkipV1Former delivers consistent reductions of approximately 25 \% in KV cache while improving perplexity relative to standard Multi-Head Attention (MHA) Transformers and some advanced variants. Moreover, we propose a recipe for uptraining existing MHA Transformer checkpoints to SkipV1Former with only 10-15\% additional compute. Finally, SkipV1Former can seamlessly combine advanced methods like Group-Query Attention and Multi-Latent Attention to achieve further KV cache savings and performance improvement. When combined with YOCO, it cuts KV cache size by nearly 50 \% while still improving performance. The code is
available at: https://github.com/Zhoutong-Wu/SkipV1Former. Zhoutong Wu, Yiming Dong, Chenheng Zhang, Cong Fang 0001, Zhouchen Lin |
NeurIPS | 3 |
| 2025 | AdaMSS: Adaptive Multi-Subspace Approach for Parameter-Efficient Fine-TuningabstractIn this paper, we propose AdaMSS, an adaptive multi-subspace approach for parameter-efficient fine-tuning of large models. Unlike traditional parameter-efficient fine-tuning methods that operate within a large single subspace of the network weights, AdaMSS leverages subspace segmentation to obtain multiple smaller subspaces and adaptively reduces the number of trainable parameters during training, ultimately updating only those associated with a small subset of subspaces most relevant to the target downstream task. By using the lowest-rank representation, AdaMSS achieves more compact expressiveness and finer tuning of the model parameters. Theoretical analyses demonstrate that AdaMSS has better generalization guarantee than LoRA, PiSSA, and other single-subspace low-rank-based methods. Extensive experiments across image classification, natural language understanding, and natural language generation tasks show that AdaMSS achieves comparable performance to full fine-tuning and outperforms other parameter-efficient fine-tuning methods in most cases, all while requiring fewer trainable parameters. Notably, on the ViT-Large model, AdaMSS achieves 4.7\% higher average accuracy than LoRA across seven tasks, using just 15.4\% of the trainable parameters. On RoBERTa-Large, AdaMSS outperforms PiSSA by 7\% in average accuracy across six tasks while reducing the number of trainable parameters by approximately 94.4\%. These results demonstrate the effectiveness of AdaMSS in parameter-efficient fine-tuning. The code for AdaMSS is available at https://github.com/jzheng20/AdaMSS. Wanglong Lu, Yiming Dong, Chaojie Ji, Yankai Cao, Zhouchen Lin |
NeurIPS | 3 |
| 2025 | On the O(sqrt(d)/T^(1/4)) Convergence Rate of RMSProp and Its Momentum Extension Measured by l_1 NormabstractAlthough adaptive gradient methods have been extensively used in deep learning, their convergence rates proved in the literature are all slower than that of SGD, particularly with respect to their dependence on the dimension. This paper considers the classical RMSProp and its momentum extension and establishes the convergence rate of $\frac{1}{T}\sum_{k=1}^TE\left[||\nabla f(\mathbf{x}^k)||_1\right]\leq O(\frac{\sqrt{d}C}{T^{1/4}})$ measured by $\ell_1$ norm without the bounded gradient assumption, where $d$ is the dimension of the optimization variable, $T$ is the iteration number, and $C$ is a constant identical to that appeared in the optimal convergence rate of SGD. Our convergence rate matches the lower bound with respect to all the coefficients except the dimension $d$. Since $||\mathbf{x}||_2\ll ||\mathbf{x}||_1\leq\sqrt{d}||\mathbf{x}||_2$ for problems with extremely large $d$, our convergence rate can be considered to be analogous to the $\frac{1}{T}\sum_{k=1}^TE\left[||\nabla f(\mathbf{x}^k)||_2\right]\leq O(\frac{C}{T^{1/4}})$ rate of SGD in the ideal case of $||\nabla f(\mathbf{x})||_1=\varTheta(\sqrt{d})||\nabla f(\mathbf{x})||_2$. Huan Li 0007, Yiming Dong, Zhouchen Lin |
J. Mach. Learn. Res. | 2 |
| 2025 | KISSColor: Kinetic and Intuitive Stroke Stretching for Vector Drawing ColorizationabstractHand-drawn vector sketches often contain implied lines, imprecise intersections, and unintended gaps, making it challenging to identify closed regions for colorization. These challenges become more pronounced as the number of strokes increases. In this paper, we present KISSColor, a novel method for inferring users' intended closed regions. Specifically, we propose intuitive stroke stretching by extending open strokes along tangent isolines of winding-number fields, which provably form geometrically aligned closed regions. Extending all open strokes can lead to overly fragmented regions due to redundant intersections. While a Mixed Integer Programming (MIP) formulation helps reduce redundancy, it is computationally expensive. To improve efficiency, we introduce kinetic stroke stretching, which grows all strokes simultaneously and prioritizes early intersections using a kinetic data structure. This approach preserves stylistic ambiguity for lines requiring long extensions. Based on the growth results, redundant regions are suppressed to minimize fragmentation. We conduct extensive experiments demonstrating the effectiveness of KISSColor, which generates more intuitive partitions, especially for imprecise sketches (see teaser figure). Our code and data will be released upon publication. Yiming Dong, Hongxu Xin, Zhiyang Dou, Rui Xu 0016, Yuan Liu 0025, Shuang-Min Chen, Shi-Qing Xin, Changhe Tu, Taku Komura, Wenping Wang 0001 |
ACM Trans. Graph. | 1 |
| 2024 | Reducing Memory Footprint in Deep Network Training by Gradient Space Reutilization
Yiming Dong, Zhouchen Lin |
PRCV (2) | 1 |
| 2021 | Efficient Equivariant NetworkabstractConvolutional neural networks (CNNs) have dominated the field of Computer Vision and achieved great success due to their built-in translation equivariance. Group equivariant CNNs (G-CNNs) that incorporate more equivariance can significantly improve the performance of conventional CNNs. However, G-CNNs are faced with two major challenges: \emph{spatial-agnostic problem} and \emph{expensive computational cost}. In this work, we propose a general framework of previous equivariant models, which includes G-CNNs and equivariant self-attention layers as special cases. Under this framework, we explicitly decompose the feature aggregation operation into a kernel generator and an encoder, and decouple the spatial and extra geometric dimensions in the computation. Therefore, our filters are essentially dynamic rather than being spatial-agnostic. We further show that our \emph{E}quivariant model is parameter \emph{E}fficient and computation \emph{E}fficient by complexity analysis, and also data \emph{E}fficient by experiments, so we call our model $E^4$-Net. Extensive experiments verify that our model can significantly improve previous works with smaller model size.Especially, under the setting of training on $1/5$ data of CIFAR10, our model improves G-CNNs by $5\%+$ accuracy,while using only $56\%$ parameters and $68\%$ FLOPs. Lingshen He, Zhengyang Shen, Yiming Dong, Yisen Wang 0001, Zhouchen Lin |
NeurIPS | 4 |
| 2021 | Gauge Equivariant TransformerabstractAttention mechanism has shown great performance and efficiency in a lot of deep learning models, in which relative position encoding plays a crucial role. However, when introducing attention to manifolds, there is no canonical local coordinate system to parameterize neighborhoods. To address this issue, we propose an equivariant transformer to make our model agnostic to the orientation of local coordinate systems (\textit{i.e.}, gauge equivariant), which employs multi-head self-attention to jointly incorporate both position-based and content-based information. To enhance expressive ability, we adopt regular field of cyclic groups as feature fields in intermediate layers, and propose a novel method to parallel transport the feature vectors in these fields. In addition, we project the position vector of each point onto its local coordinate system to disentangle the orientation of the coordinate system in ambient space (\textit{i.e.}, global coordinate system), achieving rotation invariance. To the best of our knowledge, we are the first to introduce gauge equivariance to self-attention, thus name our model Gauge Equivariant Transformer (GET), which can be efficiently implemented on triangle meshes. Extensive experiments show that GET achieves state-of-the-art performance on two common recognition tasks. Lingshen He, Yiming Dong, Yisen Wang 0001, Dacheng Tao, Zhouchen Lin |
NeurIPS | 2 |