EDBT 2026 Demo / reviewers in the wild / expert
Hyeongsoo Lim
dblp:429/9797
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 65% Speech recognition and synthesis · 12% Deep learning architectures and training · 12% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.3 | 2 | 2026 | GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection · AAAI 2026 SelFusion: Self-distillation for Diffusion Language Models · ACL (1) 2026 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
1.0 | 1 | 2026 | BiCycle: Group-wise Recursive Transformer Based on ASR Mechanism · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
black-box knowledge distillation |
1.0 | 1 | 2026 | GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection · AAAI 2026 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model |
1.0 | 1 | 2026 | SelFusion: Self-distillation for Diffusion Language Models · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression
large language model compression |
1.0 | 1 | 2026 | GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection · AAAI 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection · AAAI 2026 |
Machine learning › Deep learning architectures and training › transformer
recursive transformer |
1.0 | 1 | 2026 | BiCycle: Group-wise Recursive Transformer Based on ASR Mechanism · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation |
1.0 | 1 | 2026 | SelFusion: Self-distillation for Diffusion Language Models · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
parameter sharing |
0.3 | 1 | 2026 | BiCycle: Group-wise Recursive Transformer Based on ASR Mechanism · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
recursive transformer · 1.0rationale injection · 1.0masking · 1.0knowledge distillation · 1.0cross-attention · 1.0bidirectional distillation · 1.0attention pattern analysis · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BiCycle: Group-wise Recursive Transformer Based on ASR MechanismabstractRecursive transformer (RT) is a promising parameter-sharing technique for reducing computational burden of large-scale model. While RT has been successfully applied to large language models (LLMs), its effectiveness in automatic speech recognition (ASR) remains limited, despite the parallel trend of model scaling in the speech domain. In this paper, we reveal that conventional RT designs for LLMs are suboptimal for speech recognition, primarily because they do not fully consider the layer-wise specialization inherent in the ASR architecture, where lower layers focus on phonetic features and upper layers capture linguistic localization. To address this, we propose BiCycle, a novel RT scheme tailored for ASR. In particular, we firstly analyze attention patterns in a pre-trained ASR model to divide its layers into phonetic and linguistic groups. BiCycle then constructs an efficient RT model by transferring the pre-trained model’s weights in a step-wise manner and applies recursion separately to the phonetic and linguistic groups, preventing conflicts between their roles. Extensive experimental results confirm that the proposed method not only preserves the original ASR mechanism but also outperforms conventional RT approaches. Min Ho Jang, Eun Seo Seo, Hyeongsoo Lim |
AAAI | 4 |
| 2026 | GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale InjectionabstractKnowledge distillation (KD) is a promising compression technique for reducing the computational burden of large language models (LLMs). Depending on access to the teacher model’s internal parameters, KD is typically categorized into white-box and black-box KD. While white-box KD benefits from full access to intrinsic knowledge such as softmax distributions, black-box KD adopts a black-box LLM (e.g., GPT-4) as the teacher, which provides only text-level outputs via API calls. This limited supervision makes black-box KD generally less effective than its white-box counterpart. To bridge the gap between white-box and black-box KD, we propose GrayKD, a novel framework that can effectively distill text-level knowledge from a black-box LLM in a single-stage manner. In particular, rationales generated by the black-box LLM are injected into the student via a lightweight cross-attention module (teacher mode), enabling the model to approximate the black-box teacher’s output distribution without access to internal parameters. The student is then trained with the softmax-level knowledge provided by the teacher mode (student mode). Since both the teacher and student modes share the same backbone, the proposed teacher mode remains highly parameter-efficient, requiring only a small number of additional parameters for rationale injection. Experimental results on instruction-following tasks demonstrate that GrayKD achieves substantial performance improvements over existing KD methods. Hyeongsoo Lim, Hyung Yong Kim, Min Ho Jang, Eun Seo Seo, Youshin Lim, Shukjae Choi, Yunkyu Lim, Hanbin Lee, Byeong-Yeol Kim, Jiwon Yoon 0002 |
AAAI | 1 |
| 2026 | SelFusion: Self-distillation for Diffusion Language ModelsabstractDiffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded generation quality limits practical applicability.Although knowledge distillation (KD) can be a promising direction for improving performance, we empirically find that naively applying conventional KD yields only marginal gains, or even degrades generation quality.Based on these observations, we propose a novel self-distillation framework for DLMs, namely SelFusion.To enable effective KD without an external teacher model, SelFusion performs two forward passes with different masking levels, defining the hard mode with a larger masking probability and the easy mode with a smaller masking probability.However, the easy mode is not always more accurate than the hard mode and can be overconfident on incorrect tokens.Thus, we introduce bidirectional KD between the two modes, which can dynamically determine the distillation direction based on token-level correctness.Experimental results on instruction-following tasks show that the proposed self-distillation substantially outperforms other KD methods with external LLM and DLM teachers.In many configurations, the student trained with SelFusion even surpasses the performance of the LLM teacher, providing a practical path toward improving DLM generation quality. Hyeongsoo Lim, Eun Seo Seo, Min Ho Jang, Jiwon Yoon 0002 |
ACL (1) | 1 |