Guangtao Lyu

dblp:327/8033 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0006-4474-2964ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 26% Efficient and distributed learning · 19% Language models and text generation · 19%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
dataset distillation
1.012026
Channel-masked Asymmetric Distribution Matching for Cross-Domain Generalized Dataset Distillation · AAAI 2026
Machine learning › Transfer learning and domain adaptation
domain generalization
1.012026
Channel-masked Asymmetric Distribution Matching for Cross-Domain Generalized Dataset Distillation · AAAI 2026
Natural language and speech › Language models and text generation
knowledge editing
1.012026
Fisher-Driven Adaptive Locating for Knowledge Editing in Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation › knowledge editing
locate-then-edit
1.012026
Fisher-Driven Adaptive Locating for Knowledge Editing in Large Language Models · ACL (1) 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
Semantically Comprehensive Token Pruning in LVLMs via Maximizing Concept Coverage · ACL (1) 2026
Machine learning › Efficient and distributed learning › model compression › token pruning
visual token pruning
1.012026
Semantically Comprehensive Token Pruning in LVLMs via Maximizing Concept Coverage · ACL (1) 2026
Computer vision › Vision and language › motion-language model
motion-language pretraining
0.912025
Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization · ICLR 2025
Computer vision › Vision and language › video-language understanding
motion-language understanding
0.912025
Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization · ICLR 2025
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.912025
Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization · ICLR 2025
Machine learning › Generative modeling › motion generation
gesture generation
0.812024
LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures · ACL (1) 2024
Machine learning › Generative modeling
multimodal generation
0.812024
LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures · ACL (1) 2024
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
0.312026
Semantically Comprehensive Token Pruning in LVLMs via Maximizing Concept Coverage · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

sparse autoencoder · 1.0probing · 1.0maximum concept coverage · 1.0marginal semantic gain · 1.0low-rank modulation · 1.0fisher information · 1.0distribution matching · 1.0channel masking · 1.0masked language modeling · 0.9contrastive learning · 0.9
YearPublicationVenuePosition
2026 Channel-masked Asymmetric Distribution Matching for Cross-Domain Generalized Dataset Distillation
abstract
Dataset distillation has achieved remarkable progress as an effective approach for data compression. However, real-world data often comes from diverse domains, leading to potential mismatches between the domains of synthesized images and those of the evaluation set. Existing methods primarily assume domain alignment between them, which limits their generalization ability in the above cross-domain scenarios. In this paper, we aim to ensure that images synthesized from known domains maintain robust performance on unseen domains and propose a novel framework called Channel-masked Asymmetric Distribution Matching (CADM). During asymmetric distribution matching, domain-sensitive channels of real data are selectively masked at different layers to extract domain-invariant features that guide synthetic data optimization. To further improve synthetic data representation, we introduce a class-focused domain-agnostic regularization to capture class-relevant knowledge while ignoring domain-specific information. Experiments show that our method produces domain-robust synthetic data and substantially improves generalization performance on unseen domains.
Jiexi Yan, Guangtao Lyu, Erkun Yang, Guihai Chen, Yanhua Yang
AAAI4
2026 Semantically Comprehensive Token Pruning in LVLMs via Maximizing Concept Coverage
abstract
High-resolution visual tokens impose substantial computational burdens owing to extreme redundancy in Large Visual Language Models (LVLMs).Existing visual token pruning methods typically leverage simple metrics derived from human experience, such as attention or similarity, to rank and select tokens within a highly entangled feature space.However, these metrics lack interpretability and often introduce human bias, failing to capture the genuine semantic significance of tokens, especially amidst the inherent semantic complexity and ambiguity of visual tokens.To mitigate this limitation, we propose a novel Semantically Comprehensive Token Selection (SCTS) method for unbiased, interpretable visual token pruning via a concept-driven paradigm.To unravel the model's intrinsic semantic representation mechanism, we first introduce a Sparse Autoencoder to disentangle visual features into an interpretable space, with each dimension encoding a distinct semantic concept.We then formulate the token pruning task as a Maximum Concept Coverage problem, quantifying the Marginal Semantic Gain (MSG) of each token's contribution to uncovered concepts and iteratively selecting tokens with the highest MSG.This concept-centric approach prioritizes tokens with unique semantic contributions, guaranteeing semantic comprehensiveness while preserving robust performance even at high compression ratios.Extensive experiments across multiple LVLM architectures and benchmarks verify that SCTS consistently outperforms state-of-the-art approaches, achieving a superior trade-off between computational efficiency and semantic completeness.
Xu Yang 0019, Guangtao Lyu, Cheng Deng 0002
ACL (1)5
2026 Fisher-Driven Adaptive Locating for Knowledge Editing in Large Language Models
abstract
Large language models (LLMs) store extensive factual knowledge acquired during pretraining, yet this knowledge is inherently static and may become inaccurate or outdated, leading to knowledge hallucinations.Knowledge editing offers an efficient alternative to full retraining by enabling targeted factual updates while preserving overall model behavior.Existing locate-then-edit methods, however, rely on fixed layer selection strategies, treating the locating stage as a static design choice and failing to account for the hierarchical and instancedependent nature of knowledge representation in LLMs.In this paper, we propose FiDAL, a Fisher-driven adaptation-aware locating strategy that dynamically identifies which model components should be edited for a given knowledge update.FiDAL formulates localization as a weight-level decision problem and leverages Fisher Information to select layers that are both influential and sensitive to factual modifications.A lightweight probing stage with low-rank modulation enables efficient localization with minimal overhead.Experiments on standard benchmarks demonstrate that FiDAL consistently improves editing effectiveness and knowledge preservation across multiple editing methods.
Jiexi Yan, Guangtao Lyu, Muli Yang, Cheng Deng 0002
ACL (1)3
2025 Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization
abstract
Recently, the comprehensive understanding of human motion has been a prominent area of research due to its critical importance in many fields. However, existing methods often prioritize specific downstream tasks and roughly align text and motion features within a CLIP-like framework. This results in a lack of rich semantic information which restricts a more profound comprehension of human motions, ultimately leading to unsatisfactory performance. Therefore, we propose a novel motion-language representation paradigm to enhance the interpretability of motion representations by constructing a universal motion-language space, where both motion and text features are concretely lexicalized, ensuring that each element of features carries specific semantic meaning. Specifically, we introduce a multi-phase strategy mainly comprising Lexical Bottlenecked Masked Language Modeling to enhance the language model's focus on high-entropy words crucial for motion semantics, Contrastive Masked Motion Modeling to strengthen motion feature extraction by capturing spatiotemporal dynamics directly from skeletal motion, Lexical Bottlenecked Masked Motion Modeling to enable the motion model to capture the underlying semantic features of motion for improved cross-modal understanding, and Lexical Contrastive Motion-Language Pretraining to align motion and text lexicon representations, thereby ensuring enhanced cross-modal coherence. Comprehensive analyses and extensive experiments across multiple public datasets demonstrate that our model achieves state-of-the-art performance across various tasks and scenarios.
Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng 0002
ICLR1
2025 Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative Modeling
abstract
In dance performances, choreographers define the visual expression of movement, while cinematographers shape its final presentation through camera work. Consequently, the synthesis of camera movements informed by both music and dance has garnered increasing research interest. While recent advancements have led to notable progress in this area, existing methods predominantly operate in an offline manner—that is, they require access to the entire dance sequence before generating corresponding camera motions. This constraint renders them impractical for real-time applications, particularly in live stage performances, where immediate responsiveness is essential. To address this limitation, we introduce a more practical yet challenging task: online camera movement synthesis, in which camera trajectories must be generated using only the current and preceding segments of dance and music. In this paper, we propose TemMEGA (Temporal Masked Generative Modeling), a unified framework capable of handling both online and offline camera movement generation. TemMEGA consists of three key components. First, a discrete camera tokenizer encodes camera motions as discrete tokens via a discrete quantization scheme. Second, a consecutive memory encoder captures historical context by jointly modeling long- and short-term temporal dependencies across dance and music sequences. Finally, a temporal conditional masked transformer is employed to predict future camera motions by leveraging masked token prediction. Extensive experimental evaluations demonstrate the effectiveness of our TemMEGA, highlighting its superiority in both online and offline camera movement synthesis.
Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng 0002
NeurIPS2
2024 LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures
abstract
In response to the escalating demand for digital human representations, progress has been made in the generation of realistic human gestures from given speeches.Despite the remarkable achievements of recent research, the generation process frequently includes unintended, meaningless, or non-realistic gestures.To address this challenge, we propose a gesture translation paradigm, GesTran, which leverages large language models (LLMs) to deepen the understanding of the connection between speech and gesture and sequentially generates human gestures by interpreting gestures as a unique form of body language.The primary stage of the proposed framework employs a transformer-based auto-encoder network to encode human gestures into discrete symbols.Following this, the subsequent stage utilizes a pre-trained LLM to decipher the relationship between speech and gesture, translating the speech into gesture by interpreting the gesture as unique language tokens within the LLM.Our method has demonstrated state-of-the-art performance improvement through extensive and impartial experiments conducted on public TED and TED-Expressive datasets.
Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng 0002
ACL (1)2
2023 HelixNet: Dual Helix Cooperative Decoders for Scene Text Removal
Guangtao Lyu, Anna Zhu
PRCV (7)2
2023 FETNet: Feature erasing and transferring network for scene text removal
Guangtao Lyu, Kun Liu 0027, Anna Zhu, Seiichi Uchida, Brian Kenji Iwana
Pattern Recognit.1
2023 Corrigendum to "FETNet: Feature Erasing and Transferring Network for Scene Text Removal": Pattern Recognition Volume 140 (2023) 109531
Guangtao Lyu, Kun Liu 0027, Anna Zhu, Seiichi Uchida, Brian Kenji Iwana
Pattern Recognit.1
2022 PSSTRNet: Progressive Segmentation-Guided Scene Text Removal Network
abstract
Scene text removal (STR) is a challenging task due to the complex text fonts, colors, sizes, and background textures in scene images. However, most previous methods learn both text location and background inpainting implicitly within a single network, which weakens the text localization mecha-nism and makes a lossy background. To tackle these prob-lems, we propose a simple Progressive Segmentation-guided Scene Text Removal Network(PSSTRNet) to remove the text in the image iteratively. It contains two decoder branches, a text segmentation branch, and a text removal branch, with a shared encoder. The text segmentation branch generates text mask maps as the guidance for the regional removal branch. In each iteration, the original image, previous text removal result, and text mask are input to the network to extract the rest part of the text segments and cleaner text removal result. To get a more accurate text mask map, an update module is developed to merge the mask map in the current and previous stages. The final text removal result is obtained by adaptive fusion of results from all previous stages. A sufficient number of experiments and ablation studies conducted on the real and synthetic public datasets demonstrate our proposed method achieves state-of-the-art performance.
Guangtao Lyu, Anna Zhu
ICME1