Huanlin Gao

dblp:420/4865 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 51% Generative modeling · 22% Efficient and distributed learning · 22%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › image-text retrieval
composed image retrieval
1.012026
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment · AAAI 2026
Computer vision › Vision and language › vision-language model
contrastive vision-language model
1.012026
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment · AAAI 2026
Computer vision › Vision and language
cross-modal alignment
1.012026
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment · AAAI 2026
Computer vision › Vision and language
image-text retrieval
1.012026
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment · AAAI 2026
Machine learning › Efficient and distributed learning › inference acceleration
caching
0.912025
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation · NeurIPS 2025
Machine learning › Generative modeling › video generation
diffusion-based video generation
0.912025
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation · NeurIPS 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

monotonicity-aware contrastive loss · 1.0in-batch PCA · 1.0hierarchical decomposition · 1.0lexicographic minimax path optimization · 0.9directed graph with error-weighted edges · 0.9
YearPublicationVenuePosition
2026 HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
abstract
Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-form descriptions. In particular, they fail to capture two essential properties of language: semantic hierarchy, which reflects the multi-level compositional structure of text, and semantic monotonicity, where richer descriptions should result in stronger alignment with visual content. To address these limitations, we propose HiMo-CLIP, a representation-level framework that enhances CLIP-style models without modifying the encoder architecture. HiMo-CLIP introduces two key components: a hierarchical decomposition (HiDe) module that extracts latent semantic components from long-form text via in-batch PCA, enabling flexible, batch-aware alignment across different semantic granularities, and a monotonicity-aware contrastive loss (MoLo) that jointly aligns global and component-level representations, encouraging the model to internalize semantic ordering and alignment strength as a function of textual completeness. These components work together to produce structured, cognitively aligned cross-modal representations. Experiments on multiple image-text retrieval benchmarks show that HiMo-CLIP consistently outperforms strong baselines, particularly under long or compositional descriptions.
Ruijia Wu, Fei Shen 0004, Shaoan Zhao, Qiang Hui, Huanlin Gao, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
AAAI6
2025 LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
abstract
We present LeMiCa, a training-free and efficient acceleration framework for diffusion-based video generation. While existing caching strategies primarily focus on reducing local heuristic errors, they often overlook the accumulation of global errors, leading to noticeable content degradation between accelerated and original videos. To address this issue, we formulate cache scheduling as a directed graph with error-weighted edges and introduce a Lexicographic Minimax Path Optimization strategy that explicitly bounds the worst-case path error. This approach substantially improves the consistency of global content and style across generated frames. Extensive experiments on multiple text-to-video benchmarks demonstrate that LeMiCa delivers dual improvements in both inference speed and generation quality. Notably, our method achieves a 2.9× speedup on the Latte model and reaches an LPIPS score of 0.05 on Open-Sora, outperforming prior caching techniques. Importantly, these gains come with minimal perceptual quality degradation, making LeMiCa a robust and generalizable paradigm for accelerating diffusion-based video generation. We believe this approach can serve as a strong foundation for future research on efficient and reliable video synthesis.
Huanlin Gao, Fuyuan Shi, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
NeurIPS1