VLDB 2026 Research / reviewers in the wild / expert
Huanlin Gao
dblp:420/4865
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Vision and language · 51% Generative modeling · 22% Efficient and distributed learning · 22% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › image-text retrieval
composed image retrieval |
1.0 | 1 | 2026 | HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment · AAAI 2026 |
Computer vision › Vision and language › vision-language model
contrastive vision-language model |
1.0 | 1 | 2026 | HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment · AAAI 2026 |
Computer vision › Vision and language
cross-modal alignment |
1.0 | 1 | 2026 | HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment · AAAI 2026 |
Computer vision › Vision and language
image-text retrieval |
1.0 | 1 | 2026 | HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment · AAAI 2026 |
Machine learning › Efficient and distributed learning › inference acceleration
caching |
0.9 | 1 | 2025 | LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation · NeurIPS 2025 |
Machine learning › Generative modeling › video generation
diffusion-based video generation |
0.9 | 1 | 2025 | LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
inference acceleration |
0.9 | 1 | 2025 | LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
monotonicity-aware contrastive loss · 1.0in-batch PCA · 1.0hierarchical decomposition · 1.0lexicographic minimax path optimization · 0.9directed graph with error-weighted edges · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language AlignmentabstractContrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-form descriptions. In particular, they fail to capture two essential properties of language: semantic hierarchy, which reflects the multi-level compositional structure of text, and semantic monotonicity, where richer descriptions should result in stronger alignment with visual content. To address these limitations, we propose HiMo-CLIP, a representation-level framework that enhances CLIP-style models without modifying the encoder architecture. HiMo-CLIP introduces two key components: a hierarchical decomposition (HiDe) module that extracts latent semantic components from long-form text via in-batch PCA, enabling flexible, batch-aware alignment across different semantic granularities, and a monotonicity-aware contrastive loss (MoLo) that jointly aligns global and component-level representations, encouraging the model to internalize semantic ordering and alignment strength as a function of textual completeness. These components work together to produce structured, cognitively aligned cross-modal representations. Experiments on multiple image-text retrieval benchmarks show that HiMo-CLIP consistently outperforms strong baselines, particularly under long or compositional descriptions. Ruijia Wu, Fei Shen 0004, Shaoan Zhao, Qiang Hui, Huanlin Gao, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian |
AAAI | 6 |
| 2025 | LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video GenerationabstractWe present LeMiCa, a training-free and efficient acceleration framework for diffusion-based video generation. While existing caching strategies primarily focus on reducing local heuristic errors, they often overlook the accumulation of global errors, leading to noticeable content degradation between accelerated and original videos. To address this issue, we formulate cache scheduling as a directed graph with error-weighted edges and introduce a Lexicographic Minimax Path Optimization strategy that explicitly bounds the worst-case path error. This approach substantially improves the consistency of global content and style across generated frames. Extensive experiments on multiple text-to-video benchmarks demonstrate that LeMiCa delivers dual improvements in both inference speed and generation quality. Notably, our method achieves a 2.9× speedup on the Latte model and reaches an LPIPS score of 0.05 on Open-Sora, outperforming prior caching techniques. Importantly, these gains come with minimal perceptual quality degradation, making LeMiCa a robust and generalizable paradigm for accelerating diffusion-based video generation. We believe this approach can serve as a strong foundation for future research on efficient and reliable video synthesis. Huanlin Gao, Fuyuan Shi, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian |
NeurIPS | 1 |