VLDB 2026 Research / reviewers in the wild / expert
Tianyuan Liang
dblp:425/0742
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0001-7576-6356ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Video understanding and tracking · 38% Efficient and distributed learning · 38% Language models and text generation · 19% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization |
1.0 | 1 | 2026 | Resonating with RoPE: Spectral Quantization for High-Fidelity Key Cache Compression · ACL (1) 2026 |
Natural language and speech › Language models and text generation › language modeling
long-context language modeling |
1.0 | 1 | 2026 | Resonating with RoPE: Spectral Quantization for High-Fidelity Key Cache Compression · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | Resonating with RoPE: Spectral Quantization for High-Fidelity Key Cache Compression · ACL (1) 2026 |
Computer vision › Video understanding and tracking › action detection
temporal action localization |
1.0 | 1 | 2026 | Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action Localization · AAAI 2026 |
Computer vision › Video understanding and tracking › action detection › temporal action localization
zero-shot activity detection |
1.0 | 1 | 2026 | Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action Localization · AAAI 2026 |
Multimedia analysis and retrieval
cross-modal retrieval |
0.9 | 1 | 2025 | FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person Search · ACM Multimedia 2025 |
Multimedia analysis and retrieval › cross-modal retrieval
image-text retrieval |
0.9 | 1 | 2025 | FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person Search · ACM Multimedia 2025 |
Multimedia analysis and retrieval › visual search › person search
text-based person search |
0.9 | 1 | 2025 | FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person Search · ACM Multimedia 2025 |
Computer vision › Face, body and person analysis
person re-identification |
0.3 | 1 | 2025 | FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person Search · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 1.7large language model noise filtering · 1.7dual-template caption generation · 1.7adaptive curriculum learning · 1.7triton kernels · 1.0spectral quantization · 1.0multimodal large language model · 1.0large language model · 1.0discrete cosine transform · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action LocalizationabstractCurrent Zero-Shot Temporal Action Localization (ZSTAL) methods, whether training-based or training-free ones, still predominantly rely on a single, unified query to localize an entire action. This unified representation is fundamentally ill-suited for complex real-world activities, as it fails to capture their internal compositional structure and adapt to dynamic, multi-stage variations across videos. To address this, we regard ZSTAL as a compositional reasoning task and introduce CASCADE, a Context-Aware Staged Action DEcomposition framework. Inspired by the human cognitive process of perceiving context, decomposing events, and reconstructing instances, CASCADE follows a training-free pipeline. It first perceives the video's context by leveraging a Multimodal Large Language Model (MLLM) to both filter out irrelevant actions and then generate a rich, video-specific caption for each action present in the video. An LLM then decomposes this caption into multiple, temporally ordered stages, which serve as fine-grained queries to guide the MLLM in estimating frame-level confidence scores. Recognizing that this decomposition can fragment a single action, a novel hierarchical merging logic then reconstructs complete instances by intelligently fusing these preliminary temporal segments based on their semantic progression and coherence. Extensive experiments and ablation studies on THUMOS14 and ActivityNet-1.3 show that CASCADE not only sets a new state-of-the-art among training-free methods but, most notably, significantly outperforms all prior training-based approaches on ActivityNet-1.3. Haoyu Tang 0002, Tianyuan Liang, Han Jiang 0012, Qinghai Zheng, Yupeng Hu 0003 |
AAAI | 2 |
| 2026 | Resonating with RoPE: Spectral Quantization for High-Fidelity Key Cache CompressionabstractThe linear growth of KV cache bottlenecks long-context LLMs, yet RoPE-induced oscillations complicate Key cache quantization.To address this issue, we propose SpectrumQuant, a frequency-domain framework that utilizes the Discrete Cosine Transform (DCT) to convert these oscillations into sparse spectral representations.Specifically, our pipeline integrates dominant frequency extraction, hybrid bit-width allocation, and high-frequency preemphasis to maximize fidelity while minimizing memory footprint.To eliminate computational overhead, we develop fused Triton kernels featuring deferred inverse transformation and on-chip sparse accumulation.Extensive experiments on several benchmarks confirm SpectrumQuant achieves efficient compression with performance and latency comparable to FP16 baselines. Haoyu Tang 0002, Tianyuan Liang, Yupeng Hu 0003, Weili Guan |
ACL (1) | 3 |
| 2025 | FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person SearchabstractText-Based Person Search, which aims to retrieve target pedestrian images using natural language descriptions, has garnered significant attention in multimedia research due to its potential in suspect retrieval and missing person identification. While supervised and weakly supervised methods rely on costly annotated training data, unsupervised TBPS eliminates the need for textual descriptions or identity annotations, presenting a more practical paradigm. Current unsupervised TBPS approaches face two primary challenges: 1) Predefined attribute templates for caption generation limit linguistic diversity and real-world adaptability, and 2) Threshold-based sample selection using pre-trained vision-language models (VLMs) introduces noisy pairs due to inadequate pedestrian-specific representation. To address these limitations, we propose FACE, a unified framework featuring Dual-template Caption Generation (DCG) and Adaptive Curriculum Training (ACT). The DCG module generates high-quality captions through complementary flexible-style (natural language) and fixed-style (attribute-enumerated) templates, enhanced by LLM-based noise filtering. The ACT framework progressively refines training through a self-improving loop: initial high-confidence sample selection using VLMs bootstraps the model, while evolving feature representations enable dynamic incorporation of harder samples through curriculum learning. This dual strategy achieves mutual reinforcement between caption quality and model discriminability. Extensive experiments on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets under unsupervised settings demonstrate that our framework achieves the state-of-the-art performance. Xiaoxuan Mu, Haoyu Tang 0002, Han Jiang 0012, Tianyuan Liang, Qinghai Zheng, Jihua Zhu |
ACM Multimedia | 4 |