VLDB 2026 Research / reviewers in the wild / expert
Tianlin Hui
dblp:383/5496
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Robot manipulation · 48% Efficient and distributed learning · 32% Planning, search and constraint satisfaction · 16% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
1.7 | 2 | 2025 | Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal Fusion · ACM Multimedia 2025 Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and Memory · ICCV 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-horizon planning |
0.9 | 1 | 2025 | Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and Memory · ICCV 2025 |
Robotics › Robot manipulation
mobile manipulation |
0.9 | 1 | 2025 | Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal Fusion · ACM Multimedia 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | TF-ATM: Training-Free Adaptive Token Merging · ACM Multimedia 2025 |
Machine learning › Efficient and distributed learning › model compression › token compression
token merging |
0.9 | 1 | 2025 | TF-ATM: Training-Free Adaptive Token Merging · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
training-free token merging · 0.9reasoning · 0.9memory · 0.9latent feature aligner · 0.9joint training · 0.9domain transfer policy · 0.9LiDAR · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoME: LoRA-Driven Multimodal Extractor for RGB-X Vision TasksabstractRGB-X multimodal vision tasks present a highly promising approach to enhancing model performance in complex visual conditions. Existing multimodal frameworks are based on either the symmetric parallel network of feature fusion or the shared network of input fusion. However, parallel networks suffer from uncontrollable parameters and imbalanced optimization across modal branches, while shared networks often lead to a lack of diversity in gradient optimization. To address these challenges, we propose the LoRA-driven Multimodal Extractor (LoME), following a comprehensive analysis of existing multimodal frameworks. The low-rank properties of modal adapters for LoME ensure controllable growth in model parameters as the number of modalities increases. The dynamic parameter fusion between adapters and the shared feature extractor decouples gradient optimization directions, effectively mitigating imbalances caused by multimodal data biases while preserving complementary features. Moreover, we employ a training strategy based on dynamic rank allocation to reduce computational overhead and enhance modal diversity expression. We validate the effectiveness and generalizability of LoME across three multimodal vision tasks. LoME achieves superior performance compared to previous state-of-the-art methods on multiple datasets. For example, on the DroneVehicle dataset, our method achieves a 10.4% improvement in accuracy compared to the SOTA method, while the parameter overhead is reduced to 23% of the previous network (44.63M). The code has been open-sourced at https://github.com/zyszxhy/LoME. Weiying Xie, Tianlin Hui, Daixun Li, Jie Lei 0001, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and Memory
Daixun Li, Mingxiang Cao, Donglai Liu, Weiying Xie, Tianlin Hui, Lunkai Lin, Yunsong Li 0001 |
ICCV | 6 |
| 2025 | Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal FusionabstractVision-Language-Action (VLA) systems are crucial for autonomous decision-making in embodied intelligence. While current systems have advanced the instruction-following capabilities, their limited spatial perception often leads to suboptimal performance for mobile manipulation tasks in unstructured environments. To address this challenge, we propose Uni-Sight, an end-to-end VLA system for robust mobile manipulation. Uni-Sight unifies decision-making, perception, and control through joint training, enabling synchronized cross-component optimization. Within the system, we introduce Latent Feature Aligner (LFA) that ensures accurate target localization by aligning multi-view data. Specifically, we develop Domain Transfer Policy (DTP), a hierarchical policy constrained by LiDAR-guided spatial priors, which ensures 3D spatial understanding with limited visual coverage. Extensive experiments on 20 real-world mobile manipulation tasks demonstrate the high task success rate and robust execution performance of Uni-Sight. Our Uni-Sight achieves a 3.04× the success rate of existing methods, and exhibits superior generalization in both long-horizon and zero-shot scenes. Code and dataset are publicly available at https://github.com/trantor2nd/Uni-Sight. Daixun Li, Sibo He, Jiayun Tian, Weiying Xie, Mingxiang Cao, Donglai Liu, Tianlin Hui, Yunsong Li 0001 |
ACM Multimedia | 9 |
| 2025 | TF-ATM: Training-Free Adaptive Token Merging
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Tianlin Hui, Jitao Ma, Leyuan Fang |
ACM Multimedia | 5 |