Tianlin Hui

dblp:383/5496 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot manipulation · 48% Efficient and distributed learning · 32% Planning, search and constraint satisfaction · 16%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.722025
Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal Fusion · ACM Multimedia 2025
Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and Memory · ICCV 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-horizon planning
0.912025
Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and Memory · ICCV 2025
Robotics › Robot manipulation
mobile manipulation
0.912025
Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal Fusion · ACM Multimedia 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
TF-ATM: Training-Free Adaptive Token Merging · ACM Multimedia 2025
Machine learning › Efficient and distributed learning › model compression › token compression
token merging
0.912025
TF-ATM: Training-Free Adaptive Token Merging · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

training-free token merging · 0.9reasoning · 0.9memory · 0.9latent feature aligner · 0.9joint training · 0.9domain transfer policy · 0.9LiDAR · 0.9
YearPublicationVenuePosition
2026 LoME: LoRA-Driven Multimodal Extractor for RGB-X Vision Tasks
abstract
RGB-X multimodal vision tasks present a highly promising approach to enhancing model performance in complex visual conditions. Existing multimodal frameworks are based on either the symmetric parallel network of feature fusion or the shared network of input fusion. However, parallel networks suffer from uncontrollable parameters and imbalanced optimization across modal branches, while shared networks often lead to a lack of diversity in gradient optimization. To address these challenges, we propose the LoRA-driven Multimodal Extractor (LoME), following a comprehensive analysis of existing multimodal frameworks. The low-rank properties of modal adapters for LoME ensure controllable growth in model parameters as the number of modalities increases. The dynamic parameter fusion between adapters and the shared feature extractor decouples gradient optimization directions, effectively mitigating imbalances caused by multimodal data biases while preserving complementary features. Moreover, we employ a training strategy based on dynamic rank allocation to reduce computational overhead and enhance modal diversity expression. We validate the effectiveness and generalizability of LoME across three multimodal vision tasks. LoME achieves superior performance compared to previous state-of-the-art methods on multiple datasets. For example, on the DroneVehicle dataset, our method achieves a 10.4% improvement in accuracy compared to the SOTA method, while the parameter overhead is reduced to 23% of the previous network (44.63M). The code has been open-sourced at https://github.com/zyszxhy/LoME.
Weiying Xie, Tianlin Hui, Daixun Li, Jie Lei 0001, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and Memory
Daixun Li, Mingxiang Cao, Donglai Liu, Weiying Xie, Tianlin Hui, Lunkai Lin, Yunsong Li 0001
ICCV6
2025 Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal Fusion
abstract
Vision-Language-Action (VLA) systems are crucial for autonomous decision-making in embodied intelligence. While current systems have advanced the instruction-following capabilities, their limited spatial perception often leads to suboptimal performance for mobile manipulation tasks in unstructured environments. To address this challenge, we propose Uni-Sight, an end-to-end VLA system for robust mobile manipulation. Uni-Sight unifies decision-making, perception, and control through joint training, enabling synchronized cross-component optimization. Within the system, we introduce Latent Feature Aligner (LFA) that ensures accurate target localization by aligning multi-view data. Specifically, we develop Domain Transfer Policy (DTP), a hierarchical policy constrained by LiDAR-guided spatial priors, which ensures 3D spatial understanding with limited visual coverage. Extensive experiments on 20 real-world mobile manipulation tasks demonstrate the high task success rate and robust execution performance of Uni-Sight. Our Uni-Sight achieves a 3.04× the success rate of existing methods, and exhibits superior generalization in both long-horizon and zero-shot scenes. Code and dataset are publicly available at https://github.com/trantor2nd/Uni-Sight.
Daixun Li, Sibo He, Jiayun Tian, Weiying Xie, Mingxiang Cao, Donglai Liu, Tianlin Hui, Yunsong Li 0001
ACM Multimedia9
2025 TF-ATM: Training-Free Adaptive Token Merging
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Tianlin Hui, Jitao Ma, Leyuan Fang
ACM Multimedia5