VLDB 2026 Research / reviewers in the wild / expert
Hui Li 0010
dblp:66/3387-10
· DBLP profile ↗
22ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0002-8533-2084ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Three-Dimensional multi-object tracking based on geometry-driven and multi-domain grid feature aggregation
Hui Li 0010, Saiyu Li, Meifang Wang, Ying Gao 0005 |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Personalized dynamic lighting via LLM-Empowered multi-Agent reinforcement learning: From user intent to photometric control
Ye Tao 0002, Jiakuo Duan, Yongjie Du, Hui Li 0010, Xu Yu 0001, Xiaoni Wang |
Expert Syst. Appl. | 4 |
| 2026 | OV-Pro: Enhancing open-vocabulary 3D object detection by prototype contrastive distillation
Tongao Ge, Junyin Wang, Chenghu Du, Hui Li 0010, Huikai Liu, Shengwu Xiong 0001 |
Pattern Recognit. | 4 |
| 2026 | Panoptic-VSNet: Visual-semantic prior knowledge-driven multimodal 3D panoptic segmentation
Hui Li 0010, Yuang Ji |
Pattern Recognit. | 2 |
| 2026 | 3D Multi-Object Tracking Driven by Multi-Level Association and Intelligent Filteringabstract3D multi-object tracking has been extensively applied in areas such as autonomous driving, uncrewed aerial vehicles, and robots. However, existing 3D multi-object tracking methods still face challenges including inaccurately fitting the true motion of objects, insufficient utilization of trajectory information, and trajectory drift. To address these issues, we propose a 3D multi-object tracking framework driven by multi-level association and intelligent filtering. We design an adaptive prediction module for state estimation that reduces noise and error in prediction information through estimation and smoothing processes, thereby enhancing the stability and accuracy of state prediction. Subsequently, a multi-level trajectory integration-guided association strategy is introduced. This strategy integrates detection results, trajectory data, and potential real-world states of the objects. It minimizes incorrect associations and identity switches, thus achieving more accurate and robust data association. Finally, we propose a quality-aware intelligent filtering module for trajectory correction. This module assesses the quality of all matched detection-trajectory pairs and applies regression correction to low-quality drift detections, effectively reducing trajectory fragmentization and identity switches. On the KITTI test dataset, our method achieves 80.64% HOTA and 53.68% HOTA for the car and pedestrian categories, respectively, while reducing identity switches to 50 and 97. These results demonstrate that, compared with existing methods, our method delivers superior tracking performance and exhibits stronger robustness in handling challenges such as large inter-frame displacements, long-term occlusions, and identity switches. Hengyuan Liu, Zhong Chen 0003, Zhenzhen Du, Hui Li 0010, Xiaoxue Ai |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | Three-Dimensional Multiobject Tracking Based on Voxel Masking Encoder and Deep Hashing ParadigmabstractIn autonomous driving, accurate 3-D multiobject tracking (MOT) plays a key role in ensuring vehicle safety. However, due to the complexity of the environment, existing methods still face many challenges when dealing with long-distance objects, partial occlusions, and interference from similar categories. To tackle these challenges, we propose a 3-D MOT framework based on a voxel masking encoder (VME) and a deep hashing paradigm (DHP). We introduce a masking strategy that processes voxel features from near to far while maintaining feature sparsity, effectively capturing global contextual information between spatial features. Simultaneously, DHP is utilized to generate image hash codes and compute their hamming distance from the category hash codes. This process effectively distinguishes between object categories and thus avoids cross-category object dissociation. In addition, we propose a distance optimization matching (DOM) method that uses geometric dimensions and spatial distances to build a cost matrix, achieving more efficient and precise object associations. Results from experiments conducted on the KITTI dataset reveal that our framework delivers outstanding tracking performance, surpassing other advanced methods in tracking accuracy. The code is released at https://github.com/lsy-collab/VD-MOT. Saiyu Li, Zhong Chen 0003, Hui Li 0010, Ye Tao 0002, Ying Gao 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | SiQA: A Large Multi-Modal Question Answering Model for Structured Images Based on RAGabstractExisting Large Multimodal Models (LMMs) demonstrate excellent performance in handling visual tasks in everyday scenarios. However, they still face challenges in understanding structured images, such as flowcharts and organizational charts, which are characterized by text-rich and complex hierarchical components. In this paper, we propose SiQA, a knowledge construction and Retrieval-Augmented Generation(RAG)-based multimodal Question-Answering model designed for Structured Images. SiQA operates in three stages: Knowledge Graph (KG) generation, retrieval-augmented, and answer generation. First, a KG representing the semantics of the structured images is generated through component analysis. We then performed similarity retrieval between the KG and queries, using a node-first algorithm to construct the most relevant subgraph. Finally, after performing an encoding alignment on the multimodal information, it is fed into the LLM to generate the answer. Additionally, we introduce a new dataset, OCQA1, which includes 5,112 questions derived from 1,000 Organizational Charts. We evaluated SiQA’s structured image detection and question-answering capabilities on the FD-DETR (a flowchart dataset) and SCQA, and verified its effectiveness and strong generalization ability through comparisons with existing state-of-the-art (SOTA) methods. Jiawang Liu, Ye Tao 0002, Fei Wang 0077, Hui Li 0010, Xiugong Qin |
ICASSP | 4 |
| 2025 | Multimodal feature adaptive fusion for anchor-free 3D object detection
Yanli Wu, Junyin Wang, Hui Li 0010, Xiaoxue Ai |
Appl. Intell. | 3 |
| 2025 | Trajectory prediction based on grouped spatial-temporal encoder
Ying Gao 0005, Hui Li 0010, Xiaoya Liu, Qinghua Lin |
Frontiers Comput. Sci. | 3 |
| 2025 | Synergistic-aware cascaded association and trajectory refinement for multi-object tracking
Hui Li 0010, Su Qin, Saiyu Li, Ying Gao 0005, Yanli Wu |
Image Vis. Comput. | 1 |
| 2025 | Group commonality graph: Multimodal pedestrian trajectory prediction via deep group features
Ying Gao 0005, Hui Li 0010, Xiaoya Liu, Qinghua Lin |
Pattern Recognit. Lett. | 3 |
| 2025 | MCCA-MOT: Multimodal Collaboration-Guided Cascade Association Network for 3D Multi-Object Trackingabstract3D multi-object tracking is an important component of autonomous driving technology. Recent 3D multi-object tracking methods still suffer from issues such as information loss during the fusion of multimodal features, weak discriminative power of the association matrix, and poor robustness of single similarity measure. To address these problems, this paper proposes a Multimodal Collaboration-guided Cascade Association network for 3D multi-object tracking (MCCA-MOT). We design a point cloud feature adaptive diffusion fusion module. This module utilizes inverse distance weighting aggregation diffusion technology to address the issue of information loss during the feature fusion process. This enhances the tracking performance of small objects. Secondly, we propose a dynamic sampling feature cooperative fusion module. This module performs fine-grained local-global feature cooperative fusion based on dynamic sampling, enhancing the distinctiveness of object features. It improves the tracking capability of occluded objects. Finally, in the multi-similarity measure-driven cascading association module, we construct a more discriminative association matrix using multiple types of information and design a cascading strategy. This strategy applies different similarity measures at different association stages for objects with ambiguous features. This reduces identity switches and trajectory fragmentation. Extensive experiments on the KITTI dataset demonstrate the superiority of our method in various performance metrics. Our detailed implementations can be obtained athttps://github.com/yuanfuture/MCCA-MOT. Hui Li 0010, Hengyuan Liu, Zhenzhen Du, Zhong Chen 0003, Ye Tao 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Secrecy Performance Intelligent Prediction for Mobile Vehicular Networks: An DI-CNN ApproachabstractThe rapid expansion of Internet of Vehicles (IoV) networks has facilitated high throughput and reliable vehicular communications. Mobile vehicular networks face the challenges: diversification of network equipment, user mobility, and the broadcast nature of wireless channels, so physical layer security modeling of IoV communication systems has become important. The complexity of wireless communication channels makes real-time prediction of secrecy performance challenging. This paper presents an analysis of secrecy performance for mobile vehicular networks. To ensure data secure transmission, we have employed the decode-and-forward (DF) relaying scheme. The signal-to-noise ratio (SNR) of the effective end-to-end link is employed to obtain the mathematical expression results, which can evaluate the secrecy performance. The theoretical secrecy performance is confirmed via simulation. Then, we design a dense-inception convolution neural network (DI-CNN) model, and propose a DI-CNN-based intelligent prediction algorithm.Transformer, ShuffleNetV2, RegNet and YOLOv5 methods are employed to analyze the performance of DI-CNN algorithm. It is shown that the DI-CNN approach has a prediction accuracy that is 48.8% better than Transformer. Lingwei Xu, Huihui Tang, Hui Li 0010, Xingwang Li 0001, T. Aaron Gulliver, Khoa N. Le |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | DLFusion: Painting-Depth Augmenting-LiDAR for Multimodal Fusion 3D Object DetectionabstractSurround-view cameras combined with image depth transformation to 3D feature space and fusion with point cloud features are highly regarded. The transformation of 2D features into 3D feature space by means of predefined sampling points and depth distribution happens throughout the scene, and this process generates a large number of redundant features. In addition, multimodal feature fusion unified in 3D space often happens in the previous step of the downstream task, ignoring the interactive fusion between different scales. To this end, we design a new framework, focusing on the design that can give 3D geometric perception information to images and unify them into voxel space to accomplish multi-scale interactive fusion, and we mitigate feature alignment between modal features by geometric relationships between voxel features. The method has two main designs. First, a Segmentation-guided Image View Transformation module is used to accurately transform the pixel region containing the object into a 3D pseudo-point voxel space with the help of a depth distribution. This allows subsequent feature fusion to be performed in a unified voxel feature. Secondly, a Voxel-centric Consistent Fusion module is used to alleviate the errors caused by depth estimation, as well as to achieve better feature fusion between unified modalities. Through extensive experiments on the KITTI and nuScenes datasets, we validate the effectiveness of our camera-LIDAR fusion method. Our proposed approach shows competitive performance on both datasets and outperforms state-of-the-art methods in certain classes of 3D object detection benchmarks. https://github.com/no-Name128/DLFusion [code release] Junyin Wang, Chenghu Du, Hui Li 0010, Shengwu Xiong 0001 |
ACM Multimedia | 3 |
| 2023 | Multi-object tracking via deep feature fusion and association analysis
Hui Li 0010, Xiaoguo Liang, Yongfeng Yuan, Yuanzhi Cheng, Guanglei Zhang, Shinichi Tamura |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | Multiobject Tracking via Discriminative Embeddings for the Internet of ThingsabstractMultiobject tracking (MOT) technology can be deployed to the Internet of Things (IoT) devices to enhance the security and reliability of some video analysis applications, such as video surveillance and intelligent security system. However, since the IoT devices with limited computing capacity and storage, most existing MOT methods are difficult to deploy to IoT devices and exhibit poor tracking robustness in scenes with frequent occlusions, severe crowded, and scale variations. To alleviate the aforementioned issues, we propose a regression-based online MOT method. First, an object-aware embedding extraction module (OAEM) is designed to extract preliminary discriminative embedding of the object in the current frame. Then, an embedding aggregation module (EAM) is proposed to obtain high-quality aggregated embedding in temporal. Finally, the temporal embedding combined with the preliminary embedding extracted from the current frame to obtain a refined embedding feature for subsequent position prediction and association. Importantly, we achieve a beneficial interaction between embedding extraction, position prediction and association task. The proposed method does not suffer from significant memory consumption. Therefore, our method is a potential solution for intelligent video analysis on IoT devices. To evaluate the proposed method, we have conducted numerous experiments on the MOT16, MOT17, and MOT20 benchmark data sets. Results demonstrate that the proposed method can provide a more robust tracking performance compared to other optimal methods. Hui Li 0010, Xiaoguo Liang, Lingwei Xu, T. Aaron Gulliver |
IEEE Internet Things J. | 1 |
| 2022 | Cross-domain recommendation based on latent factor alignment
Xu Yu 0001, Qiang Hu 0002, Hui Li 0010, Junwei Du |
Neural Comput. Appl. | 3 |
| 2022 | GSTA: Pedestrian trajectory prediction based on global spatio-temporal association of graph attention network
Yun Liu 0007, Hui Li 0010, Ye Tao 0002 |
Pattern Recognit. Lett. | 3 |
| 2021 | Efficient and accurate object detection for 3D point clouds in intelligent visual internet of things
Hui Li 0010, Junyin Wang, Lingwei Xu, Ye Tao 0002 |
Multim. Tools Appl. | 1 |
| 2021 | Object detection method based on global feature augmentation and adaptive regression in IoT
Hui Li 0010, Lingwei Xu, Junyin Wang |
Neural Comput. Appl. | 1 |
| 2021 | QoS intelligent prediction for mobile video networks: a GR approach
Lingwei Xu, Han Wang 0005, Hui Li 0010, Wenzhong Lin, T. Aaron Gulliver |
Neural Comput. Appl. | 3 |
| 2015 | Pedestrian detection algorithm based on video sequences and laser point cloud
Hui Li 0010, Yun Liu 0007, Shengwu Xiong 0001, Lin Wang 0019 |
Frontiers Comput. Sci. | 1 |