VLDB 2026 Research / reviewers in the wild / expert
Lei Yang 0063
dblp:50/2484-63
· DBLP profile ↗
16ranked-venue papers
6as first author
7since 2021 · last 2024
0009-0007-0873-5369ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 5 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Dynamic Point Cloud Dataset for MPEG Point Cloud Compression and Performance AnalysisabstractRecent years witnessed the development in MPEG point cloud compression (PCC). However, the exploration of inter-frame coding may be impeded due to the lack of dynamic point clouds (point cloud sequences). To promote the development of PCC technology, we propose Dynamic3D , a dynamic 3D point cloud dataset with high-quality real-captured 3D persons and objects. There are several appealing properties: 1) Dynamic scenes: It contains five sequences and each sequence comprises 600 frames with temporal variation; 2) Complex content: instead of a single person or object in the existing dataset from MPEG, our established dataset contains multiple persons or both person and objects; 3) Realistic capture: the color industrial cameras and infrared cameras are used for data acquisition. This dataset provides the vast exploration space for PCC, especially the elimination of temporal redundancy. Extensive simulations are conducted on this dataset by using the reference software of MPEG G-PCC and V-PCC, i.e., (GeS-TM and TMC2), delivering observations, analysis and opportunities for the future research of PCC. Lili Zhao 0001, Qian Yin 0002, Lancao Ren, Lei Yang 0063, Chuanmin Jia, Siwei Ma 0001 |
DCC | 4 |
| 2024 | Towards Robust Visual Localization Using Multi-View Images and HD Vector MapabstractRobust and accurate localization is highly desired in intelligent driving and robotic navigation. Existing methods highly rely on feature maps and complex parameter tuning, while suffering from ineffective data association, heavy computation, high dependency on training data and low robustness. In this paper, we propose a high-robust and cost-effective visual localization system, which jointly exploits the semantic information of Bird’s-Eye-View (BEV) representation from multi-view images and the vectorized High Definition (HD) map. We formulate the visual localization as cross-modal data association issue and innovatively project the vectorized landmarks of HD map into BEV semantic map. Finally, the highly accurate vehicle’s pose can be estimated by pose optimization based on direct image alignment. Extensive simulations experimented on nuScenes dataset show that the proposed method can deliver robust and accurate localization results under various scenarios. In addition, the proposed system is convenient for large-scale deployment and has been tested on the commercial test car. Lili Zhao 0001, Zhili Liu, Qian Yin 0002, Lei Yang 0063, Meng Guo 0007 |
ICIP | 4 |
| 2024 | ISCom: Interest-Aware Semantic Communication Scheme for Point Cloud Video Streaming on Metaverse XR DevicesabstractIn the metaverse era, point cloud video (PCV) streaming on mobile XR devices is pivotal. While most current methods focus on PCV compression from traditional 3-DoF video services, emerging AI techniques extract vital semantic information, producing content resembling the original. However, these are early-stage and computationally intensive. To enhance the inference efficacy of AI-based approaches, accommodate dynamic environments, and facilitate applicability to metaverse XR devices, we present ISCom, an interest-aware semantic communication scheme for lightweight PCV streaming. ISCom is featured with a region-of-interest (ROI) selection module, a lightweight encoder-decoder training module, and a learning-based scheduler to achieve real-time PCV decoding and rendering on resource-constrained devices. ISCom’s dual-stage ROI selection provides significantly reduces data volume according to real-time interest. The lightweight PCV encoder-decoder training is tailored to resource-constrained devices and adapts to the heterogeneous computing capabilities of devices. Furthermore, We provide a deep reinforcement learning (DRL)-based scheduler to select optimal encoder-decoder model for various devices adaptivelly, considering the dynamic network environments and device computing capabilities. Our extensive experiments demonstrate that ISCom outperforms baselines on mobile devices, achieving a minimum rendering frame rate improvement of 10 FPS and up to 22 FPS. Furthermore, our method significantly reduces memory usage by 41.7% compared to the state-of-the-art AITransfer method. These results highlight the effectiveness of ISCom in enabling lightweight PCV streaming and its potential to improve immersive experiences for emerging metaverse application. Yakun Huang, Boyuan Bai, Yuanwei Zhu, Xiuquan Qiao, Xiang Su 0001, Lei Yang 0063, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 6 |
| 2023 | Attention-Based Global-Local Graph Learning for Dynamic Facial Expression Recognition
Ningwei Xie, Meng Guo 0007, Lei Yang 0063, Yafei Gong |
ICIG (1) | 4 |
| 2023 | Learning Spatial-Temporal Embeddings for Sequential Point Cloud Frame InterpolationabstractA point cloud sequence is usually acquired at a low frame rate owing to the limitations from the sensing equipment. Consequently, the immersive experience of the virtual reality might be greatly degraded. To tackle this issue, a point cloud frame interpolation process can be used to increase the frame rate of the acquired point cloud sequence by generating new frames between the consecutive ones. However, it is still challenging for deep neural networks to synthesize high-fidelity point clouds, especially for those with complex geometric details and large motion. In this paper, a novel frame interpolation network is proposed, which jointly exploits the spatial features and flows. The key success of our method lies in the developed spatial-temporal feature propagation module and temporal-aware feature-to-point mapping module. The former effectively embeds the spatial features and scene flows into a spatial-temporal feature representation (STFR). The latter generates a much improved target frame from STFR. Extensive experimental results have demonstrated that our method has achieved the best performance in most cases. Lili Zhao 0001, Zhuoqun Sun, Lancao Ren, Qian Yin 0002, Lei Yang 0063, Meng Guo 0007 |
ICIP | 5 |
| 2022 | Attention-Based Fusion of Directed Rotation Graphs for Skeleton-Based Dynamic Hand Gesture Recognition
Ningwei Xie, Lei Yang 0063, Meng Guo 0007 |
PRCV (1) | 3 |
| 2021 | Deep Multi-Patch Matching Network for Visible Thermal Person Re-IdentificationabstractVisible Thermal Person Re-Identification(VTReID) is a cross-modality retrieval problem in computer vision. Accurate VTReID is very challenging due to large modality discrepancies. In this work, we design a novelMulti-Patch Matching Network(MPMN) framework to simultaneously mitigate the heterogeneity of coarse-grained and fine-grained visual semantics. In view of cross-modality matching, we verify that aligning modality distributions of the original features is likely to suffer from the selective alignment behavior, i.e., only focuses on easiest dimensions or subspaces. Inspired by adversarial learning, we propose a newMulti-Patch Modality Alignment(MPMA) loss to jointly balance and reduce the modality discrepancies of multi-patch features by mining hard subspaces and abandoning easy subspaces. Since multi-patch features are potentially complementary to each other, the semantic correlations between different patches should be exploited during training. Motivated by knowledge distillation, we put forward a newCross-Patch Correlation Distillation(CPCD) loss to transfer the semantic knowledges across different patches. To balance multi-patch tasks, an effectivePatch-Aware Priority Attention(PAPA) method is further introduced to dynamically prioritize hard patch tasks during training. This paper experimentally demonstrates the effectiveness of the proposed methods, achieving superior performance over the state-of-the-art methods on RegDB and SYSU-MM01 datasets. Pingyu Wang, Zhicheng Zhao 0001, Yanyun Zhao, Haiying Wang 0005, Lei Yang 0063 |
IEEE Trans. Multim. | 6 |
| 2020 | Deep hard modality alignment for visible thermal person re-identification
Pingyu Wang, Zhicheng Zhao 0001, Yanyun Zhao, Lei Yang 0063 |
Pattern Recognit. Lett. | 5 |
| 2013 | Categorization of Multiple Objects in a Scene Using a Biased Sampling Strategy
Lei Yang 0063, Nanning Zheng 0001, Yang Yang 0025, Jie Yang 0001 |
Int. J. Comput. Vis. | 1 |
| 2011 | A Handwritten Character Extraction Algorithm for Multi-language Document ImageabstractIn this paper, we propose a novel method for extracting handwritten characters from multi-language document images, which may contain various types of characters, e.g. Chinese, English, Japanese or their mixture. Firstly, text patches in document image are segmented based on connected component analysis. Rules for merging connected components are chosen according to the results of language identification. Then features are extracted for each basic analysis unit-text patch. Genetic algorithm is applied for feature fusion and patch type classification. Finally, a Markov Random Field model is utilized as a post-processing step to further correct the misclassification of text patch type by considering the document context. Experimental results show that the proposed algorithm can apparently improve the performance of handwritten character extraction. Yonghong Song, Guilin Xiao, Yuanlin Zhang 0001, Lei Yang 0063, Liuliu Zhao |
ICDAR | 4 |
| 2011 | A unified context assessing model for object categorization
Lei Yang 0063, Nanning Zheng 0001, Jie Yang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2009 | Categorization of Multiple Objects in a Scene without Semantic Segmentation
Lei Yang 0063, Nanning Zheng 0001, Yang Yang 0066, Jie Yang 0001 |
ACCV (1) | 1 |
| 2009 | A biased sampling strategy for object categorizationabstractIn this paper, we present a biased sampling strategy for object class modeling, which can effectively circumvent the scene matching problem commonly encountered in statistical image-based object categorization. The method optimally combines the bottom-up, biologically inspired saliency information with loose, top-down class prior information to form a probabilistic distribution for feature sampling. When sampling over different positions and scales of patches, the weak spatial coherency is preserved by a segment-based analysis. We evaluate the proposed sampling strategy within the bag-of-features (BoF) object categorization framework on three public data sets. Our technique outperforms other state-of-the-art sampling technologies, and leads to a better performance in object categorization on VOC2008 dataset. Lei Yang 0063, Nanning Zheng 0001, Jie Yang 0001 |
ICCV | 1 |
| 2009 | PFID: Pittsburgh fast-food image datasetabstractWe introduce the first visual dataset of fast foods with a total of 4,545 still images, 606 stereo pairs, 303 3600videos for structure from motion, and 27 privacy-preserving videos of eating events of volunteers. This work was motivated by research on fast food recognition for dietary assessment. The data was collected by obtaining three instances of 101 foods from 11 popular fast food chains, and capturing images and videos in both restaurant conditions and a controlled lab setting. We benchmark the dataset using two standard approaches, color histogram and bag of SIFT features in conjunction with a discriminative classifier. Our dataset and the benchmarks are designed to stimulate research in this area and will be released freely to the research community. Kapil Dhingra, Lei Yang 0063, Rahul Sukthankar, Jie Yang 0001 |
ICIP | 4 |
| 2009 | A new tracking method for small infrared targetsabstractWe report on a new approach to tracking small infrared targets. The method improves on existing target trackers by combining mean-shift tracker with Kalman filtering and by updating the tracking parameters through the measurement of the complexity of the target region. We have further developed a nonlinear algorithm to improve the robustness of the traditional mean-shift tracker for small infrared targets. Experimental results demonstrate a superior performance of our method compared to existing target trackers, particularly in the environment of strong measurement noise and large variation of illumination. Lei Yang 0063, Weiping Lu, Jie Yang 0001 |
ICIP | 1 |
| 2008 | Layered object categorizationabstractIn this paper, we propose a novel framework of object categorization, namely layered object categorization, which takes advantage of hierarchical category information and performs object categorization at different levels. The proposed hierarchical structure of object categories is built bottom-up and top-down simultaneously accordingly to cognitive rules. First, part-based models are learnt to evaluate structure similarities at the basic level and objects are divided into basic categories. Then the decision cues for object categorization at different layers are optimally selected. Prior knowledge about inter-category relationships is utilized to infer objectspsila higher inclusive concept labels, while the most discriminative visual details of each category at the lower specific levels are selected automatically. We evaluate the proposed method with a hierarchical database and show promising results. The layered object categorization provides an efficient way for dynamically adapting the object categorization results to different applications. Lei Yang 0063, Jie Yang 0001, Nanning Zheng 0001, Hong Cheng 0002 |
ICPR | 1 |