VLDB 2026 Research / reviewers in the wild / expert
Feixiang Lu
dblp:220/2687
· DBLP profile ↗
14ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0003-3952-4402ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | 3D Human Pose Estimation via Non-causal Retentive Networks
Kaili Zheng, Feixiang Lu, Yihao Lv, Liangjun Zhang, Chenyi Guo, Ji Wu 0002 |
ECCV (33) | 2 |
| 2024 | P2ANet: A Large-Scale Benchmark for Dense Action Detection from Table Tennis Match Broadcasting VideosabstractWhile deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark P 2 ANet for P ing P ong- A ction detection, which consists of 2,721 video clips collected from the broadcasting videos of professional table tennis matches in World Table Tennis Championships and Olympiads. We work with a crew of table tennis professionals and referees on a specially designed annotation toolbox to obtain fine-grained action labels (in 14 classes) for every ping-pong action that appeared in the dataset, and formulate two sets of action detection problems— action localization and action recognition . We evaluate a number of commonly seen action recognition (e.g., TSM, TSN, Video SwinTransformer, and Slowfast) and action localization models (e.g., BSN, BSN++, BMN, TCANet), using P 2 ANet for both problems, under various settings. These models can only achieve 48% area under the AR-AN curve for localization and 82% top-one accuracy for recognition since the ping-pong actions are dense with fast-moving subjects but broadcasting videos are with only 25 FPS. The results confirm that P 2 ANet is still a challenging task and can be used as a special benchmark for dense action detection from videos. We invite readers to examine our dataset by visiting the following link: https://github.com/Fred1991/P2ANET . Jiang Bian 0003, Xuhong Li 0002, Tao Wang 0011, Qingzhong Wang, Feixiang Lu, Dejing Dou, Haoyi Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2023 | A Survey on Video Action Recognition in Sports: Datasets, Methods and ApplicationsabstractTo understand human behaviors, action recognition based on videos is a common approach. Compared with image-based action recognition, videos provide much more information, reducing the ambiguity of actions. In the last decade, many works focus on datasets, novel models and learning approaches have improved video action recognition to a higher level. However, there are challenges and unsolved problems, in particular in sports analytics where data collection and labeling are more sophisticated, requiring people with domain knowledge and even sport professionals to annotate data. In addition, the actions could be extremely fast and it becomes difficult to recognize them. Moreover, in team sports like football and basketball, one action could involve multiple players, and to correctly recognize them, we need to analyze all players, which is relatively complicated. In this paper, we present a survey on video action recognition for sports analytics. We introduce more than ten types of sports, including team sports, such as football, basketball, volleyball, hockey and individual sports, such as figure skating, gymnastics, table tennis, tennis, diving and badminton. Then we compare numerous existing frameworks for sports analysis to present status quo of video action recognition in both team sports and individual sports. Finally, we discuss the challenges and unsolved problems in this area and to facilitate sports analytics, we develop a toolbox using PaddlePaddle, which supports football, basketball, table tennis and figure skating action recognition. Qingzhong Wang, Jiang Bian 0003, Feixiang Lu, Dejing Dou, Haoyi Xiong |
IEEE Trans. Multim. | 5 |
| 2022 | P^3-Net: Part Mobility Parsing from Point Cloud Sequences via Learning Explicit Point CorrespondenceabstractUnderstanding an articulated 3D object with its movable parts is an essential skill for an intelligent agent. This paper presents a novel approach to parse 3D part mobility from point cloud sequences. The key innovation is learning explicit point correspondence from a raw unordered point cloud sequence. We propose a novel deep network called P^3-Net to parallelize trajectory feature extraction and point correspondence establishment, performing joint optimization between them. Specifically, we design a Match-LSTM module to reaggregate point features among different frames by a point correspondence matrix, a.k.a. the matching matrix. To obtain this matrix, an attention module is proposed to calculate the point correspondence. Moreover, we implement a Gumbel-Sinkhorn module to reduce the many-to-one relationship for better point correspondence. We conduct comprehensive evaluations on public benchmarks, including the motion dataset and the PartNet dataset. Results demonstrate that our approach outperforms SOTA methods on various 3D parsing tasks of part mobility, including motion flow prediction, motion part segmentation, and motion attribute (i.e. axis & range) estimation. Moreover, we integrate our approach into a robot perception module to validate its robustness. Yahao Shi, Feixiang Lu |
AAAI | 3 |
| 2022 | Rec2Real: Semantics-Guided Photo-Realistic Image Synthesis Using Rough Urban Reconstruction Models
Feixiang Lu, Tiancheng Xu, Liangjun Zhang |
CGI | 2 |
| 2022 | Part-Level Car Parsing and Reconstruction in Single Street View ImagesabstractPart information has been proven to be resistant to occlusions and viewpoint changes, which are main difficulties in car parsing and reconstruction. However, in the absence of datasets and approaches incorporating car parts, there are limited works that benefit from it. In this paper, we propose the first part-aware approach for joint part-level car parsing and reconstruction in single street view images. Without labor-intensive part annotations on real images, our approach simultaneously estimates pose, shape, and semantic parts of cars. There are two contributions in this paper. First, our network introduces dense part information to facilitate pose and shape estimation, which is further optimized with a novel 3D loss. To obtain part information in real images, a class-consistent method is introduced to implicitly transfer part knowledge from synthesized images. Second, we construct the first high-quality dataset containing 348 car models with physical dimensions and part annotations. Given these models, 60K synthesized images with randomized configurations are generated. Experimental results demonstrate that part knowledge can be effectively transferred with our class-consistent method, which significantly improves part segmentation performance on real street views. By fusing dense part information, our pose and shape estimation results achieve the state-of-the-art performance on the ApolloCar3D and outperform previous approaches by large margins in terms of both A3DP-Abs and A3DP-Rel. Qichuan Geng, Hong Zhang 0009, Feixiang Lu, Xinyu Huang 0001, Sen Wang 0003, Zhong Zhou, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | AutoShape: Real-Time Shape-Aware Monocular 3D Object DetectionabstractExisting deep learning-based approaches for monocular 3D object detection in autonomous driving often model the object as a rotated 3D cuboid while the object’s geometric shape has been ignored. In this work, we propose an approach for incorporating the shape-aware 2D/3D constraints into the 3D detection framework. Specifically, we employ the deep neural network to learn distinguished 2D keypoints in the 2D image domain and regress their corresponding 3D coordinates in the local 3D object coordinate first. Then the 2D/3D geometric constraints are built by these correspondences for each object to boost the detection performance. For generating the ground truth of 2D/3D keypoints, an automatic model-fitting approach has been proposed by fitting the deformed 3D object model and the object mask in the 2D image. The proposed framework has been verified on the public KITTI dataset and the experimental results demonstrate that by using additional geometrical constraints the detection performance has been significantly improved as compared to the baseline method. More importantly, the proposed framework achieves state-of-the-art performance with real time. Data and code will be available at https://github.com/zongdai/AutoShape Zongdai Liu, Dingfu Zhou, Feixiang Lu, Liangjun Zhang |
ICCV | 3 |
| 2021 | Robust 2D/3D Vehicle Parsing in Arbitrary Camera Views for CVISabstractWe present a novel approach to robustly detect and perceive vehicles in different camera views as part of a cooperative vehicle-infrastructure system (CVIS). Our formulation is designed for arbitrary camera views and makes no assumptions about intrinsic or extrinsic parameters. First, to deal with multi-view data scarcity, we propose a part-assisted novel view synthesis algorithm for data augmentation. We train a part-based texture inpainting network in a self-supervised manner. Then we render the textured model into the background image with the target 6-DoF pose. Second, to handle various camera parameters, we present a new method that produces dense mappings between image pixels and 3D points to perform robust 2D/3D vehicle parsing. Third, we build the first CVIS dataset for bench-marking, which annotates more than 1540 images (14017 instances) from real-world traffic scenarios. We combine these novel algorithms and datasets to develop a robust approach for 2D/3D vehicle parsing for CVIS. In practice, our approach outperforms SOTA methods on 2D detection, in-stance segmentation, and 6-DoF pose estimation by 3.8%, 4.3%, and 2.9%, respectively. Feixiang Lu, Zongdai Liu, Liangjun Zhang, Dinesh Manocha |
ICCV | 2 |
| 2020 | IAFA: Instance-Aware Feature Aggregation for 3D Object Detection from a Single Image
Dingfu Zhou, Xibin Song, Yuchao Dai, Junbo Yin, Feixiang Lu, Miao Liao, Liangjun Zhang |
ACCV (1) | 5 |
| 2020 | 3D Part Guided Image Editing for Fine-Grained Object UnderstandingabstractHolistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving vehicle can be success in dealing with emergency cases. However, existing visual models tackle rarely on these situations, but focus on bounding box detection. In this paper, we fill this important missing piece in autonomous driving by solving two critical issues. First, for dealing with data scarcity, we propose an effective training data generation process by fitting a 3D car model with dynamic parts to cars in real images. This allows us to directly edit the real images using the aligned 3D parts, yielding effective training data for learning robust deep neural networks (DNNs). Secondly, to benchmark the quality of 3D part understanding, we collected a large dataset in real driving scenario with cars in uncommon states (CUS), i.e. with door or trunk opened etc., which demonstrates that our trained network with edited images largely outperforms other baselines in terms of 2D detection and instance segmentation accuracy. Zongdai Liu, Feixiang Lu, Peng Wang 0001, Liangjun Zhang, Ruigang Yang |
CVPR | 2 |
| 2020 | DVI: Depth Guided Video Inpainting for Autonomous Driving
Miao Liao, Feixiang Lu, Dingfu Zhou, Wei Li 0111, Ruigang Yang |
ECCV (21) | 2 |
| 2020 | InstanceFusion: Real-time Instance-level 3D Reconstruction Using a Single RGBD CameraabstractAbstract We present InstanceFusion, a robust real‐time system to detect, segment, and reconstruct instance‐level 3D objects of indoor scenes with a hand‐held RGBD camera. It combines the strengths of deep learning and traditional SLAM techniques to produce visually compelling 3D semantic models. The key success comes from our novel segmentation scheme and the efficient instance‐level data fusion, which are both implemented on GPU. Specifically, for each incoming RGBD frame, we take the advantages of the RGBD features, the 3D point cloud, and the reconstructed model to perform instance‐level segmentation. The corresponding RGBD data along with the instance ID are then fused to the surfel‐based models. In order to sufficiently store and update these data, we design and implement a new data structure using the OpenGL Shading Language. Experimental results show that our method advances the state‐of‐the‐art (SOTA) methods in instance segmentation and data fusion by a big margin. In addition, our instance segmentation improves the precision of 3D reconstruction, especially in the loop closure. InstanceFusion system runs 20.5Hz on a consumer‐level GPU, which supports a number of augmented reality (AR) applications (e.g., 3D model registration, virtual interaction, AR map) and robot applications (e.g., navigation, manipulation, grasping). To facilitate future research and reproduce our system more easily, the source code, data, and the trained model are released on Github: https://github.com/Fancomi2017/InstanceFusion . Feixiang Lu, Haotian Peng, Xinhang Yang, Ruizhi Cao, Liangjun Zhang, Ruigang Yang |
Comput. Graph. Forum | 1 |
| 2018 | Reconstructing non-rigid object with large movement using a single depth camera
Feixiang Lu, Feng Lu 0005, Yu Zhang 0035, Xiaowu Chen 0001, Qinping Zhao |
Comput. Aided Geom. Des. | 1 |
| 2018 | Real-time 3D scene reconstruction with dynamically moving object using a single depth camera
Feixiang Lu, Yu Zhang 0035, Qinping Zhao |
Vis. Comput. | 1 |