EDBT 2026 Demo / reviewers in the wild / expert
Qian Bao
dblp:48/505
· DBLP profile ↗
20ranked-venue papers
7as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-authorArtificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Learning Monocular Regression of 3D People in Crowds via Scene-Aware Blending and De-OcclusionabstractIn this study, we address the challenge of estimating 3D body pose, shape, and depth relationships from single RGB images in crowded scenes. The difficulty lies in the limited availability of in-the-wild training samples, which feature densely populated scenes. To mitigate this issue, we introduce a synthesis-based approach that fuses multiple human samples into a single composite scene. Our innovative scene-aware blending technique maintains human-scene consistency by positioning individuals within plausible locations and adjusting their scales to conform to 3D settings. Furthermore, our method enables flexible per-subject occlusion management during the blending process, bolstering the robustness of 3D human body representations through a novel de-occlusion training scheme. We present a one-stage model, CBD, designed to learn monocular regression of 3D people in crowds by leveraging blending and de-occlusion techniques. Our quantitative and qualitative evaluations on four benchmark datasets reveal that CBD surpasses existing state-of-the-art approaches in terms of 3D human pose and mesh regression accuracy, thereby establishing it as a promising solution for monocular 3D human mesh recovery in densely populated scenes. Yu Sun 0030, Lubing Xu, Qian Bao, Wu Liu 0005, Wenpeng Gao, Yili Fu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | TRACE: 5D Temporal Regression of Avatars with Dynamic Cameras in 3D EnvironmentsabstractAlthough the estimation of 3D human pose and shape (HPS) is rapidly progressing, current methods still cannot reliably estimate moving humans in global coordinates, which is critical for many applications. This is particularly challenging when the camera is also moving, entangling human and camera motion. To address these issues, we adopt a novel 5D representation (space, time, and identity) that enables end-to-end reasoning about people in scenes. Our method, called TRACE, introduces several novel architectural components. Most importantly, it uses two new “maps” to reason about the 3D trajectory of people over time in camera, and world, coordinates. An additional memory unit enables persistent tracking of people even during long occlusions. TRACE is the first one-stage method to jointly recover and track 3D humans in global coordinates from dynamic cameras. By training it end-to-end, and using full image information, TRACE achieves state-of-the-art performance on tracking and HPS benchmarks. The code11https://www.yusun.work/TRACE/TRACE.html and dataset22https://github.com/Arthur151/DynaCam are released for research purposes. Yu Sun 0030, Qian Bao, Wu Liu 0005, Tao Mei 0001, Michael J. Black |
CVPR | 2 |
| 2023 | A dual-population based bidirectional coevolution algorithm for constrained multi-objective optimization problems
Qian Bao, Maocai Wang, Guangming Dai, Xiaoyu Chen 0002, Zhiming Song, Shuijia Li |
Expert Syst. Appl. | 1 |
| 2022 | Putting People in their Place: Monocular Regression of 3D People in DepthabstractGiven an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene contains people of very different sizes, e.g. from infants to adults. To solve this, we need several things. First, we develop a novel method to infer the poses and depth of multiple people in a single image. While previous work that estimates multiple people does so by reasoning in the image plane, our method, called BEV, adds an additional imaginary Bird's-Eye-View representation to explicitly reason about depth. BEV reasons simultaneously about body centers in the image and in depth and, by combing these, estimates 3D body position. Unlike prior work, BEV is a single-shot method that is end-to-end differentiable. Second, height varies with age, making it impossible to resolve depth without also estimating the age of people in the image. To do so, we exploit a 3D body model space that lets BEV infer shapes from infants to adults. Third, to train BEV, we need a new dataset. Specifically, we create a “Relative Human” (RH) dataset that includes age labels and relative depth relationships between the people in the images. Extensive experiments on RH and AGORA demonstrate the effectiveness of the model and training scheme. BEV out-performs existing methods on depth reasoning, child shape estimation, and robustness to occlusion. The code11https://github.com/Arthur151/ROMP and dataset22https://github.com/Arthur151/Relative_Human are released for research purposes. Yu Sun 0030, Wu Liu 0005, Qian Bao, Yili Fu 0001, Tao Mei 0001, Michael J. Black |
CVPR | 3 |
| 2022 | Genre-Conditioned Long-Term 3D Dance Generation Driven by MusicabstractDancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre. In this paper, we focus on generating long-term 3D dance from music with a specific genre. Specifically, we construct a pure transformer-based architecture to correlate motion features and music features. To utilize the genre information, we propose to embed the genre categories into the transformer decoder so that it can guide every frame. Moreover, different from previous inference schemes, we introduce the motion queries to output the dance sequence in parallel that significantly improves the efficiency. Extensive experiments on AIST++[1] dataset show that our model outperforms state-of-the-art methods with a much faster inference speed. Yuhang Huang 0006, Junjie Zhang 0002, Qian Bao, Dan Zeng 0001, Zhineng Chen, Wu Liu 0005 |
ICASSP | 4 |
| 2022 | Learning Monocular Mesh Recovery of Multiple Body Parts Via SynthesisabstractIn this paper, we focus on simultaneously recovering the 3D mesh of multiple body parts from a single RGB image. One of the main challenges is that available datasets with full-body 3D annotations are very limited. This results in poor generalization ability of existing learning-based methods. Existing optimization-based methods iteratively fit the 3D mesh to the 2d pose, which is very time-consuming. To address these limitations, we propose to integrate multiple 3D single-body-part datasets to create a highly diverse whole-body 3D motion space for learning from controllable synthetics. Compared with the learning-based approaches, the proposed method greatly alleviates the reliance on training data. Compared with the optimization-based approaches, the proposed method is a hundred times faster. Our proposed method also outperforms previous state-of-the-art methods on CMU Panoptic dataset. Yu Sun 0030, Qian Bao, Wu Liu 0005, Wenpeng Gao, Yili Fu 0001 |
ICASSP | 3 |
| 2022 | WOC: A Handy Webcam-based 3D Online ChatroomabstractWe develop WOC, a webcam-based 3D virtual online chatroom for multi-person interaction, which captures the 3D motion of users and drives their individual 3D virtual avatars in real-time. Compared to the existing wearable equipment-based solution, WOC offers convenient and low-cost 3D motion capture with a single camera. To promote the immersive chat experience, WOC provides high-fidelity virtual avatar manipulation, which also supports the user-defined characters. With the distributed data flow service, the system delivers highly synchronized motion and voice for all users. Deployed on the website and no installation required, users can freely experience the virtual online chat at https://yanch.cloud/. Chuanhang Yan, Yu Sun 0030, Qian Bao, Jinhui Pang, Wu Liu 0005, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2021 | Monocular, One-stage, Regression of Multiple 3D PeopleabstractThis paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a One-stage fashion for Multiple 3D People (termed ROMP). The approach is conceptually simple, bounding box-free, and able to learn a per-pixel representation in an end-to-end manner. Our method simultaneously predicts a Body Center heatmap and a Mesh Parameter map, which can jointly describe the 3D body mesh on the pixel level. Through a body-center-guided sampling process, the body mesh parameters of all people in the image are easily extracted from the Mesh Parameter map. Equipped with such a fine-grained representation, our one-stage framework is free of the complex multi-stage process and more robust to occlusion. Compared with state-of-the-art methods, ROMP achieves superior performance on the challenging multi-person benchmarks, including 3DPW and CMU Panoptic. Experiments on crowded/occluded datasets demonstrate the robustness under various types of occlusion. The code, released at https://github.com/Arthur151/ROMP, is the first real-time implementation of monocular multi-person 3D mesh regression. Yu Sun 0030, Qian Bao, Wu Liu 0005, Yili Fu 0001, Michael J. Black, Tao Mei 0001 |
ICCV | 2 |
| 2021 | Neural Architecture Search for Joint Human Parsing and Pose EstimationabstractHuman parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process while pose estimation is usually a regression task, it is non-trivial to extract discriminative features for both tasks while modeling their correlation in the joint learning fashion. Recent studies have shown that Neural Architecture Search (NAS) has the ability to allocate efficient feature connections for specific tasks automatically. With the spirit of NAS, we propose to search for an efficient network architecture (NPPNet) to tackle two tasks at the same time. On the one hand, to extract task-specific features for the two tasks and lay the foundation for the further searching of feature interaction, we propose to search their encoder-decoder architectures, respectively. On the other hand, to ensure two tasks fully communicate with each other, we propose to embed NAS units in both multi-scale feature interaction and high-level feature fusion to establish optimal connections between two tasks. Experimental results on both parsing and pose estimation benchmark datasets have demonstrated that the searched model achieves state-of-the-art performances on both tasks.1 Dan Zeng 0001, Yuhang Huang 0006, Qian Bao, Junjie Zhang 0002, Chi Su, Wu Liu 0005 |
ICCV | 3 |
| 2021 | Pose-Guided Tracking-by-Detection: Robust Multi-Person Pose TrackingabstractMulti-person pose tracking task aims to estimate and track person keypoints in videos. Most of the previous methods follow the general track-by-detection strategy that ignores the consistent pose information during the whole framework. Thus, they often suffer from missing detections or inaccurate human association in challenging scenes with motion blur or person occlusion. To handle those problems, we propose a pose-guided tracking-by-detection framework that fuses pose information into both video human detection and human association procedures. In the video human detection stage, we adopt the pose-guided person location prediction exploiting the temporal information to make up missing detections. Technically, pose heatmaps are utilized to cope with the person-specific intra-class distractors. Furthermore, in the human association stage, we propose an appearance discriminative model based on the hierarchical pose-guided graph convolutional networks (PoseGCN). The PoseGCN-based model exploits human structural relations to boost person representation. Extensive experiments show the superiority of our method on the challenging pose tracking benchmark. Our proposed method ranks first on the PoseTrack leaderboard.11http://posetrack.net/leaderboard.php till the submission date (22-Aug-2019) of this paper. Our code has been publicly available at https://github.com/human-centric982/PGPT. Qian Bao, Wu Liu 0005, Yuhao Cheng, Boyan Zhou, Tao Mei 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Smart Director: An Event-Driven Directing System for Live BroadcastingabstractLive video broadcasting normally requires a multitude of skills and expertise with domain knowledge to enable multi-camera productions. As the number of cameras keeps increasing, directing a live sports broadcast has now become more complicated and challenging than ever before. The broadcast directors need to be much more concentrated, responsive, and knowledgeable, during the production. To relieve the directors from their intensive efforts, we develop an innovative automated sports broadcast directing system, called Smart Director, which aims at mimicking the typical human-in-the-loop broadcasting process to automatically create near-professional broadcasting programs in real-time by using a set of advanced multi-view video analysis algorithms. Inspired by the so-called “three-event” construction of sports broadcast [ 14 ], we build our system with an event-driven pipeline consisting of three consecutive novel components: (1) the Multi-View Event Localization to detect events by modeling multi-view correlations, (2) the Multi-View Highlight Detection to rank camera views by the visual importance for view selection, and (3) the Auto-Broadcasting Scheduler to control the production of broadcasting videos. To our best knowledge, our system is the first end-to-end automated directing system for multi-camera sports broadcasting, completely driven by the semantic understanding of sports events. It is also the first system to solve the novel problem of multi-view joint event detection by cross-view relation modeling. We conduct both objective and subjective evaluations on a real-world multi-camera soccer dataset, which demonstrate the quality of our auto-generated videos is comparable to that of the human-directed videos. Thanks to its faster response, our system is able to capture more fast-passing and short-duration events which are usually missed by human directors. Yingwei Pan, Qian Bao, Ning Zhang 0023, Ting Yao 0003, Jingen Liu, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Pose-native Network Architecture Search for Multi-person Human Pose EstimationabstractMulti-person pose estimation has achieved great progress in recent years, even though, the precise prediction for occluded and invisible hard keypoints remains challenging. Most of the human pose estimation networks are equipped with an image classification-based pose encoder for feature extraction and a handcrafted pose decoder for high-resolution representations. However, the pose encoder might be sub-optimal because of the gap between image classification and pose estimation. The widely used multi-scale feature fusion in pose decoder is still coarse and cannot provide sufficient high-resolution details for hard keypoints. Neural Architecture Search (NAS) has shown great potential in many visual tasks to automatically search efficient networks. In this work, we present the Pose-native Network Architecture Search (PoseNAS) to simultaneously design a better pose encoder and pose decoder for pose estimation. Specifically, we directly search a data-oriented pose encoder with stacked searchable cells, which can provide an optimum feature extractor for the pose specific task. In the pose decoder, we exploit scale-adaptive fusion cells to promote rich information exchange across the multi-scale feature maps. Meanwhile, the pose decoder adopts a Fusion-and-Enhancement manner to progressively boost the high-resolution representations that are non-trivial for the precious prediction of hard keypoints. With the exquisitely designed search space and search strategy, PoseNAS can simultaneously search all modules in an end-to-end manner. PoseNAS achieves state-of-the-art performance on three public datasets, MPII, COCO, and PoseTrack, with small-scale parameters compared with the existing methods. Our best model obtains 76.7% mAP and 75.9% mAP on the COCO validation set and test set with only 33.6M parameters. Code and implementation are available at https://github.com/for-code0216/PoseNAS. Qian Bao, Wu Liu 0005, Ling-Yu Duan, Tao Mei 0001 |
ACM Multimedia | 1 |
| 2019 | POINet: Pose-Guided Ovonic Insight Network for Multi-Person Pose TrackingabstractMulti-person pose tracking aims to jointly estimate and track multi-person keypoints in the unconstrained videos. The most popular solution to this task follows the tracking-by-detection strategy that relies on human detection and data association. While human detection has been boosted by deep learning, existing works mainly exploit several separated stages with hand-crafted metrics to realize data association, leading to great uncertainty and feeble adaption in complex scenes. To handle these problems, we propose an end-to-end pose-guided ovonic insight network (POINet) for the data association in multi-person pose tracking, which jointly learns feature extraction, similarity estimation, and identity assignment. Specifically, we design a pose-guided representation network to integrate pose information into hierarchical convolutional features, generating a pose-aligned person representation for person, which helps handle partial occlusions. Moreover, we propose an ovonic insight network to adaptively encode the cross-frame identity transformation, which can cope with the tough tracking cases of person leaving and entering the scene. In general, the proposed POINet provides a new insight to realize multi-person pose tracking in an end-to-end fashion. Extensive experiments conducted on the PoseTrack benchmark demonstrate that our POINet outperforms the state-of-the-art methods. Weijian Ruan, Wu Liu 0005, Qian Bao, Jun Chen 0001, Yuhao Cheng, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2019 | RotateView: A Video Composition System for Interactive Product DisplayabstractProduct display in e-commerce commonly uses static images or videos. In this paper, we design a novel video composition system “RotateView” which displays real-world products videos on the mobile phone in an interactive and 3D-like way. The system estimates the rotation direction of products using optical flow, segments products using an improved version of online adaptive video object segmentation, and adjusts motion and color for composition. The rotation direction estimation methods are evaluated on a video dataset of 1500 rotated videos that are uploaded by sellers, and on our collected real-world video object segmentation (RVOS) dataset. On our RVOS dataset, experimental results show an improvement in segmentation over the state of the art. Moreover, the proposed algorithms for motion and color adjustment are evaluated using extensive user studies. The RVOS dataset can be downloaded viahttps://realvos.github.io/for further studies and code will be made available. Shan An, Si Liu 0001, Zhibiao Huang, Guangfu Che, Qian Bao, Zhaoqi Zhu, Dennis Z. Weng |
IEEE Trans. Multim. | 5 |
| 2017 | Holographic SAR Tomography Image Reconstruction by Combination of Adaptive Imaging and Sparse Bayesian InferenceabstractIn this letter, we propose an imaging algorithm for the holographic synthetic aperture radar tomography in the circumstance of sparse and nonuniform elevation circular passes. Considering the anisotropic behavior of scatterers and the off-grid effect of sparse signal recovery, the algorithm combines the 2-D adaptive imaging method for circular SAR and the sparse Bayesian inference-based method for elevation reconstruction. For each circular pass, the azimuth-range 2-D image can be formed by the adaptive imaging method, which depends on the preretrieved maximum azimuth response angle and the azimuth persistence width. To deal with the off-grid effect in elevation reconstruction, which is caused by the deviation between the true scatterers and the discretized imaging grids, the off-grid sparse Bayesian inference method jointly estimates the scatterers and elevation off-grid error by applying their hierarchical priors. Compared with the conventional compressive sensing method that does not concern the off-grid effect, the proposed algorithm can provide more accurate 3-D reconstruction for pointlike targets, which is verified by the real-data experiments. Qian Bao, Yun Lin 0002, Wen Hong, Wenjie Shen, Xueming Peng |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | Gridless sparse recovery methods for DLSLA 3-D SAR crosstrack reconstructionabstractDownward looking sparse linear array three-dimensional synthetic aperture radar (DLSLA 3-D SAR) can obtain 3-D scene properties and has broad application prospects. However, the reconstruction of cross-track dimension usually suffers from incomplete observation, which is caused by the non-uniformly and sparsely distributed virtual antenna phase centers. By formulating the cross-track reconstruction into the problem of sparse signal recovery, we introduce two kinds of gridless sparse recovery (GL-SR) methods to DLSLA 3-D SAR cross-track imaging, i.e., atomic norm minimization (ANM) and gridless SPICE (GLS). Compared with the conventional grid-based sparse recovery (GB-SR) methods, which assume that the scatterers are exactly on the discretized grids, the GL-SR methods can avoid the off-grid effect. Experiments compare the performance of GB-SR and GL-SR methods for DLSLA 3-D SAR cross-track reconstruction. Qian Bao, Wen Hong, Yun Lin 0002, Bingchen Zhang, Weixian Tan |
IGARSS | 1 |
| 2016 | DLSLA 3-D SAR imaging algorithm for off-grid targets based on pseudo-polar formatting and atomic norm minimization
Qian Bao, Kuoye Han, Xueming Peng, Wen Hong, Bingchen Zhang, Weixian Tan |
Sci. China Inf. Sci. | 1 |
| 2016 | DLSLA 3-D SAR Imaging Based on Reweighted Gridless Sparse Recovery MethodabstractDownward-looking sparse-linear-array 3-D synthetic aperture radar (DLSLA 3-D SAR) cross-track reconstruction usually suffers from incomplete observation and limited resolution. The incomplete observation is caused by the sparse and nonuniform distribution of the equivalent antenna phase centers (APCs) due to the array elements' installation location restriction, loss, or deviation. Sparse recovery methods provide a solution with improved resolution from the incomplete observation for the 3-D imaging scene that behaves with spatial sparsity. However, conventional grid-based sparse recovery (GB-SR) methods are under the assumption that the scatterers are located on the discretized grids; otherwise, the off-grid effect or basis mismatch problem will occur. In this letter, we propose a reweighted scheme-based gridless sparse recovery (GL-SR) method, i.e., reweighted gridless sparse iterative covariance-based estimation (RGLS), for DLSLA 3-D SAR cross-track imaging. The proposed method possesses the merits of gridless SPICE (GLS), i.e., free of off-grid effect and user parameters, and has a statistically more appealing property than GLS by adopting the reweighted scheme. As seen from the experiments that compare the performance of GB-SR and GL-SR methods for DLSLA 3-D SAR cross-track reconstruction, the proposed method performs outstandingly under the circumstance of sparse and nonuniform APCs' distribution. Qian Bao, Xueming Peng, Zhirui Wang 0003, Yun Lin 0002, Wen Hong |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | Airborne DLSLA 3-D SAR Image Reconstruction by Combination of Polar Formatting and L1 RegularizationabstractAirborne downward-looking sparse linear array 3-D synthetic aperture radar (DLSLA 3-D SAR) operates downward-looking observation and obtains the 3-D microwave scattering information of the observed scene. The cross-track physical sparse linear array is often configured to obtain uniform virtual phase centers in order to adopt the frequency-domain algorithm. However, the virtual phase centers usually have to be nonuniformly and sparsely distributed due to the array elements' installation locations restricted by the airborne platform and the airborne wing tremor effect. In this state, the frequency-domain algorithm cannot be directly used. In this paper, a DLSLA 3-D SAR image reconstruction algorithm that combines polar formatting and L1regularization is presented. Wave propagation and along-track dimensional imaging are first finished after polar formatting and wavefront curvature phase error compensation; then, cross-track dimensional imaging is completed with the L1regularization technique. The proposed algorithm is applicable to airborne DLSLA 3-D SAR imaging under nonuniformly and sparsely distributed virtual phase centers condition. The proposed algorithm was verified by 3-D distributed scene simulation experiment (P-band circular SAR image was selected as radar cross-section input, and X-band digital elevation model of the same area was selected as the coordinate positions of the scene) and the field experiment. Image reconstruction results and image reconstruction performances, such as normalized radar cross section, height errors, and orthographic projection image grayscale distribution, are demonstrated and analyzed with different signal-to-noise ratios, different array sparsity, and the incomplete compensated residual oscillation error 3-D distributed scene simulation experiments. Simulation and field experimental results show the good performance in focusing and the robustness of the proposed algorithm. Xueming Peng, Weixian Tan, Wen Hong, Chenglong Jiang, Qian Bao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | Statistical Analysis of the Effects of Virtual Element Position Errors on Airborne Down-Looking LASAR 3-D ImagingabstractIn order to achieve 3-D imaging with an airborne down-looking linear-array synthetic aperture radar (LASAR), a uniform virtual antenna array may be obtained by aperture synthesis of the cross-track sparse multiple-input-multiple-output array. However, the actual 3-D imaging quality is unavoidably degraded by errors in the virtual element position. In this letter, we investigate the effects of these errors on the forms and the degrees of image quality degradation by decomposing the error-related stochastic processes via an orthogonal transform based on discrete Legendre polynomials. It should be noted that these analyses are helpful for designing a LASAR system and providing a reference for specifying the requisite precision of measurement devices and calibration methods. Finally, we briefly consider the use of calibration methods to eliminate the effects of errors. Kuoye Han, Qian Bao, Weixian Tan, Wen Hong |
IEEE Geosci. Remote. Sens. Lett. | 2 |