Hualian Sheng

dblp:268/7077 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-2405-9325ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Model
Hualian Sheng, Sijia Cai, Bing Deng, Qiao Liang 0002, Wen Li 0001, Jieping Ye, Shuhang Gu
ICCV2
2025 EchoShot: Multi-Shot Portrait Video Generation
abstract
Video diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge multiple shots with identity consistency and flexible content controllability. In this work, we propose EchoShot, a native and scalable multi-shot framework for portrait customization built upon a foundation video diffusion model. To start with, we propose shot-aware position embedding mechanisms within the video diffusion transformer architecture to model inter-shot variations and establish intricate correspondence between multi-shot visual content and their textual descriptions. This simple yet effective design enables direct training on multi-shot video data without introducing additional computational overhead. To facilitate model training within multi-shot scenarios, we construct PortraitGala, a large-scale and high-fidelity human-centric video dataset featuring cross-shot identity consistency and fine-grained captions such as facial attributes, outfits, and dynamic motions. To further enhance applicability, we extend EchoShot to perform reference image-based personalized multi-shot generation and long video synthesis with infinite shot counts. Extensive evaluations demonstrate that EchoShot achieves superior identity consistency as well as attribute-level controllability in multi-shot portrait video generation. Notably, the proposed framework demonstrates potential as a foundational paradigm for general multi-shot video modeling. Project page: https://johnneywang.github.io/EchoShot-webpage.
Jiahao Wang 0004, Hualian Sheng, Sijia Cai, Weizhan Zhang, Caixia Yan, Yachuang Feng, Bing Deng, Jieping Ye
NeurIPS2
2025 CT3D++: Improving 3D Object Detection with Keypoint-Induced Channel-wise Transformer
Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Qiao Liang 0002, Minjian Zhao, Jieping Ye
Int. J. Comput. Vis.1
2024 RoScenes: A Large-Scale Multi-view 3D Dataset for Roadside Perception
Xiaosu Zhu, Hualian Sheng, Sijia Cai, Bing Deng, Shaopeng Yang, Qiao Liang 0002, Ken Chen 0005, Lianli Gao, Jingkuan Song, Jieping Ye
ECCV (41)2
2024 ARIoU: Anchor-free Rotation-decoupling IoU-based optimization for 3D object detection
Chenyiming Wen, Hualian Sheng, Ming-Min Zhao, Minjian Zhao
Neurocomputing2
2023 PDR: Progressive Depth Regularization for Monocular 3D Object Detection
abstract
Accurately predicting object depth is a key challenge in monocular 3D detection task. The perspective projection principle used by most state-of-the-art approaches demands a complex balance between the ratio-form depth estimation and 2D-3D geometric regularizations, and thus can lead to sub-optimal solutions. In this paper, we propose a novel synergistic scheme that can achieve better trade-off among these competing objectives. Our main proposal is a progressive depth regularization (PDR) architecture that splits the overall training process into three sequential depth estimation steps to gradually remove the unwanted deviations induced by the over-regularization. Specifically, our model first learns the coarse depth with the conventional perspective projection and combines the coarse-to-fine generation to reduce the search space of 2D projection height prediction. We then deactivate individual supervision on 2D projection height prediction and introduces a new auxiliary 3D physical height prediction to relax the 2D and 3D regularizations, respectively. Consequently, our PDR leads to more precise depth estimation by mitigating the inherent ambiguities in the geometric priors of perspective projection through progressive regularization relaxation. Extensive experiments on both KITTI and Rope3D benchmark show that our PDR delivers strong performance gains as compared to the previous methods.
Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Minjian Zhao, Gim Hee Lee
IEEE Trans. Circuits Syst. Video Technol.1
2022 Balanced and Hierarchical Relation Learning for One-shot Object Detection
abstract
Instance-level feature matching is significantly important to the success of modern one-shot object detectors. Re-cently, the methods based on the metric-learning paradigm have achieved an impressive process. Most of these works only measure the relations between query and target objects on a single level, resulting in suboptimal performance overall. In this paper, we introduce the balanced and hierarchical learning for our detector. The contributions are two-fold: firstly, a novel Instance-level Hierarchical Relation (IHR) module is proposed to encode the contrastive-level, salient-level, and attention-level relations simultane-ously to enhance the query-relevant similarity representation. Secondly, we notice that the batch training of the IHR module is substantially hindered by the positive-negative sample imbalance in the one-shot scenario. We then in-troduce a simple but effective Ratio-Preserving Loss (RPL) to protect the learning of rare positive samples and sup-press the effects of negative samples. Our loss can adjust the weight for each sample adaptively, ensuring the desired positive-negative ratio consistency and boosting query-related IHR learning. Extensive experiments show that our method outperforms the state-of-the-art method by 1.6% and 1.3% on PASCAL VOC and MS COCO datasets for unseen classes, respectively. The code will be available at https://github.com/hero-y/BHRL.
Hanqing Yang 0002, Sijia Cai, Hualian Sheng, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Yu Zhang 0018
CVPR3
2022 Rethinking IoU-based Optimization for Single-stage 3D Object Detection
Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Minjian Zhao, Gim Hee Lee
ECCV (9)1
2021 Improving 3D Object Detection with Channel-wise Transformer
abstract
Though 3D object detection from point clouds has achieved rapid progress in recent years, the lack of flexible and high-performance proposal refinement remains a great hurdle for existing state-of-the-art two-stage detectors. Previous works on refining 3D proposals have relied on human-designed components such as keypoints sampling, set abstraction and multi-scale feature fusion to produce powerful 3D object representations. Such methods, however, have limited ability to capture rich contextual dependencies among points. In this paper, we leverage the high-quality region proposal network and a Channel-wise Transformer architecture to constitute our two-stage 3D object detection framework (CT3D) with minimal hand-crafted design. The proposed CT3D simultaneously performs proposal-aware embedding and channel-wise context aggregation for the point features within each proposal. Specifically, CT3D uses proposal’s keypoints for spatial contextual modelling and learns attention propagation in the encoding module, mapping the proposal to point embeddings. Next, a new channel-wise decoding module enriches the query-key interaction via channel-wise re-weighting to effectively merge multi-level contexts, which contributes to more accurate object predictions. Extensive experiments demonstrate that our CT3D method has superior performance and excellent scalability. Remarkably, CT3D achieves the AP of 81.77% in the moderate car category on the KITTI test 3D detection benchmark, outperforms state-of-the-art 3D detectors.
Hualian Sheng, Sijia Cai, Yuan Liu 0017, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Minjian Zhao
ICCV1
2020 Energy Efficiency Optimization for Beamspace Massive MIMO Systems with Low-Resolution ADCs
abstract
In this article, we propose a sparse hybrid combining (SHC) scheme for the uplink transmission of beamspace massive multiple-input multiple-output (MIMO) system with low-resolution analog to digital converters (LADCs), to alleviate the performance bottleneck caused by the multi-user interference and quantization noise, with reduced hardware cost and power consumption. To this end, we formulate the optimization of the proposed SHC scheme as a system energy efficiency maximization problem under some practical constraints. The resulting problem contains the highly coupled nonconvex objective function, as well as the discrete binary constraints. By exploiting some fractional programming (FP) techniques and introducing auxiliary variables, we first recast the original challenging problem into a more tractable yet equivalent form. We then develop an efficient double-loop iterative algorithm based on the penalty dual decomposition (PDD) method to find its local stationary solutions. Finally, simulation results verify the effectiveness of the proposed SHC scheme by numerical examples in terms of the achieved system energy efficiency.
Hualian Sheng, Xihan Chen, Kaiming Shen, Xiongfei Zhai, An Liu 0001, Minjian Zhao
WCNC1