EDBT 2026 Demo / reviewers in the wild / expert
Yubin Hu 0001
dblp:266/8226-1
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0001-6107-2858ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data SynthesisabstractIn the field of autonomous driving, sensor simulation is essential for generating rare and diverse scenarios that are difficult to capture in real-world environments. Current solutions fall into two categories: 1) CG-based methods, such as CARLA, which lack diversity and struggle to scale to the vast array of rare cases required for robust perception training; and 2) learning-based approaches, such as NeuSim, which are limited to specific object categories (vehicles) and require extensive multi-sensor data, hindering their applicability to generic objects. To address these limitations, we propose SynthDrive, a scalable "real2sim2real" system that leverages 3D generation to automate asset mining, generation, and rare-case data synthesis.Our framework introduces two key innovations: 1) Automated Rare-Case Mining and Synthesis. Given a text prompt describing specific objects, SynthDrive automatically mines image data from the Internet and then generates corresponding high-fidelity 3D assets, which eliminates the need for costly manual data collection. By integrating these assets into existing street-view data, our pipeline produces photorealistic rare-case data, supporting rapid scaling to diverse assets including irregular obstacles and temporary traffic facilities. 2) High-Fidelity 3D Generation. We propose a hybrid asset generation pipeline that combines a geometry-aware LRM, iterative mesh optimization, and an improved texture fusion algorithm. Our approach achieves 0.0164 Chamfer Distance on the GSO dataset, outperforming InstantMesh by 14.1% in geometry accuracy, and achieves 19.05 PSNR (vs 16.84) for texture quality. This enables fine geometry details and high-resolution texture generation, which is essential for perception model training. Experiments demonstrate that SynthDrive-generated data improves the performance of downstream perception tasks (2D and 3D detection on rare objects) by 2-4% mAP. SynthDrive greatly lowers the data production cost and improves the diversity for corner-case data generation, showcasing extensive potential applications in the field of autonomous driving. Zhengqing Chen, Ruohong Mei, Qingjie Wang, Yubin Hu 0001, Wei Yin 0006, Weiqiang Ren, Qian Zhang 0009 |
IROS | 5 |
| 2025 | Indoor Scene Reconstruction With Fine-Grained Details Using Hybrid Representation and Normal Prior EnhancementabstractThe reconstruction of indoor scenes from multi-view RGB images is challenging due to the coexistence of flat and texture-less regions alongside delicate and fine-grained regions. Recent methods leverage neural radiance fields aided by predicted surface normal priors to recover the scene geometry. These methods excel in producing complete and smooth results for floor and wall areas. However, they struggle to capture complex surfaces with high-frequency structures due to the inadequate neural representation and the inaccurately predicted normal priors. This work aims to reconstruct high-fidelity surfaces with fine-grained details by addressing the above limitations. To improve the capacity of the implicit representation, we propose a hybrid architecture to represent low-frequency and high-frequency regions separately. To enhance the normal priors, we introduce a simple yet effective image sharpening and denoising technique, coupled with a network that estimates the pixel-wise uncertainty of the predicted surface normal vectors. Identifying such uncertainty can prevent our model from being misled by unreliable surface normal supervisions that hinder the accurate reconstruction of intricate geometries. Experiments on the benchmark datasets show that our method outperforms existing methods in terms of reconstruction quality. Furthermore, the proposed method also generalizes well to real-world indoor scenarios captured by our hand-held mobile phones. Yubin Hu 0001, Matthieu Lin, Yu-Hui Wen, Wang Zhao 0001, Yong-Jin Liu 0001, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion ModelabstractOcclusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based in-painting model, to reconstruct complete surfaces for the hidden parts of objects. Specifically, we utilize a pre-trained diffusion model to fill in the hidden areas of 2D images. Then we use these in-painted images to optimize a neural implicit surface representation for each instance for 3D reconstruction. Since creating the in-painting masks needed for this process is tricky, we adopt a human-in-the-loop strategy that involves very little human engagement to generate high-quality masks. Moreover, some parts of objects can be totally hidden because the videos are usually shot from limited perspectives. To ensure recovering these invisible areas, we develop a cascaded network architecture for predicting signed distance field, making use of different frequency bands of positional encoding and maintaining overall smoothness. Besides the commonly used rendering loss, Eikonal loss, and silhouette loss, we adopt a CLIP-based semantic consistency loss to guide the surface from unseen camera angles. Experiments on ScanNet scenes show that our proposed framework achieves state-of-the-art accuracy and completeness in object-level reconstruction from scene-level RGB-D videos. Code: https://github.com/THU-LYJ-Lab/O2-Recon. Yubin Hu 0001, Wang Zhao 0001, Matthieu Lin, Yu-Hui Wen, Ying He 0001, Yong-Jin Liu 0001 |
AAAI | 1 |
| 2024 | Exploring Temporal Feature Correlation for Efficient and Stable Video Semantic SegmentationabstractThis paper tackles the problem of efficient and stable video semantic segmentation. While stability has been under-explored, prevalent work in efficient video semantic segmentation uses the keyframe paradigm. They efficiently process videos by only recomputing the low-level features and reusing high-level features computed at selected keyframes. In addition, the reused features stabilize the predictions across frames, thereby improving video consistency. However, dynamic scenes in the video can easily lead to misalignments between reused and recomputed features, which hampers performance. Moreover, relying on feature reuse to improve prediction consistency is brittle; an erroneous alignment of the features can easily lead to unstable predictions. Therefore, the keyframe paradigm exhibits a dilemma between stability and performance. We address this efficiency and stability challenge using a novel yet simple Temporal Feature Correlation (TFC) module. It uses the cosine similarity between two frames’ low-level features to inform the semantic label’s consistency across frames. Specifically, we selectively reuse label-consistent features across frames through linear interpolation and update others through sparse multi-scale deformable attention. As a result, we no longer directly reuse features to improve stability and thus effectively solve feature misalignment. This work provides a significant step towards efficient and stable video semantic segmentation. On the VSPW dataset, our method significantly improves the prediction consistency of image-based methods while being as fast and accurate. Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Yangguang Li 0001, Andrew Zhao, Gao Huang 0001, Yong-Jin Liu 0001 |
AAAI | 3 |
| 2024 | NGP-RT: Fusing Multi-level Hash Features with Lightweight Attention for Real-Time Novel View Synthesis
Yubin Hu 0001, Jingwei Huang 0001, Yong-Jin Liu 0001 |
ECCV (22) | 1 |
| 2024 | MMPI: a Flexible Radiance Field Representation by Multiple Multi-plane Images BlendingabstractThis paper presents a flexible representation of neural radiance fields based on multi-plane images (MPI), for high-quality view synthesis of complex scenes. MPI with Normalized Device Coordinate (NDC) parameterization is widely used in NeRF learning for its simple definition, easy calculation, and powerful ability to represent unbounded scenes. However, existing NeRF works that adopt MPI representation for novel view synthesis can only handle simple forward-facing unbounded scenes (e.g., the scenes in the LLFF dataset), where the input cameras are all observing in similar directions with small relative translations. Hence, extending these MPIbased methods to more complex scenes like large-range or even 360-degree scenes is very challenging. In this paper, we explore the potential of MPI and show that MPI can synthesize high-quality novel views of complex scenes with diverse camera distributions and view directions, which are not only limited to simple forward-facing scenes. Our key idea is to encode the neural radiance field with multiple MPIs facing different directions and blend them with an adaptive blending operation. For each region of the scene, the blending operation gives larger blending weights to those advantaged MPIs with stronger local representation abilities while giving lower weights to those with weaker representation abilities. Such blending operation automatically modulates the multiple MPIs to appropriately represent the diverse local density and color information. Experiments on the KITTI dataset and ScanNet dataset demonstrate that our proposed MMPI synthesizes high-quality images from diverse camera pose distributions and is fast to train, outperforming the previous fast-training NeRF methods for novel view synthesis. Moreover, we show that MMPI can encode extremely long trajectories and produce novel view renderings, demonstrating its potential in applications like autonomous driving. Our demo video is available at https://youtube.com/watch?v=mbNKwN5urC8. Peng Wang 0099, Yubin Hu 0001, Wang Zhao 0001, Ran Yi 0002, Yong-Jin Liu 0001, Wenping Wang 0001 |
ICRA | 3 |
| 2024 | SS3DM: Benchmarking Street-View Surface Reconstruction with a Synthetic 3D Mesh DatasetabstractReconstructing accurate 3D surfaces for street-view scenarios is crucial for applications such as digital entertainment and autonomous driving simulation. However, existing street-view datasets, including KITTI, Waymo, and nuScenes, only offer noisy LiDAR points as ground-truth data for geometric evaluation of reconstructed surfaces. These geometric ground-truths often lack the necessary precision to evaluate surface positions and do not provide data for assessing surface normals. To overcome these challenges, we introduce the SS3DM dataset, comprising precise \textbf{S}ynthetic \textbf{S}treet-view \textbf{3D} \textbf{M}esh models exported from the CARLA simulator. These mesh models facilitate accurate position evaluation and include normal vectors for evaluating surface normal. To simulate the input data in realistic driving scenarios for 3D reconstruction, we virtually drive a vehicle equipped with six RGB cameras and five LiDAR sensors in diverse outdoor scenes. Leveraging this dataset, we establish a benchmark for state-of-the-art surface reconstruction methods, providing a comprehensive evaluation of the associated challenges. For more information, visit our homepage at https://ss3dm.top. Yubin Hu 0001, Kairui Wen, Yong-Jin Liu 0001 |
NeurIPS | 1 |
| 2024 | AlphaTablets: A Generic Plane Representation for 3D Planar Reconstruction from Monocular VideosabstractWe introduce AlphaTablets, a novel and generic representation of 3D planes that features continuous 3D surface and precise boundary delineation. By representing 3D planes as rectangles with alpha channels, AlphaTablets combine the advantages of current 2D and 3D plane representations, enabling accurate, consistent and flexible modeling of 3D planes. We derive differentiable rasterization on top of AlphaTablets to efficiently render 3D planes into images, and propose a novel bottom-up pipeline for 3D planar reconstruction from monocular videos. Starting with 2D superpixels and geometric cues from pre-trained models, we initialize 3D planes as AlphaTablets and optimize them via differentiable rendering. An effective merging scheme is introduced to facilitate the growth and refinement of AlphaTablets. Through iterative optimization and merging, we reconstruct complete and accurate 3D planes with solid surfaces and clear boundaries. Extensive experiments on the ScanNet dataset demonstrate state-of-the-art performance in 3D planar reconstruction, underscoring the great potential of AlphaTablets as a generic 3D plane representation for various applications. Wang Zhao 0001, Shaohui Liu, Yubin Hu 0001, Yushi Bai, Yu-Hui Wen, Yong-Jin Liu 0001 |
NeurIPS | 4 |
| 2024 | Text-image conditioned diffusion for consistent text-to-3D generation
Yushi Bai, Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Qi Wang 0079, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Aided Geom. Des. | 5 |
| 2024 | Gaussian in the Dark: Real-Time View Synthesis From Inconsistent Dark Images Using Gaussian SplattingabstractAbstract 3D Gaussian Splatting has recently emerged as a powerful representation that can synthesize remarkable novel views using consistent multi‐view images as input. However, we notice that images captured in dark environments where the scenes are not fully illuminated can exhibit considerable brightness variations and multi‐view inconsistency, which poses great challenges to 3D Gaussian Splatting and severely degrades its performance. To tackle this problem, we propose Gaussian‐DK. Observing that inconsistencies are mainly caused by camera imaging, we represent a consistent radiance field of the physical world using a set of anisotropic 3D Gaussians, and design a camera response module to compensate for multi‐view inconsistencies. We also introduce a step‐based gradient scaling strategy to constrain Gaussians near the camera, which turn out to be floaters, from splitting and cloning. Experiments on our proposed benchmark dataset demonstrate that Gaussian‐DK produces high‐quality renderings without ghosting and floater artifacts and significantly outperforms existing methods. Furthermore, we can also synthesize light‐up images by controlling exposure levels that clearly show details in shadow areas. Zhen-Hui Dong, Yubin Hu 0001, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Graph. Forum | 3 |
| 2024 | PVP-Recon: Progressive View Planning via Warping Consistency for Sparse-View Surface ReconstructionabstractNeural implicit representations have revolutionized dense multi-view surface reconstruction, yet their performance significantly diminishes with sparse input views. A few pioneering works have sought to tackle this challenge by leveraging additional geometric priors or multi-scene generalizability. However, they are still hindered by the imperfect choice of input views, using images under empirically determined viewpoints. We propose PVP-Recon , a novel and effective sparse-view surface reconstruction method that progressively plans the next best views to form an optimal set of sparse viewpoints for image capturing. PVP-Recon starts initial surface reconstruction with as few as 3 views and progressively adds new views which are determined based on a novel warping score that reflects the information gain of each newly added view. This progressive view planning progress is interleaved with a neural SDF-based reconstruction module that utilizes multi-resolution hash features, enhanced by a progressive training scheme and a directional Hessian loss. Quantitative and qualitative experiments on three benchmark datasets show that our system achieves high-quality reconstruction with a constrained input budget and outperforms existing baselines. Matthieu Lin, Jenny Sheng, Ruoyu Fan, Yiheng Han, Yubin Hu 0001, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001, Wenping Wang 0001 |
ACM Trans. Graph. | 7 |
| 2023 | DarkFeat: Noise-Robust Feature Detector and Descriptor for Extremely Low-Light RAW ImagesabstractLow-light visual perception, such as SLAM or SfM at night, has received increasing attention, in which keypoint detection and local feature description play an important role. Both handcraft designs and machine learning methods have been widely studied for local feature detection and description, however, the performance of existing methods degrades in the extreme low-light scenarios in a certain degree, due to the low signal-to-noise ratio in images. To address this challenge, images in RAW format that retain more raw sensing information have been considered in recent works with a denoise-then-detect scheme. However, existing denoising methods are still insufficient for RAW images and heavily time-consuming, which limits the practical applications of such scheme. In this paper, we propose DarkFeat, a deep learning model which directly detects and describes local features from extreme low-light RAW images in an end-to-end manner. A novel noise robustness map and selective suppression constraints are proposed to effectively mitigate the influence of noise and extract more reliable keypoints. Furthermore, a customized pipeline of synthesizing dataset containing low-light RAW image matching pairs is proposed to extend end-to-end training. Experimental results show that DarkFeat achieves state-of-the-art performance on both indoor and outdoor parts of the challenging MID benchmark, outperforms the denoise-then-detect methods and significantly reduces computational costs up to 70%. Code is available at https://github.com/THU-LYJ-Lab/DarkFeat. Yubin Hu 0001, Wang Zhao 0001, Jisheng Li, Yong-Jin Liu 0001, Yuxing Han 0001, Jiangtao Wen |
AAAI | 2 |
| 2023 | Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosabstractVideo semantic segmentation (VSS) is a computationally expensive task due to the per-frame prediction for videos of high frame rates. In recent work, compact models or adaptive network strategies have been proposed for efficient VSS. However, they did not consider a crucial factor that affects the computational cost from the input side: the input resolution. In this paper, we propose an altering resolution framework called AR-Seg for compressed videos to achieve efficient VSS. AR-Seg aims to reduce the computational cost by using low resolution for non-keyframes. To prevent the performance degradation caused by downsampling, we design a Cross Resolution Feature Fusion (CR-eFF) module, and supervise it with a novel Feature Similarity Training (FST) strategy. Specifically, CReFF first makes use of motion vectors stored in a compressed video to warp features from high-resolution keyframes to low-resolution non-keyframes for better spatial alignment, and then selectively aggregates the warped features with local attention mechanism. Furthermore, the proposed FST supervises the aggregated features with high-resolution features through an explicit similarity loss and an implicit constraint from the shared decoding layer. Extensive experiments on CamVid and Cityscapes show that AR-Seg achieves state-of-the-art performance and is compatible with different segmentation backbones. On CamVid, AR-Seg saves 67% computational cost (measured in GFLOPs) with the PSPNet18 back-bone while maintaining high segmentation accuracy. Code: https://github.com/THU-LYJ-Lab/AR-Seg. Yubin Hu 0001, Yanghao Li, Jisheng Li, Yuxing Han 0001, Jiangtao Wen, Yong-Jin Liu 0001 |
CVPR | 1 |
| 2022 | Vision Perception Unit: Next-Generation Smart CMOS Image SensorabstractAs we reach the end of Moore’s Law and Dennard Scaling, it has become highly desirable to design a highly integrated and optimized pipeline specifically for computer vision. A new generation of integrated "smart" visual processors that streamline an end-to-end optimized visual information acquisition and processing pipeline (VIAPP) becomes necessary to lower the cost, power consumption, and latency.We describe a new paradigm for VIAPP as Vision Perception Unit (VPU), wherein electric signals generated by photons are amplified before converting to the digital signals to emulate an initial layer of a convolutional neural network (CNN). The outputs from these layers are then converted to digital signals and processed by following layers of a deep CNN. Wenqi Ji, Yuxing Han 0001, Jiangtao Wen, Yubin Hu 0001, Futang Wang, Jun Zhang 0006 |
HCS | 4 |
| 2021 | Learning To Compose 6-DOF Omnidirectional Videos Using Multi-Sphere ImagesabstractOmnidirectional video is an essential component of Virtual Reality. Although various methods have been proposed to generate content that can be viewed with six degrees of freedom (6-DoF), existing systems usually involve complex depth estimation, image inpainting or stitching pre-processing. In this paper, we propose a system that uses a 3D ConvNet to generate a multi-sphere images (MSI) representation that can be experienced in 6-DoF VR. The system utilizes conventional omnidirectional VR camera footage directly without the need for a depth map or segmentation mask, thereby significantly simplifying the overall complexity of the 6-DoF omnidirectional video composition. By using a newly designed weighted sphere sweep volume (WSSV) fusing technique, our approach is compatible with most panoramic VR camera setups. A ground truth generation approach for high-quality artifact-free 6-DoF contents is proposed and can be used by the research and development community for 6-DoF content generation. Jisheng Li, Yubin Hu 0001, Yuxing Han 0001, Jiangtao Wen |
ICIP | 3 |
| 2021 | Extending 6-DoF VR Experience Via Multi-Sphere Images InterpolationabstractThree-degrees-of-freedom (3-DoF) omnidirectional imaging has been widely used in various applications ranging from street maps to 3-DoF VR live broadcasting. Although allowing for navigating viewpoints rotationally inside a virtual world, it does not provide motion parallax key for human 3D perception. Recent research mitigates this problem by introducing 3 transitional degrees of freedom (6-DoF) using multi-sphere images (MSI) which is beginning to show promises in handling occlusions and reflective objects. However, the design of MSI naturally limits the range of authentic 6-DoF experiences, as existing mechanisms for MSI rendering cannot fully utilize multi-layer information when synthesizing novel views between multiple MSIs. To tackle this problem and extend the 6-DoF range, we propose an MSI interpolation pipeline that utilizes adjacent MSIs' 3D information embedded inside their layers. In this work, we describe an MSI projection scheme along with an MSI interpolation network to predict intermediate MSIs in order to facilitate the need for extended range. We demonstrate that our system significantly improves the range of 6-DoF experience compared with other MSI-based methods. With extensive experiments, we show our algorithm outperforms state-of-the-art methods both qualitatively and quantitatively in synthesizing novel view panoramas. Jisheng Li, Jinghui Jiao, Yubin Hu 0001, Yuxing Han 0001, Jiangtao Wen |
ACM Multimedia | 4 |