Youquan Liu

dblp:91/6914 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
abstract
Generative world models have become essential data engines for autonomous driving, yet most focus on videos or occupancy grids and overlook the unique challenges of LiDAR. Extending LiDAR generation to dynamic 4D modeling requires addressing controllability, temporal coherence, and standardized evaluation. We present LiDARCrafter, a unified framework for controllable 4D LiDAR generation and editing. Free-form language instructions are converted into ego-centric scene graphs that guide a tri-branch diffusion model to generate object geometry, motion, and structural priors. An autoregressive module further produces temporally coherent and stable LiDAR sequences with improved global consistency. To enable fair comparison, we introduce a comprehensive benchmark covering scene-, object-, and sequence-level metrics for rigorous and reproducible evaluation. Experiments on nuScenes show that LiDARCrafter achieves state-of-the-art fidelity, controllability, and temporal consistency, paving the way for scalable data augmentation and realistic simulation in diverse scenarios. Code have been publicly available at https://lidarcrafter.github.io.
Alan Liang, Youquan Liu, Dongyue Lu, Lingdong Kong, Huaici Zhao, Wei Tsang Ooi
AAAI2
2026 La La LiDAR: Large-Scale Layout Generation from LiDAR Data
abstract
Controllable generation of realistic LiDAR scenes is crucial for applications such as autonomous driving and robotics. While recent diffusion-based models achieve high-fidelity LiDAR generation, they lack explicit control over foreground objects and spatial relationships, limiting their usefulness for scenario simulation and safety validation. To address these limitations, we propose Large-scale Layout-guided LiDAR generation model ("La La LiDAR"), a novel layout-guided generative framework that introduces semantic-enhanced scene graph diffusion with relation-aware contextual conditioning for structured LiDAR layout generation, followed by foreground-aware control injection for complete scene generation. This enables customizable control over object placement while ensuring spatial and semantic consistency. To support our structured LiDAR generation, we introduce Waymo-SG and nuScenes-SG, two large-scale LiDAR scene graph datasets, along with new evaluation metrics for layout synthesis. Extensive experiments demonstrate that La La LiDAR achieves state-of-the-art performance in both LiDAR generation and downstream perception tasks, establishing a new benchmark for controllable 3D scene generation.
Youquan Liu, Lingdong Kong, Weidong Yang 0001, Xin Li 0110, Alan Liang, Runnan Chen, Ben Fei, Tongliang Liu
AAAI1
2026 LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving
abstract
Recent advancements in vision foundation models (VFMs) have revolutionized visual perception in 2D, yet their potential for 3D scene understanding, particularly in autonomous driving applications, remains underexplored. In this paper, we introduce LargeAD, a versatile and scalable framework designed for large-scale 3D pretraining across diverse real-world driving datasets. Our framework leverages VFMs to extract semantically rich superpixels from 2D images, which are aligned with LiDAR point clouds to generate high-quality contrastive samples. This alignment facilitates cross-modal representation learning, enhancing the semantic consistency between 2D and 3D data. We introduce several key innovations: (i) VFM-driven superpixel generation for detailed semantic representation, (ii) a VFM-assisted contrastive learning strategy to align multimodal features, (iii) superpoint temporal consistency to maintain stable representations across time, and (iv) multi-source data pretraining to generalize across various LiDAR configurations. Our approach achieves substantial gains over state-of-the-art methods in linear probing and fine-tuning for LiDAR-based segmentation and object detection. Extensive experiments on 11 large-scale multi-sensor datasets highlight our superior performance, demonstrating adaptability, efficiency, and robustness in real-world autonomous driving scenarios.
Lingdong Kong, Xiang Xu 0009, Youquan Liu, Jun Cen, Runnan Chen, Liang Pan, Kai Chen 0026, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 LLM4AV: Large Language Model-Driven Autonomous Vehicle Scenario Generation
Henri Patrick Kanimba Ntwali, Youquan Liu, Junyan Ma
WISA2
2025 Perspective-Invariant 3D Object Detection
abstract
With the rise of robotics, LiDAR-based 3D object detection has garnered significant attention in both academia and industry. However, existing datasets and methods predominantly focus on vehicle-mounted platforms, leaving other autonomous platforms underexplored. To bridge this gap, we introduce Pi3DET, the first benchmark featuring LiDAR data and 3D bounding box annotations collected from multiple platforms: vehicle, quadruped, and drone, thereby facilitating research in 3D object detection for non-vehicle platforms as well as cross-platform 3D detection. Based on Pi3DET, we propose a novel cross-platform adaptation framework that transfers knowledge from the well-studied vehicle platform to other platforms. This framework achieves perspective-invariant 3D detection through robust alignment at both geometric and feature levels. Additionally, we establish a benchmark to evaluate the resilience and robustness of current 3D detectors in cross-platform scenarios, providing valuable insights for developing adaptive 3D perception systems. Extensive experiments validate the effectiveness of our approach on challenging cross-platform tasks, demonstrating substantial gains over existing adaptation methods. We hope this work paves the way for generalizable and unified 3D perception systems across diverse and complex environments. Our Pi3DET dataset, cross-platform benchmark suite, and annotation toolkit have been made publicly available.
Ao Liang, Lingdong Kong, Dongyue Lu, Youquan Liu, Huaici Zhao, Wei Tsang Ooi
ICCV4
2025 Lightweight Object Tracking and Localization for Assembly Guidance on AR Helmet
Youquan Liu, Ruizhi Wan
ICIG (3)2
2025 3EED: Ground Everything Everywhere in 3D
abstract
Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce 3EED, a multi-platform, multi-modal 3D grounding benchmark featuring RGB and LiDAR data from vehicle, drone, and quadruped platforms. We provide over 128,000 objects and 22,000 validated referring expressions across diverse outdoor scenes -- 10x larger than existing datasets. We develop a scalable annotation pipeline combining vision-language model prompting with human verification to ensure high-quality spatial grounding. To support cross-platform learning, we propose platform-aware normalization and cross-modal alignment techniques, and establish benchmark protocols for in-domain and cross-platform evaluations. Our findings reveal significant performance gaps, highlighting the challenges and opportunities of generalizable 3D grounding. The 3EED dataset and benchmark toolkit are released to advance future research in language-driven 3D embodied perception.
Yuhao Dong, Tianshuai Hu, Alan Liang, Youquan Liu, Dongyue Lu, Liang Pan, Lingdong Kong, Junwei Liang 0001, Ziwei Liu 0002
NeurIPS5
2025 Spiral: Semantic-Aware Progressive LiDAR Scene Generation and Understanding
abstract
Leveraging diffusion models, 3D LiDAR scene generation has achieved great success in both range-view and voxel-based representations. While recent voxel-based approaches can generate both geometric structures and semantic labels, existing range-view methods are limited to producing unlabeled LiDAR scenes. Relying on pretrained segmentation models to predict the semantic maps often results in suboptimal cross-modal consistency. To address this limitation while preserving the advantages of range-view representations, such as computational efficiency and simplified network design, we propose Spiral, a novel range-view LiDAR diffusion model that simultaneously generates depth, reflectance images, and semantic maps. Furthermore, we introduce novel semantic-aware metrics to evaluate the quality of the generated labeled range-view data. Experiments on SemanticKITTI and nuScenes datasets demonstrate that Spiral achieves state-of-the-art performance with the smallest parameter size, outperforming two-step methods that combine the best available generative and segmentation models. Additionally, we validate that Spiral’s generated range images can be effectively used for synthetic data augmentation in the downstream segmentation training, significantly reducing the labeling effort on LiDAR data.
Dekai Zhu, Yixuan Hu, Youquan Liu, Dongyue Lu, Lingdong Kong, Slobodan Ilic
NeurIPS3
2025 Visual Foundation Models Boost Cross-Modal Unsupervised Domain Adaptation for 3D Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) is vital for alleviating the workload of labeling 3D point cloud data and mitigating the absence of labels when facing an unseen domain. Various methods have recently emerged to utilize images along with point clouds to enhance the performance of cross-domain 3D segmentation. However, the pseudo labels, which are generated from models trained on the source domain and provide additional supervised signals for the target domain, are inadequate when utilized for 3D segmentation due to their inherent noisiness and consequently restrict the accuracy of neural networks. With the advent of 2D Visual Foundation Models (VFMs) and their abundant knowledge prior, we propose a novel pipeline VFMSeg to further enhance the cross-modal UDA framework by leveraging these models. In this work, we study how to harness the knowledge priors learned by VFMs to produce more accurate labels for unlabeled target domains and improve overall performance. We first utilize a multi-modal VFM, which is pre-trained on large-scale image-text pairs, to provide supervised labels (VFM-PL) for images and point clouds from the target domain. Then, we adopt another VFM to generate fine-grained 2D masks for guiding the generation of augmented images and point clouds, which mix the data from source and target domains like view frustums (FrustumMixing). Finally, we merge class-wise prediction across modalities to produce more accurate annotations for unlabeled target domains. Our method is evaluated on various autonomous driving datasets and the results demonstrate a significant improvement in 3D segmentation task. Our code is available athttps://github.com/EtronTech/VFMSeg
Weidong Yang 0001, Lingdong Kong, Youquan Liu, Qingyuan Zhou, Rui Zhang 0103, Zhijun Li 0001, Wenming Chen 0001, Ben Fei
IEEE Trans. Intell. Transp. Syst.4
2024 OpenESS: Event-Based Semantic Scene Understanding with Open Vocabularies
abstract
Event-based semantic segmentation (ESS) is a fundamental yet challenging task for event camera sensing. The difficulties in interpreting and annotating event data limit its scalability. While domain adaptation from images to event data can help to mitigate this issue, there exist data representational differences that require additional effort to resolve. In this work, for the first time, we synergize information from image, text, and event-data domains and introduce OpenESS to enable scalable ESS in an open-world, annotation-efficient manner. We achieve this goal by transferring the semantically rich CLIP knowledge from image-text pairs to event streams. To pursue better cross-modality adaptation, we propose a frame-to-event contrastive distillation and a text-to-event semantic consistency regularization. Experimental results on popular ESS benchmarks showed our approach outperforms existing methods. Notably, we achieve 53.93% and 43.31% mIoU on DDD17 and DSEC-Semantic without using either event or frame labels.
Lingdong Kong, Youquan Liu, Lai Xing Ng, Benoit Cottereau, Wei Tsang Ooi
CVPR2
2024 Multi-Space Alignments Towards Universal LiDAR Segmentation
abstract
A unified and versatile LiDAR segmentation model with strong robustness and generalizability is desirable for safe autonomous driving perception. This work presents M3Net, a one-of-a-kind framework for fulfilling multitask, multi-dataset, multimodality LiDAR segmentation in a universal manner using just a single set of parameters. To better exploit data volume and diversity, we first combine large-scale driving datasets acquired by different types of sensors from diverse scenes and then conduct alignments in three spaces, namely data, feature, and label spaces, during the training. As a result, M3Net is capable of taming heterogeneous data for training state-of-the-art LiDAR segmentation models. Extensive experiments on twelve LiDAR segmentation datasets verify our effectiveness. Notably, using a shared set of parameters, M3Net achieves 75.1%,83.1%, and 72.4% mIoU scores, respectively, on the official benchmarks of SemanticKITTI, nuScenes, and Waymo Open.
Youquan Liu, Lingdong Kong, Xiaoyang Wu 0002, Runnan Chen, Xin Li 0110, Liang Pan, Ziwei Liu 0002, Yuexin Ma
CVPR1
2024 Learning to Adapt SAM for Segmenting Cross-Domain Point Clouds
Xidong Peng, Runnan Chen, Feng Qiao 0001, Lingdong Kong, Youquan Liu, Yujing Sun 0001, Xinge Zhu, Yuexin Ma
ECCV (43)5
2023 CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIP
abstract
Contrastive Language-Image Pre-training (CLIP) achieves promising results in 2D zero-shot and few-shot learning. Despite the impressive performance in 2D, applying CLIP to help the learning in 3D scene understanding has yet to be explored. In this paper, we make the first attempt to investigate how CLIP knowledge benefits 3D scene understanding. We propose CLIP2Scene, a simple yet effective framework that transfers CLIP knowledge from 2D image-text pre-trained models to a 3D point cloud network. We show that the pre-trained 3D network yields impressive performance on various downstream tasks, i.e., annotation-free and fine-tuning with labelled data for semantic segmentation. Specifically, built upon CLIP, we design a Semantic-driven Cross-modal Contrastive Learning framework that pre-trains a 3D network via semantic and spatial-temporal consistency regularization. For the former, we first leverage CLIP's text semantics to select the positive and negative point samples and then employ the contrastive loss to train the 3D network. In terms of the latter, we force the consistency between the temporally coherent point cloud features and their corresponding image features. We conduct experiments on SemanticKITTI, nuScenes, and ScanNet. For the first time, our pre-trained network achieves annotation-free 3D semantic segmentation with 20.8% and 25.08% mIoU on nuScenes and ScanNet, respectively. When fine-tuned with 1% or 100% labelled data, our method significantly outperforms other self-supervised methods, with improvements of 8% and 1% mIoU, respectively. Furthermore, we demonstrate the generalizability for handling cross-domain datasets. Code is publicly available11https://github.com/runnanchen/CLIP2Scene..
Runnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu, Yuexin Ma, Yikang Li 0002, Yuenan Hou, Yu Qiao 0001, Wenping Wang 0001
CVPR2
2023 LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal Fusion
abstract
LiDAR-camera fusion methods have shown impressive performance in 3D object detection. Recent advanced multi-modal methods mainly perform global fusion, where image features and point cloud features are fused across the whole scene. Such practice lacks fine-grained region-level information, yielding suboptimal fusion performance. In this paper, we present the novel Local-to-Global fusion network (LoGoNet), which performs LiDAR-camerafusion at both local and global levels. Concretely, the Global Fusion (GoF) of LoGoNet is built upon previous literature, while we exclusively use point centroids to more precisely represent the position of voxel features, thus achieving better crossmodal alignment. As to the Local Fusion (LoF), we first divide each proposal into uniform grids and then project these grid centers to the images. The image features around the projected grid points are sampled to be fused with position-decorated point cloud features, maximally uti-lizing the rich contextual information around the proposals. The Feature Dynamic Aggregation (FDA) module is further proposed to achieve information interaction between these locally and globally fused features, thus producing more informative multi-modal features. Extensive experiments on both Waymo Open Dataset (WOD) and KITTI datasets show that LoGoNet outperforms all state-of-the-art 3D detection methods. Notably, LoGoNet ranks 1st on Waymo 3D object detection leaderboard and obtains 81.02 mAPH (L2) detection performance. It is noteworthy that, for the first time, the detection performance on three classes surpasses 80 APH (L2) simultaneously. Code will be available at https://github.com/sankin97/LoGoNet.
Xin Li 0110, Tao Ma 0002, Yuenan Hou, Botian Shi, Yuchen Yang 0003, Youquan Liu, Xingjiao Wu, Qin Chen 0001, Yikang Li 0002, Yu Qiao 0001, Liang He 0001
CVPR6
2023 SCPNet: Semantic Scene Completion on Point Cloud
abstract
Training deep models for semantic scene completion (SSC) is challenging due to the sparse and incomplete input, a large quantity of objects of diverse scales as well as the inherent label noise for moving objects. To address the above-mentioned problems, we propose the following three solutions: 1) Redesigning the completion sub-network. We design a novel completion sub-network, which consists of several Multi-Path Blocks (MPBs) to aggregate multi-scale features and is free from the lossy downsampling operations. 2) Distilling rich knowledge from the multi-frame model. We design a novel knowledge distillation objective, dubbed Dense-to-Sparse Knowledge Distillation (DSKD). It transfers the dense, relation-based semantic knowledge from the multi-frame teacher to the single-frame student, significantly improving the representation learning of the single-frame model. 3) Completion label rectification. We propose a simple yet effective label rectification strategy, which uses off-the-shelf panoptic segmentation labels to remove the traces of dynamic objects in completion labels, greatly improving the performance of deep models especially for those moving objects. Extensive experiments are conducted in two public SSC benchmarks, i.e., SemanticKITTI and SemanticPOSS. Our SCPNet ranks 1st on SemanticKITTI semantic scene completion challenge and surpasses the competitive S3CNet [3] by 7.2 mIoU. SCP-Net also outperforms previous completion algorithms on the SemanticPOSS dataset. Besides, our method also achieves competitive results on SemanticKITTI semantic segmentation tasks, showing that knowledge learned in the scene completion is beneficial to the segmentation task.
Zhaoyang Xia, Youquan Liu, Xin Li 0110, Xinge Zhu, Yuexin Ma, Yikang Li 0002, Yuenan Hou, Yu Qiao 0001
CVPR2
2023 Rethinking Range View Representation for LiDAR Segmentation
abstract
LiDAR segmentation is crucial for autonomous driving perception. Recent trends favor point- or voxel-based methods as they often yield better performance than the traditional range view representation. In this work, we unveil several key factors in building powerful range view models. We observe that the "many-to-one" mapping, semantic incoherence, and shape deformation are possible impediments against effective learning from range view projections. We present RangeFormer – a full-cycle framework comprising novel designs across network architecture, data augmentation, and post-processing – that better handles the learning and processing of LiDAR point clouds from the range view. We further introduce a Scalable Training from Range view (STR) strategy that trains on arbitrary low-resolution 2D range images, while still maintaining satisfactory 3D segmentation accuracy. We show that, for the first time, a range view method is able to surpass the point, voxel, and multi-view fusion counterparts in the competing LiDAR semantic and panoptic segmentation benchmarks, i.e., SemanticKITTI, nuScenes, and ScribbleKITTI.
Lingdong Kong, Youquan Liu, Runnan Chen, Yuexin Ma, Xinge Zhu, Yikang Li 0002, Yuenan Hou, Yu Qiao 0001, Ziwei Liu 0002
ICCV2
2023 Robo3D: Towards Robust and Reliable 3D Perception against Corruptions
abstract
The robustness of 3D perception systems under natural corruptions from environments and sensors is pivotal for safety-critical applications. Existing large-scale 3D perception datasets often contain data that are meticulously cleaned. Such configurations, however, cannot reflect the reliability of perception models during the deployment stage. In this work, we present Robo3D, the first comprehensive benchmark heading toward probing the robustness of 3D detectors and segmentors under out-of-distribution scenarios against natural corruptions that occur in real-world environments. Specifically, we consider eight corruption types stemming from severe weather conditions, external disturbances, and internal sensor failure. We uncover that, although promising results have been progressively achieved on standard benchmarks, state-of-the-art 3D perception models are at risk of being vulnerable to corruptions. We draw key observations on the use of data representations, augmentation schemes, and training strategies, that could severely affect the model's performance. To pursue better robustness, we propose a density-insensitive training framework along with a simple flexible voxelization strategy to enhance the model resiliency. We hope our benchmark and approach could inspire future research in designing more robust and reliable 3D perception models. Our robustness benchmark suite is publicly available1.
Lingdong Kong, Youquan Liu, Xin Li 0110, Runnan Chen, Jiawei Ren 0001, Liang Pan, Kai Chen 0026, Ziwei Liu 0002
ICCV2
2023 UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase
abstract
Point-, voxel-, and range-views are three representative forms of point clouds. All of them have accurate 3D measurements but lack color and texture information. RGB images are a natural complement to these point cloud views and fully utilizing the comprehensive information of them benefits more robust perceptions. In this paper, we present a unified multi-modal LiDAR segmentation network, termed UniSeg, which leverages the information of RGB images and three views of the point cloud, and accomplishes semantic segmentation and panoptic segmentation simultaneously. Specifically, we first design the Learnable cross-Modal Association (LMA) module to automatically fuse voxel-view and range-view features with image features, which fully utilize the rich semantic information of images and are robust to calibration errors. Then, the enhanced voxel-view and range-view features are transformed to the point space, where three views of point cloud features are further fused adaptively by the Learnable cross-View Association module (LVA). Notably, UniSeg achieves promising results in three public benchmarks, i.e., SemanticKITTI, nuScenes, and Waymo Open Dataset (WOD); it ranks 1st on two challenges of two benchmarks, including the LiDAR semantic segmentation challenge of nuScenes and panoptic segmentation challenges of SemanticKITTI. Besides, we construct the OpenPCSeg codebase, which is the largest and most comprehensive outdoor LiDAR segmentation codebase. It contains most of the popular outdoor LiDAR segmentation algorithms and provides reproducible implementations. The OpenPCSeg codebase will be made publicly available at https://github.com/PJLab-ADG/PCSeg.
Youquan Liu, Runnan Chen, Xin Li 0110, Lingdong Kong, Yuchen Yang 0003, Zhaoyang Xia, Yeqi Bai, Xinge Zhu, Yuexin Ma, Yikang Li 0002, Yu Qiao 0001, Yuenan Hou
ICCV1
2023 RangePerception: Taming LiDAR Range View for Efficient and Accurate 3D Object Detection
abstract
LiDAR-based 3D detection methods currently use bird's-eye view (BEV) or range view (RV) as their primary basis. The former relies on voxelization and 3D convolutions, resulting in inefficient training and inference processes. Conversely, RV-based methods demonstrate higher efficiency due to their compactness and compatibility with 2D convolutions, but their performance still trails behind that of BEV-based methods. To eliminate this performance gap while preserving the efficiency of RV-based methods, this study presents an efficient and accurate RV-based 3D object detection framework termed RangePerception. Through meticulous analysis, this study identifies two critical challenges impeding the performance of existing RV-based methods: 1) there exists a natural domain gap between the 3D world coordinate used in output and 2D range image coordinate used in input, generating difficulty in information extraction from range images; 2) native range images suffer from vision corruption issue, affecting the detection accuracy of the objects located on the margins of the range images. To address the key challenges above, we propose two novel algorithms named Range Aware Kernel (RAK) and Vision Restoration Module (VRM), which facilitate information flow from range image representation and world-coordinate 3D detection results. With the help of RAK and VRM, our RangePerception achieves 3.25/4.18 higher averaged L1/L2 AP compared to previous state-of-the-art RV-based method RangeDet, on Waymo Open Dataset. For the first time as an RV-based 3D detection method, RangePerception achieves slightly superior averaged AP compared with the well-known BEV-based method CenterPoint and the inference speed of RangePerception is 1.3 times as fast as CenterPoint.
Yeqi Bai, Ben Fei, Youquan Liu, Tao Ma 0002, Yuenan Hou, Botian Shi, Yikang Li 0002
NeurIPS3
2023 Towards Label-free Scene Understanding by Vision Foundation Models
abstract
Vision foundation models such as Contrastive Vision-Language Pre-training (CLIP) and Segment Anything (SAM) have demonstrated impressive zero-shot performance on image classification and segmentation tasks. However, the incorporation of CLIP and SAM for label-free scene understanding has yet to be explored. In this paper, we investigate the potential of vision foundation models in enabling networks to comprehend 2D and 3D worlds without labelled data. The primary challenge lies in effectively supervising networks under extremely noisy pseudo labels, which are generated by CLIP and further exacerbated during the propagation from the 2D to the 3D domain. To tackle these challenges, we propose a novel Cross-modality Noisy Supervision (CNS) method that leverages the strengths of CLIP and SAM to supervise 2D and 3D networks simultaneously. In particular, we introduce a prediction consistency regularization to co-train 2D and 3D networks, then further impose the networks' latent space consistency using the SAM's robust feature representation. Experiments conducted on diverse indoor and outdoor datasets demonstrate the superior performance of our method in understanding 2D and 3D open environments. Our 2D and 3D network achieves label-free semantic segmentation with 28.4\% and 33.5\% mIoU on ScanNet, improving 4.7\% and 7.9\%, respectively. For nuImages and nuScenes datasets, the performance is 22.1\% and 26.8\% with improvements of 3.5\% and 6.0\%, respectively. Code is available. (https://github.com/runnanchen/Label-Free-Scene-Understanding)
Runnan Chen, Youquan Liu, Lingdong Kong, Nenglun Chen, Xinge Zhu, Yuexin Ma, Tongliang Liu, Wenping Wang 0001
NeurIPS2
2023 Segment Any Point Cloud Sequences by Distilling Vision Foundation Models
abstract
Recent advancements in vision foundation models (VFMs) have opened up new possibilities for versatile and efficient visual perception. In this work, we introduce Seal, a novel framework that harnesses VFMs for segmenting diverse automotive point cloud sequences. Seal exhibits three appealing properties: i) Scalability: VFMs are directly distilled into point clouds, obviating the need for annotations in either 2D or 3D during pretraining. ii) Consistency: Spatial and temporal relationships are enforced at both the camera-to-LiDAR and point-to-segment regularization stages, facilitating cross-modal representation learning. iii) Generalizability: Seal enables knowledge transfer in an off-the-shelf manner to downstream tasks involving diverse point clouds, including those from real/synthetic, low/high-resolution, large/small-scale, and clean/corrupted datasets. Extensive experiments conducted on eleven different point cloud datasets showcase the effectiveness and superiority of Seal. Notably, Seal achieves a remarkable 45.0% mIoU on nuScenes after linear probing, surpassing random initialization by 36.9% mIoU and outperforming prior arts by 6.1% mIoU. Moreover, Seal demonstrates significant performance gains over existing methods across 20 different few-shot fine-tuning tasks on all eleven tested point cloud datasets. The code is available at this link.
Youquan Liu, Lingdong Kong, Jun Cen, Runnan Chen, Liang Pan, Kai Chen 0026, Ziwei Liu 0002
NeurIPS1
2020 Tangible Multi-card Projector-Based Interaction with Physics
Songxue Wang, Youquan Liu, Junxiu Guo
ICEC2
2017 A novel surface tension formulation for SPH fluid simulation
Meng Yang 0011, Xiaosheng Li, Youquan Liu, Gang Yang 0007, Enhua Wu
Vis. Comput.3
2013 A unified smoke control method based on signed distance field
Ben Yang, Youquan Liu, Lihua You, Xiaogang Jin 0001
Comput. Graph.2
2013 Animating turbulent water by vortex shedding in PIC/FLIP
Jian Zhu 0001, Youquan Liu, Yuanzhang Chang, Enhua Wu
Sci. China Inf. Sci.2
2012 Abstract line drawings from photographs using flow-based filters
Shandong Wang, Enhua Wu, Youquan Liu, Xuehui Liu, Yanyun Chen
Comput. Graph.3
2012 Physically based object withering simulation
abstract
ABSTRACT This paper presents a finite element method‐based framework for an object withering simulation modeled with heterogeneous material, such as fruits drying or decay. We introduce diffusion procedures for both the moisture content and decay spread, which are solved directly on a tetrahedral mesh representation of the fruit flesh. Then, we use the moisture content to control shrinkage through the initial strain, which is integrated into the Lagrangian dynamic equation, and solved with the finite element method. For the complex structure of the object, another fine triangle mesh is used to represent the skin, and its deformation is solved by a thin shell technique. To couple the motion between different layers of the fruit, a tracking force is used to pull the skin and drive its deformation together with the volume mesh. In comparison with the previous work, our method provides temporally and spatially varying parameters to model the complex phenomena of object withering. Moreover, the water diffusivity can also be given by user input to present various material properties of the cut section and skin‐covered area. Our algorithm is easy to implement and highly efficient in generating a realistic appearance for the withering effect. For a medium‐scale model, we can achieve interactive simulation. Copyright © 2012 John Wiley & Sons, Ltd.
Youquan Liu, Yanyun Chen, Wen Wu 0001, Nelson L. Max, Enhua Wu
Comput. Animat. Virtual Worlds1
2012 Interactive coupling between a tree and raindrops
abstract
ABSTRACT This paper presents a novel approach for simulating the dynamic coupling between a tree and raindrops based on physical deformation and fluid simulation. By the approach, tree animation in the rain can be simulated in a two‐resolution way: branch motion and leaf motion. The branch is represented by the Euler–Bernoulli beam model, and the leaf petiole is represented by the three‐prism elastic model. Interaction coupling liquid motion on the hydrophilic surface with a flexible petiole is well implemented by a special design. To simplify the computation process, instead of the computation‐intensive three‐dimensional Navier–Stokes equations, shallow water equations are used to simulate the water dynamics together with the whole leaf deformation. Simulation has been also made to various phenomena incurred from the interactive coupling. These include, among others, part of impacting raindrops splashing into the air with the remaining flowing along the slant of the leaf and merging into larger ones or hanging on the blade boundary, with the leaf rebounding and vibrating after the drops fall off the leaf. A level‐of‐detail approach is exploited to accelerate rendering in views of different distances. The experimental results illustrate that the approach can be applied to efficiently generate realistic details of the interactive coupling between a tree and raindrops. Copyright © 2012 John Wiley & Sons, Ltd.
Meng Yang 0011, Longsheng Jiang, Xiaosheng Li, Youquan Liu, Xuehui Liu, Enhua Wu
Comput. Animat. Virtual Worlds4
2009 Time-Varying Simulation for Image-Based Carpets
abstract
The paper presents a novel approach for simulating realistic time-varied carpets. By the approach, a 3D carpet is constructed first by image-based techniques from a single photo input, through a texel structure, established to generate realistic carpet with its pattern guided by the captured image. Secondly, a time-varying simulation model is proposed to capture various aspects of the time-dependent appearance of carpets such as dust accumulation, color fading and fiber change. Additionally, a time-varying map is provided to control the specific weathering degrees at different voxels in time through a hierarchal model. Experimental results show that the realistic time-varying simulation is successfully achieved with the proposed techniques.
Shaohui Jiao, Youquan Liu, Enhua Wu
ICIG2
2009 A particle-based method for viscoelastic fluids animation
abstract
We present a particle-based method for viscoelastic fluids simulation. In the method, based on the traditional Navier-Stokes equation, an additional elastic stress term is introduced to achieve viscoelastic flow behaviors, which have both fluid and solid features. Benefiting from the Lagrangian nature of Smoothed Particle Hydrodynamics, large flow deformation can be handled more easily and naturally. And also, by changing the viscosity and elastic stress coefficient of the particles according to the temperature variation, the melting and flowing phenomena, such as lava flow and wax melting, are achieved. The temperature evolution is determined with the heat diffusion equation. The method is effective and efficient, and has good controllability. Different kinds of viscoelastic fluid behaviors can be obtained easily by adjusting the very few experimental parameters.
Yuanzhang Chang, Kai Bao, Youquan Liu, Jian Zhu 0001, Enhua Wu
VRST3
2007 Simulation and interaction of fluid dynamics
Enhua Wu, Hongbin Zhu, Xuehui Liu, Youquan Liu
Vis. Comput.4
2006 Simulation of Fluid Dynamics and Interactions
abstract
Through interaction with surroundings, the fluids may change their properties such as shapes, temperature vastly, and the same would happen to the surroundings simultaneously. On the other hand, different surroundings characterize different interactions, and may change the shapes and motions of the fluids in different ways. Therefore, it is of importance in physically-based simulation of fluids to build physically correct models to represent the varying interactions between fluids and the environments. In this paper, we make a simple summation on the interactions, and in particular focus on those most interesting to us, and model them with various physical solutions. In some of the methods, advantage is taken with the graphics processing unit (GPU) to achieve real-time computation for medial-scale simulation
Enhua Wu, Hongbin Zhu, Xuehui Liu, Youquan Liu
CW4
2006 Simulation of miscible binary mixtures based on lattice Boltzmann method
abstract
Abstract Miscible fluid mixtures, like pouring honey into water, Coca Cola into strong wine, are common phenomena in our daily life. While two miscible fluids are mixed together, their appearances in terms of colors and shapes will change due to their mixing interaction. The interaction between the mixture components could be regarded as a combination of the diffusing process and demixing process. If the former dominates the interaction, it is miscible; otherwise, it is immiscible. The complex microscopic interplay between the mixture components makes the simulation highly challenging. So far, there have been some dedicated research in computer graphics dealing with immiscible mixtures, but few works have been done focusing on miscible mixtures. In this paper, for the first time, we introduce a two‐fluid lattice Boltzmann method (LBM), called TFLBM, applied to miscible binary mixtures. Different from other similar methods, the viscous and diffusing properties of the fluid in our work are considered separately, so that the physical insight is exposed more clearly and rationally. In addition, the operation of LBM is mostly a linear local computation, and graphics processing unit (GPU) has been utilized to achieve real‐time simulation. Copyright © 2006 John Wiley & Sons, Ltd.
Hongbin Zhu, Xuehui Liu, Youquan Liu, Enhua Wu
Comput. Animat. Virtual Worlds3
2004 Real-Time 3D Fluid Simulation on GPU with Complex Obstacles
abstract
In this paper, we solve the 3D fluid dynamics problem in a complex environment by taking advantage of the parallelism and programmability of GPU. In difference from other methods, innovation is made in two aspects. Firstly, more general boundary conditions could be processed on GPU in our method. By the method, we generate the boundary from a 3D scene with solid clipping, making the computation run on GPU despite of the complexity of the whole geometry scene. Then by grouping the voxels into different types according to their positions relative to the obstacles and locating the voxel that determines the value of the current voxel, we modify the values on the boundaries according to the boundary conditions. Secondly, more compact structure in data packing with flat 3D textures is designed at the fragment processing level to enhance parallelism and reduce execution passes. The scalar variables including density and temperature are packed into four channels of texels to accelerate the computation of 3D Navier-Stokes equations (NSEs). The test results prove the efficiency of our method, and as a result, it is feasible to run middle-scale problems of 3D fluid dynamics in an interactive speed for more general environment with complex geometry on PC platform.
Youquan Liu, Xuehui Liu, Enhua Wu
PG1
2004 An improved study of real-time fluid simulation on GPU
abstract
Abstract Taking advantage of the parallelism and programmability of GPU, we solve the fluid dynamics problem completely on GPU. Different from previous methods, the whole computation is accelerated in our method by packing the scalar and vector variables into four channels of texels. In order to be adaptive to the arbitrary boundary conditions, we group the grid nodes into different types according to their positions relative to obstacles and search the node that determines the value of the current node. Then we compute the texture coordinates offsets according to the type of the boundary condition of each node to determine the corresponding variables and achieve the interaction of flows with obstacles set freely by users. The test results prove the efficiency of our method and exhibit the potential of GPU for general‐purpose computations. Copyright © 2004 John Wiley & Sons, Ltd.
Enhua Wu, Youquan Liu, Xuehui Liu
Comput. Animat. Virtual Worlds2