Xin Li 0110

dblp:09/1365-110 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
23since 2021 · last 2026
0009-0001-4575-8603ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
abstract
Alignment plays a crucial role in Large Language Models (LLMs) in aligning with human preferences on a specific task/domain. Traditional alignment methods suffer from catastrophic forgetting, where models lose previously learned values when adapting to new preferences or domains. We introduce LifeAlign, a novel framework for lifelong alignment that enables LLMs to maintain consistent human preference alignment across sequential learning tasks without forgetting previously learned values. Our approach consists of two key innovations. First, we propose a focalized preference optimization strategy that aligns LLMs with new preferences while preventing the erosion of alignment acquired from previous tasks. Second, we develop a short-to-long memory consolidation mechanism that merges denoised short-term preference representations into stable long-term memory using intrinsic dimensionality reduction, enabling efficient storage and retrieval of alignment patterns across diverse domains. We evaluate LifeAlign across multiple sequential alignment tasks spanning different domains and preference types. Experimental results demonstrate that our method achieves superior performance in maintaining both preference alignment quality and knowledge retention compared to existing lifelong learning approaches.
Junsong Li, Jie Zhou 0015, Bihao Zhan, Yutao Yang, Qianjun Pan, Shilian Chen, Tianyu Huai, Xin Li 0110, Qin Chen 0001, Liang He 0001
AAAI8
2026 La La LiDAR: Large-Scale Layout Generation from LiDAR Data
abstract
Controllable generation of realistic LiDAR scenes is crucial for applications such as autonomous driving and robotics. While recent diffusion-based models achieve high-fidelity LiDAR generation, they lack explicit control over foreground objects and spatial relationships, limiting their usefulness for scenario simulation and safety validation. To address these limitations, we propose Large-scale Layout-guided LiDAR generation model ("La La LiDAR"), a novel layout-guided generative framework that introduces semantic-enhanced scene graph diffusion with relation-aware contextual conditioning for structured LiDAR layout generation, followed by foreground-aware control injection for complete scene generation. This enables customizable control over object placement while ensuring spatial and semantic consistency. To support our structured LiDAR generation, we introduce Waymo-SG and nuScenes-SG, two large-scale LiDAR scene graph datasets, along with new evaluation metrics for layout synthesis. Extensive experiments demonstrate that La La LiDAR achieves state-of-the-art performance in both LiDAR generation and downstream perception tasks, establishing a new benchmark for controllable 3D scene generation.
Youquan Liu, Lingdong Kong, Weidong Yang 0001, Xin Li 0110, Alan Liang, Runnan Chen, Ben Fei, Tongliang Liu
AAAI4
2026 Control-Communication Co-Design for Cloud-Fog Automation Over 5G-TSN: A Coupling Loop Method Under Network Uncertainty
abstract
The cloud–fog automation (CFA) paradigm accelerates Industry 4.0 by enabling fully automated industrial systems through interconnected wired and wireless devices. However, dynamic uncertainties such as environment-induced network delays and disturbances challenge the stability and efficiency of these systems. To address these issues, this paper proposes a control–communication co-design framework named C3L (Control– Communication Coupling Loop), which integrates 5G and Time-Sensitive Networking (TSN) to coordinate transmission and control across cloud and fog layers. A hierarchical control strategy is introduced based on a delay threshold derived from Linear Matrix Inequality (LMI) analysis, enabling dynamic selection between cloud and fog control units for both stability and responsiveness. To ensure timely and reliable control input delivery under uncertain networks, we develop a deterministic transmission mechanism centered on a novel metric, Packet Loss Tolerance (PLT), which quantifies how many consecutive losses the system can endure while maintaining stability. Additionally, a new optimization criterion, Pareto Deviation Value (PDV), is proposed to avoid combining control and communication costs with incompatible physical units. Simulation results demonstrate that the proposed method enhances control robustness and communication efficiency in complex industrial networks.
Xuanzhao Lu, Qimin Xu, Meihan Lin, Xin Li 0110, Cailian Chen, Xin-Ping Guan
IEEE Internet Things J.4
2025 DriveArena: A Closed-Loop Generative Simulation Platform for Autonomous Driving
abstract
This paper presented DriveArena, the first high-fidelity closed-loop simulation system designed for driving agents navigating in real scenarios. DriveArena features a flexible, modular architecture, allowing for the seamless interchange of its core components: Traffic Manager, a traffic simulator capable of generating realistic traffic flow on any worldwide street map, and World Dreamer, a high-fidelity conditional generative model with infinite autoregression. This powerful synergy empowers any driving agent capable of processing real-world images to navigate in DriveArena's simulated environment. The agent perceives its surroundings through images generated by World Dreamer and output trajectories. These trajectories are fed into Traffic Manager, achieving realistic interactions with other vehicles and producing a new scene layout. Finally, the latest scene layout is relayed back into World Dreamer, perpetuating the simulation cycle. This iterative process fosters closed-loop exploration within a highly realistic environment, providing a valuable platform for developing and evaluating driving agents across diverse and challenging scenarios. DriveArena signifies a substantial leap forward in leveraging generative image data for the driving simulation platform, opening insights for closed-loop autonomous driving. Code will be available soon on GitHub: https://github.com/PJLab-ADG/DriveArena
Xuemeng Yang, Licheng Wen, Tiantian Wei, Yukai Ma, Jianbiao Mei, Xin Li 0110, Wenjie Lei, Daocheng Fu, Pinlong Cai, Min Dou, Liang He 0001, Yong Liu 0007, Botian Shi, Yu Qiao 0001
ICCV6
2025 UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
abstract
Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified formulation. Specifically, we constructed a dataset labeled with manipulation-related key attributes, comprising 900 articulated objects from 19 categories and 600 tools from 12 categories. Furthermore, we leverage MLLMs to infer object-centric representations for manipulation tasks, including affordance recognition and reasoning about 3D motion constraints. Comprehensive experiments in both simulation and real-world settings indicate that UniAff significantly improves the generalization of robotic manipulation for tools and articulated objects. We hope that UniAff will serve as a general baseline for unified robotic manipulation tasks in the future. Images, videos, dataset and code are published on the project website at:https://sites.google.com/view/uni-aff/home.
Qiaojun Yu, Siyuan Huang 0004, Xibin Yuan, Zhengkai Jiang 0001, Ce Hao, Xin Li 0110, Haonan Chang, Junbo Wang 0004, Liu Liu 0012, Hongsheng Li 0001, Peng Gao 0007, Cewu Lu
ICRA6
2025 Cost-Effective Topology Design for Network Planning in Industrial Time-Sensitive Networking
abstract
Time Sensitive Networking (TSN) has been widely considered as a promising networking technology in industrial fields as its capibility of deterministic transmission. One of the main challenges in TSN application is the complexity of network planning, including topology design, flow routing and scheduling schemes. In these three tasks, topology design plays a vital role in reducing costs and supporting the feasibility of routing and scheduling schemes. While recent researchers make some progress in routing and scheduling algorithms, most studies lack effective approaches for optimizing network topology, limiting the practical applicability of TSN. This paper addresses this problem by presenting a joint design method (JDM) for low-cost TSN topology design while also ensuring the feasibility of flow routing and scheduling. A unified mathematical model is developed to integrate TSN topology, routing and scheduling into a joint optimization problem, minimizing the overall network cost. On this basis, a one-hot vectorization technique is applied to linearize scheduling constraints to enhance the computational efficiency. Simulation results show that compared with other methods, the proposed JDM generates TSN topology at the lowest cost, ensuring flows deterministic transmission within minutes.
Yingxiu Chen, Xin Li 0110, Lei Xu 0043, Shihui Duan, Qimin Xu, Cailian Chen
INDIN4
2025 SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
abstract
Automating garment manipulation poses a significant challenge for assistive robotics due to the diverse and de-formable nature of garments. Traditional approaches typically require separate models for each garment type, which limits scalability and adaptability. In contrast, this paper presents a unified approach using vision-language models (VLMs) to improve keypoint prediction across various garment categories. By interpreting both visual and semantic information, our model enables robots to manage different garment states with a single model. We created a large-scale synthetic dataset using advanced simulation techniques, allowing scalable training without extensive real-world data. Experimental results indicate that the VLM-based method significantly enhances keypoint detection accuracy and task success rates, providing a more flexible and general solution for robotic garment manipulation. In addition, this research also underscores the potential of VLMs to unify various garment manipulation tasks within a single framework, paving the way for broader applications in home automation and assistive robotics in the future.
Xin Li 0110, Siyuan Huang 0004, Qiaojun Yu, Zhengkai Jiang 0001, Ce Hao, Yimeng Zhu, Hongsheng Li 0001, Peng Gao 0007, Cewu Lu
IROS1
2025 Capacity Analysis-Based Topology Planning and Traffic Scheduling for Time-Sensitive Networking
abstract
With the ability to provide deterministic transmission, time sensitive networking (TSN) has been widely used in various industrial scenarios. However, most of the existing TSN research focuses on traffic scheduling over predefined network topology. In industrial applications, optimizing network topology can reduce the number of network devices and the length of cables thereby lowering material and management costs, yet considering topology within TSN scheduling greatly increases the problem’s complexity. In this article, we first incorporate topology planning and traffic scheduling together into the TSN network design problem (TDP). A mathematical model for TDP is formulated, with the objective of minimizing the total weight and cost of a TSN network while satisfying the end-to-end deterministic transmission requirements. Subsequently, we establish the metric of capacity of TSN flow groups (CoG), and the proposed CoG estimation method enables feasibility assessment of TDP solutions. Since CoG measures a network’s capacity to accommodate TSN flows, we utilize it as the evaluation metric within our heuristic algorithm, CoG analysis based TSN network design algorithm (CATDA), reducing ineffective searches and enhancing the solution efficiency of TDP. Experiments show that compared to other algorithms, the proposed CATDA achieves the lowest-cost TSN network design solution and performs over 100 times faster than other algorithms.
Xin Li 0110, Lei Xu 0043, Qimin Xu, Cailian Chen, Xin-Ping Guan
IEEE Trans. Ind. Informatics2
2025 Scalable Scheduling in Time-Sensitive Networking: An Efficient Stream Conflict Detection Method
abstract
As an emerging communication technology, time-sensitive networking (TSN) holds the potential to enable real-time and deterministic interactions for streams within the Industrial Internet of Things. However, effectively and promptly scheduling large-scale streams in the TSN network poses a significant challenge due to high computational complexity. In this article, we conduct a schedulability analysis to preprocess the stream set with given routing paths, avoiding invalid searches and providing optimized guidance for stream routing. To accelerate the feasibility validation of potential solutions, an efficient stream conflict detection approach is proposed leveraging stream grouping with correlation analysis to compress the detection space. Integrating the above preprocess and efficient conflict detection, we develop a scalable scheduling algorithm with an incremental schedule synthesis to enhance scalability while ensuring low slot occupancy for all links. Evaluation results demonstrate that the proposed algorithm significantly reduces synthesis time and achieves low slot occupancy of all links compared to existing scheduling methods.
Lei Xu 0043, Cailian Chen, Yanzhou Zhang, Xin Li 0110, Shouliang Wang, Qimin Xu, Xin-Ping Guan
IEEE Trans. Ind. Informatics4
2024 Multi-Space Alignments Towards Universal LiDAR Segmentation
abstract
A unified and versatile LiDAR segmentation model with strong robustness and generalizability is desirable for safe autonomous driving perception. This work presents M3Net, a one-of-a-kind framework for fulfilling multitask, multi-dataset, multimodality LiDAR segmentation in a universal manner using just a single set of parameters. To better exploit data volume and diversity, we first combine large-scale driving datasets acquired by different types of sensors from diverse scenes and then conduct alignments in three spaces, namely data, feature, and label spaces, during the training. As a result, M3Net is capable of taming heterogeneous data for training state-of-the-art LiDAR segmentation models. Extensive experiments on twelve LiDAR segmentation datasets verify our effectiveness. Notably, using a shared set of parameters, M3Net achieves 75.1%,83.1%, and 72.4% mIoU scores, respectively, on the official benchmarks of SemanticKITTI, nuScenes, and Waymo Open.
Youquan Liu, Lingdong Kong, Xiaoyang Wu 0002, Runnan Chen, Xin Li 0110, Liang Pan, Ziwei Liu 0002, Yuexin Ma
CVPR5
2024 DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
abstract
Recent advancements in autonomous driving have relied on data-driven approaches, which are widely adopted but face challenges including dataset bias, overfitting, and uninterpretability. Drawing inspiration from the knowledge-driven nature of human driving, we explore the question of how to instill similar capabilities into autonomous driving systems and summarize a paradigm that integrates an interactive environment, a driver agent, as well as a memory component to address this question. Leveraging large language models (LLMs) with emergent abilities, we propose the DiLu framework, which combines a Reasoning and a Reflection module to enable the system to perform decision-making based on common-sense knowledge and evolve continuously. Extensive experiments prove DiLu's capability to accumulate experience and demonstrate a significant advantage in generalization ability over reinforcement learning-based methods. Moreover, DiLu is able to directly acquire experiences from real-world datasets which highlights its potential to be deployed on practical autonomous driving systems. To the best of our knowledge, we are the first to leverage knowledge-driven capability in decision-making for autonomous vehicles. Through the proposed DiLu framework, LLM is strengthened to apply knowledge and to reason causally in the autonomous driving domain. Project page: https://pjlab-adg.github.io/DiLu/
Licheng Wen, Daocheng Fu, Xin Li 0110, Xinyu Cai, Tao Ma 0002, Pinlong Cai, Min Dou, Botian Shi, Liang He 0001, Yu Qiao 0001
ICLR3
2024 A Timeslot Clustering-Based Hybrid Traffic Scheduling in Time-Sensitive Networking
abstract
The development of Industrial Internet of Things necessitates deterministic transmission of hybrid traffic with varying real-time requirements. Time-Sensitive Networking offers schemes combining scheduling mechanisms like the Time-Aware Shaper (TAS) and Cyclic Queuing and Forwarding (CQF). Current methods utilize the combination of TAS&CQF to schedule Time-Triggered (TT) and Audio-Video Bridging (AVB) traffic under corresponding mechanisms, respectively. However, the prior TT traffic orchestration restricts the scheduling solution space of AVB traffic, decreasing the effectiveness of combination and finally resulting in low schedulability. In this paper, we focus on enhancing the schedulability of TT&AVB scheduling by integrating TAS and Cyclic Specified Queuing and Forwarding (CSQF), which brings larger scheduling solution space to AVB traffic. The integration of TAS&CSQF is proposed by analyzing the impact of TAS on CSQF to expand the allocatable resource for AVB traffic. Then TAS&CSQF scheduling model is formulated to ensure deterministic requirements. Within this model, Timeslot Clustering Optimization (TCO) is proposed to optimize the scheduling solution space of AVB traffic and slice timeslot by refining the scheduling of TAS. Based on TCO, Flow Congestion Metric (FCM) is designed to improve schedulability by the congestion level of each link. An FCM-based algorithm is further presented to generate TT&AVB solutions incrementally. Simulation results demonstrate that our method reduces scheduling time cost by over 168 times and improves schedulability by 32% for 3000 flows compared to existing methods.
Shouliang Wang, Qimin Xu, Xin Li 0110, Cailian Chen
INDIN4
2024 Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving
abstract
Autonomous driving has advanced significantly due to sensors, machine learning, and artificial intelligence improvements. However, prevailing methods struggle with intricate scenarios and causal relationships, hindering adaptability and interpretability in varied environments. To address the above problems, we introduce LeapAD, a novel paradigm for autonomous driving inspired by the human cognitive process. Specifically, LeapAD emulates human attention by selecting critical objects relevant to driving decisions, simplifying environmental interpretation, and mitigating decision-making complexities. Additionally, LeapAD incorporates an innovative dual-process decision-making module, which consists of an Analytic Process (System-II) for thorough analysis and reasoning, along with a Heuristic Process (System-I) for swift and empirical processing. The Analytic Process leverages its logical reasoning to accumulate linguistic driving experience, which is then transferred to the Heuristic Process by supervised fine-tuning. Through reflection mechanisms and a growing memory bank, LeapAD continuously improves itself from past mistakes in a closed-loop environment. Closed-loop testing in CARLA shows that LeapAD outperforms all methods relying solely on camera input, requiring 1-2 orders of magnitude less labeled data. Experiments also demonstrate that as the memory bank expands, the Heuristic Process with only 1.8B parameters can inherit the knowledge from a GPT-4 powered Analytic Process and achieve continuous performance improvement. Project page: https://pjlab-adg.github.io/LeapAD
Jianbiao Mei, Yukai Ma, Xuemeng Yang, Licheng Wen, Xinyu Cai, Xin Li 0110, Daocheng Fu, Bo Zhang 0069, Pinlong Cai, Min Dou, Botian Shi, Liang He 0001, Yong Liu 0007, Yu Qiao 0001
NeurIPS6
2024 Robust depth completion based on Semantic Aggregation
Zhichao Fu, Xin Li 0110, Tianyu Huai, Daoguo Dong, Liang He 0001
Appl. Intell.2
2024 Cross-domain document layout analysis using document style guide
Xingjiao Wu, Luwei Xiao, Xiangcheng Du, Yingbin Zheng, Xin Li 0110, Tianlong Ma, Cheng Jin 0001, Liang He 0001
Expert Syst. Appl.5
2024 Determinacy-Oriented Task Offloading Scheduling Against DoS Attack for TSN-Based Edge Computing Architecture
abstract
The increasing scale of industrial production is leading to a greater demand for communication and computing capabilities, thereby increasing the likelihood of resource competition and conflict. This can cause stochastic overall task latency (including communication and computing latency), resulting in the occurrence of overdue tasks. To guarantee deterministic delays, the integration of edge computing (EC) with time-sensitive networking (TSN) emerges as a promising technology. However, the deterministic feature of TSN increases vulnerabilities to attacks within this integration. Particularly, uncertain denial-of-service (DoS) attacks can exhaust system resources and disrupt determinacy, causing prolonged delays, or communication failures. To this end, this article proposes an attack-tolerant TSN-based edge computing (TSN-EC) architecture to guarantee task determinacy. Based on the architecture, a task-level no-wait scheduling mechanism of TSN is proposed under packet switching mode, which ensures deterministic communication delays. A robust and deterministic task offloading scheduling (RDTOS) strategy is developed to minimize the number of overdue tasks by identifying the worst-case scenario of uncertain DoS attacks, considering each task's importance. To reduce computational complexity, a two-layer decomposition algorithm is proposed by further decomposing the master problem of the conventional C-CG algorithm. Experimental results conducted on a TSN-EC testbed demonstrate the superiority of the RDTOS strategy in enhancing security and providing overall task determinacy compared to related algorithms.
Xin Li 0110, Yingxiu Chen, Meihan Lin, Yonghui Liang, Cailian Chen, Qimin Xu, Xin-Ping Guan
IEEE Trans. Ind. Informatics1
2023 LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal Fusion
abstract
LiDAR-camera fusion methods have shown impressive performance in 3D object detection. Recent advanced multi-modal methods mainly perform global fusion, where image features and point cloud features are fused across the whole scene. Such practice lacks fine-grained region-level information, yielding suboptimal fusion performance. In this paper, we present the novel Local-to-Global fusion network (LoGoNet), which performs LiDAR-camerafusion at both local and global levels. Concretely, the Global Fusion (GoF) of LoGoNet is built upon previous literature, while we exclusively use point centroids to more precisely represent the position of voxel features, thus achieving better crossmodal alignment. As to the Local Fusion (LoF), we first divide each proposal into uniform grids and then project these grid centers to the images. The image features around the projected grid points are sampled to be fused with position-decorated point cloud features, maximally uti-lizing the rich contextual information around the proposals. The Feature Dynamic Aggregation (FDA) module is further proposed to achieve information interaction between these locally and globally fused features, thus producing more informative multi-modal features. Extensive experiments on both Waymo Open Dataset (WOD) and KITTI datasets show that LoGoNet outperforms all state-of-the-art 3D detection methods. Notably, LoGoNet ranks 1st on Waymo 3D object detection leaderboard and obtains 81.02 mAPH (L2) detection performance. It is noteworthy that, for the first time, the detection performance on three classes surpasses 80 APH (L2) simultaneously. Code will be available at https://github.com/sankin97/LoGoNet.
Xin Li 0110, Tao Ma 0002, Yuenan Hou, Botian Shi, Yuchen Yang 0003, Youquan Liu, Xingjiao Wu, Qin Chen 0001, Yikang Li 0002, Yu Qiao 0001, Liang He 0001
CVPR1
2023 SCPNet: Semantic Scene Completion on Point Cloud
abstract
Training deep models for semantic scene completion (SSC) is challenging due to the sparse and incomplete input, a large quantity of objects of diverse scales as well as the inherent label noise for moving objects. To address the above-mentioned problems, we propose the following three solutions: 1) Redesigning the completion sub-network. We design a novel completion sub-network, which consists of several Multi-Path Blocks (MPBs) to aggregate multi-scale features and is free from the lossy downsampling operations. 2) Distilling rich knowledge from the multi-frame model. We design a novel knowledge distillation objective, dubbed Dense-to-Sparse Knowledge Distillation (DSKD). It transfers the dense, relation-based semantic knowledge from the multi-frame teacher to the single-frame student, significantly improving the representation learning of the single-frame model. 3) Completion label rectification. We propose a simple yet effective label rectification strategy, which uses off-the-shelf panoptic segmentation labels to remove the traces of dynamic objects in completion labels, greatly improving the performance of deep models especially for those moving objects. Extensive experiments are conducted in two public SSC benchmarks, i.e., SemanticKITTI and SemanticPOSS. Our SCPNet ranks 1st on SemanticKITTI semantic scene completion challenge and surpasses the competitive S3CNet [3] by 7.2 mIoU. SCP-Net also outperforms previous completion algorithms on the SemanticPOSS dataset. Besides, our method also achieves competitive results on SemanticKITTI semantic segmentation tasks, showing that knowledge learned in the scene completion is beneficial to the segmentation task.
Zhaoyang Xia, Youquan Liu, Xin Li 0110, Xinge Zhu, Yuexin Ma, Yikang Li 0002, Yuenan Hou, Yu Qiao 0001
CVPR3
2023 Robo3D: Towards Robust and Reliable 3D Perception against Corruptions
abstract
The robustness of 3D perception systems under natural corruptions from environments and sensors is pivotal for safety-critical applications. Existing large-scale 3D perception datasets often contain data that are meticulously cleaned. Such configurations, however, cannot reflect the reliability of perception models during the deployment stage. In this work, we present Robo3D, the first comprehensive benchmark heading toward probing the robustness of 3D detectors and segmentors under out-of-distribution scenarios against natural corruptions that occur in real-world environments. Specifically, we consider eight corruption types stemming from severe weather conditions, external disturbances, and internal sensor failure. We uncover that, although promising results have been progressively achieved on standard benchmarks, state-of-the-art 3D perception models are at risk of being vulnerable to corruptions. We draw key observations on the use of data representations, augmentation schemes, and training strategies, that could severely affect the model's performance. To pursue better robustness, we propose a density-insensitive training framework along with a simple flexible voxelization strategy to enhance the model resiliency. We hope our benchmark and approach could inspire future research in designing more robust and reliable 3D perception models. Our robustness benchmark suite is publicly available1.
Lingdong Kong, Youquan Liu, Xin Li 0110, Runnan Chen, Jiawei Ren 0001, Liang Pan, Kai Chen 0026, Ziwei Liu 0002
ICCV3
2023 UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase
abstract
Point-, voxel-, and range-views are three representative forms of point clouds. All of them have accurate 3D measurements but lack color and texture information. RGB images are a natural complement to these point cloud views and fully utilizing the comprehensive information of them benefits more robust perceptions. In this paper, we present a unified multi-modal LiDAR segmentation network, termed UniSeg, which leverages the information of RGB images and three views of the point cloud, and accomplishes semantic segmentation and panoptic segmentation simultaneously. Specifically, we first design the Learnable cross-Modal Association (LMA) module to automatically fuse voxel-view and range-view features with image features, which fully utilize the rich semantic information of images and are robust to calibration errors. Then, the enhanced voxel-view and range-view features are transformed to the point space, where three views of point cloud features are further fused adaptively by the Learnable cross-View Association module (LVA). Notably, UniSeg achieves promising results in three public benchmarks, i.e., SemanticKITTI, nuScenes, and Waymo Open Dataset (WOD); it ranks 1st on two challenges of two benchmarks, including the LiDAR semantic segmentation challenge of nuScenes and panoptic segmentation challenges of SemanticKITTI. Besides, we construct the OpenPCSeg codebase, which is the largest and most comprehensive outdoor LiDAR segmentation codebase. It contains most of the popular outdoor LiDAR segmentation algorithms and provides reproducible implementations. The OpenPCSeg codebase will be made publicly available at https://github.com/PJLab-ADG/PCSeg.
Youquan Liu, Runnan Chen, Xin Li 0110, Lingdong Kong, Yuchen Yang 0003, Zhaoyang Xia, Yeqi Bai, Xinge Zhu, Yuexin Ma, Yikang Li 0002, Yu Qiao 0001, Yuenan Hou
ICCV3
2023 DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds
abstract
Existing offboard 3D detectors always follow a modular pipeline design to take advantage of unlimited sequential point clouds. We have found that the full potential of off-board 3D detectors is not explored mainly due to two reasons: (1) the onboard multi-object tracker cannot generate sufficient complete object trajectories, and (2) the motion state of objects poses an inevitable challenge for the object-centric refining stage in leveraging the long-term temporal context representation. To tackle these problems, we propose a novel paradigm of offboard 3D object detection, named DetZero. Concretely, an offline tracker coupled with a multi-frame detector is proposed to focus on the completeness of generated object tracks. An attention-mechanism refining module is proposed to strengthen contextual information interaction across long-term sequential point clouds for object refining with decomposed regression methods. Extensive experiments on Waymo Open Dataset show our DetZero outperforms all state-of-the-art onboard and offboard 3D detection methods. Notably, DetZero ranks 1st place on Waymo 3D object detection leaderboard1with 85.15 mAPH (L2) detection performance. Further experiments validate the application of taking the place of human labels with such high-quality results. Our empirical study leads to rethinking conventions and interesting findings that can guide future research on offboard 3D object detection.
Tao Ma 0002, Xuemeng Yang, Hongbin Zhou, Xin Li 0110, Botian Shi, Yuchen Yang 0003, Zhizheng Liu, Liang He 0001, Yu Qiao 0001, Yikang Li 0002, Hongsheng Li 0001
ICCV4
2022 Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection
Xin Li 0110, Botian Shi, Yuenan Hou, Xingjiao Wu, Tianlong Ma, Yikang Li 0002, Liang He 0001
ECCV (38)1
2022 Integrated Localization and Tracking for AUV With Model Uncertainties via Scalable Sampling-Based Reinforcement Learning Approach
abstract
This article studies the joint localization and tracking issue for the autonomous underwater vehicle (AUV), with the constraints of asynchronous time clock in cyberchannels and model uncertainty in physical channels. More specifically, we develop a reinforcement learning (RL)-based asynchronous localization algorithm to localize the position of AUV, where the time clock of AUV is not required to be well synchronized with the real time. Based on the estimated position, a scalable sampling strategy called multivariate probabilistic collocation method with orthogonal fractional factorial design (M-PCM-OFFD) is employed to evaluate the time-varying uncertain model parameters of AUV. After that, an RL-based tracking controller is designed to drive AUV to the desired target point. Besides that, the performance analyses for the integration solution are also presented. Of note, the advantages of our solution are highlighted as: 1) the RL-based localization algorithm can avoid local optimal in traditional least-square methods; 2) the M-PCM-OFFD-based sampling strategy can address the model uncertainty and reduce the computational cost; and 3) the integration design of localization and tracking can reduce the communication energy consumption. Finally, simulation and experiment demonstrate that the proposed localization algorithm can effectively eliminate the impact of asynchronous clock, and more importantly, the integration of M-PCM-OFFD in the RL-based tracking controller can find accurate optimization solutions with limited computational costs.
Jing Yan 0001, Xin Li 0110, Xian Yang 0002, Xiaoyuan Luo, Changchun Hua, Xin-Ping Guan
IEEE Trans. Syst. Man Cybern. Syst.2
2020 A Real-Time Deep Network for Crowd Counting
abstract
Automatic analysis of highly crowded people has attracted extensive attention from computer vision research. Previous approaches for crowd counting have already achieved promising performance across various benchmarks. However, to deal with the real situation, we hope the model run as fast as possible while keeping accuracy. In this paper, we propose a compact convolutional neural network for crowd counting which learns a more efficient model with a small number of parameters. With three parallel filters executing the convolutional operation on the input image simultaneously at the front of the network, our model could achieve nearly real-time speed and save more computing resources. Experiments on two benchmarks show that our proposed method not only takes a balance between performance and efficiency which is more suitable for actual scenes but also is superior to existing light-weight models in speed.
Xiaowen Shi, Xin Li 0110, Caili Wu, Shuchen Kong, Jing Yang 0023, Liang He 0001
ICASSP2
2020 Margin Guidance Network for Arbitrary-shaped Scene Text Detection
abstract
Segmentation-based scene text detection approaches have been adopted to arbitrary-shaped texts and have achieved a great progress. However, false detection always easily exist when the arbitrary-shaped texts are close to each other. In this paper, we propose the Margin Guidance Network (MGN) that mainly based on the margin constraint residual module (MCRM) to address aforementioned problem. The MCRM considers the margins between multiple text instance masks to guide the training of network and improve the performance on text detection. The MCRM contains two prediction branch, the one can generate the multiple different scale of masks for a text instance and the other branch is used to generate multiple margins between the above masks. Experimental results on three public benchmarks including ICDAR2015, CTW1500 and Total-Text have demonstrated that the proposed MGN achieves the state-of-the-art results.
Xin Li 0110, Xingjiao Wu, Tianlong Ma, Zhao Zhou, Luhui Chen, Liang He 0001
ICTAI1