Kaicheng Yu

dblp:198/0861 · DBLP profile ↗
← Back
37ranked-venue papers
7as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 14 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HEV Generative Sandbox: A Framework for Assessing Domain-Specific Social Risks Through Human-LLM Simulation
abstract
Deploying Large Language Models (LLMs) in specialized domains introduces significant societal and compliance risks, including bias amplification, misinformation propagation, and privacy violations. These risks predominantly emerge from the dynamic interactions between LLMs and humans in specific contexts. Different domains face unique distribution of hazards, and varying interaction modalities introduce distinct levels of exposure and vulnerability. However, current risk assessment frameworks lack a systematic methodology to capture this dynamic interplay. In this work, we introduce the HEV Generative Sandbox, a novel risk evaluation framework that simulates human-LLM behavior to quantify domain-contextual risks across three interdependent dimensions: 1) Hazard (H): Domain-specific threats inherent to a given context; 2) Exposure (E): The extent to which the LLM and its users are subjected to hazardous scenarios; 3) Vulnerability (V): The susceptibility of the system to risk due to human interaction or model weaknesses. Our approach pioneers "domain-rooted scenario generation", wherein we sample contextual distributions from domain-specific corpora and simulate diverse inputs. By unifying dynamic scenario simulation, causal risk decomposition, and closed-loop evaluation, the HEV Generative Sandbox provides a scalable, domain-sensitive methodology for responsible LLM deployment. This work contributes to advancing the safe deployment of LLMs by providing a comprehensive and automated risk evaluation framework.
Zhiyi Hou, Xiaoang Xu, Shuo Wang 0013, Huijia Wu, Kaicheng Yu, Yang Yu 0011, ChengXiang Zhai
AAAI6
2026 CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving
abstract
End-to-end planning methods are the de-facto standard of the current autonomous driving system, while the robustness of the data-driven approaches suffers due to the notorious long-tail problem (i.e., rare but safety-critical failure cases). In this work, we explore whether recent diffusion-based video generation methods (a.k.a. world models), paired with structured 3D layouts, can enable a fully automated pipeline to self-correct such failure cases. We first introduce an agent to simulate the role of product manager, dubbed PM-Agent, which formulates data requirements to collect data similar to the failure cases. Then, we use a generative model that can simulate both data collection and annotation. However, existing generative models struggle to generate high-fidelity data conditioned on 3D layouts. To address this, we propose DriveSora, which can generate spatiotemporally consistent videos aligned with the 3D annotations requested by PM-Agent. We integrate these components into our self-correcting agentic system, CorrectAD. Importantly, our pipeline is end-to-end model agnostic and can be applied to improve any end-to-end planner. Evaluated on both nuScenes and a more challenging in-house dataset across multiple end-to-end planners, CorrectAD corrects 62.5% and 49.8% of failure cases, reducing collision rates by 39% and 27%, respectively.
Enhui Ma, Junpeng Jiang, Kun Zhan, Xueyang Zhang, Xianpeng Lang, Di Lin 0002, Kaicheng Yu
AAAI14
2026 Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
abstract
Jiahuan Zhang, Shunwen Bai, Tianheng Wang, KaiWen Guo, Zijia Song, Hanqing WU, Guozheng Rao, Kai Han, Kaicheng Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shunwen Bai, Tianheng Wang, Zijia Song, Guozheng Rao, Kaicheng Yu
ACL (1)9
2026 MG-LDM: A multimodal guided latent diffusion model with Mamba-based temporal encoding for inverse topological design of tissue engineering skin substitutes
Kaicheng Yu, Zexue Lin, Qiang Gao 0016, Guoyin Shang
Adv. Eng. Informatics1
2026 TO-MambaLDM: Multi-physics informed Mamba-enhanced latent diffusion for cooling structure topology optimization
Kaicheng Yu, Lijie Su, Swee Leong Sing
Adv. Eng. Informatics2
2026 TriTopo-LGDM: A reverse design method for trabecular bone scaffolds integrating topology optimization and latent graph diffusion models
Kaicheng Yu, Lijie Su, Yifeng Yao, Swee Leong Sing
Adv. Eng. Informatics2
2026 A transformer-based framework for cross-material in situ monitoring in extrusion-based bioprinting
Kaicheng Yu, Yifeng Yao, Qiang Gao 0016, Guoyin Shang, Swee Leong Sing
Adv. Eng. Informatics2
2026 Aero-MambaNet: A unified mesh-mamba framework for physics-consistent aerodynamic surrogate modeling
Kaicheng Yu, Lijie Su, Swee Leong Sing
Expert Syst. Appl.2
2026 Magnetic navigation of photoacoustic/ultrasound catheters via vision-ultrasound fusion servo control for embodied medical robots
Kaicheng Yu, Dongjian Wu, Zhiping Yu, Haibo Cong, Mingjian Sun
Pattern Recognit. Lett.1
2026 Dynamic Registration-Based Photoacoustic Endoscopic Temperature Imaging for Precision Interventional Thermal Therapy and Monitoring
abstract
Atherosclerosis is a major cause of cardiovascular disease. Photothermal ablation provides a minimally invasive therapeutic approach but remains constrained by the lack of reliable temperature monitoring. Conventional thermometry provides only single-point readings and is easily affected by contact and flow, limiting spatial accuracy. Although array-based photoacoustic thermometry allows non-contact detection, it remains bulky and unsuitable for catheter integration. This study realizes real-time intravascular temperature imaging and monitoring through dynamic registration-corrected photoacoustic endoscopy (PETI-DRC). The system employs a miniaturized dual-modality catheter achieving photoacoustic/ultrasound imaging and angularly registered temperature mapping compatible with clinical interventions. Experiments on phantoms and ex vivo rabbit aortas quantitatively verified frame registration accuracy and temperature estimation reliability through comparison with thermocouple references and microscopic validation. Results show improved inter-frame consistency (average SSIM = 0.9684) and high temperature accuracy (RMSE $= 0.65~^{\circ }$ C), enabling stable reconstruction of angularly registered temperature images and reliable tracking of thermal dynamics for real-time temperature monitoring. This method provides accurate thermal feedback for safer and more effective photothermal therapy and offers a promising direction for future intravascular interventions.
Dongjian Wu, Kaicheng Yu, Haokun Zhang, Jianwu Zhu, Xiaojing Gong, Mingjian Sun
IEEE Trans. Medical Imaging2
2025 SR-LLM: Rethinking the Structured Representation in Large Language Model
abstract
Jiahuan Zhang, Tianheng Wang, Ziyi Huang, Yulong Wu, Hanqing Wu, DongbaiChen DongbaiChen, Linfeng Song, Yue Zhang, Guozheng Rao, Kaicheng Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Tianheng Wang, DongbaiChen DongbaiChen, Linfeng Song, Yue Zhang 0031, Guozheng Rao, Kaicheng Yu
ACL (1)10
2025 CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
abstract
Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need for a real-world benchmark to evaluate the ability of LLM agents to exploit web application vulnerabilities. However, existing benchmarks fall short as they are limited to abstracted Capture-the-Flag competitions or lack comprehensive coverage. Building a benchmark for real-world vulnerabilities involves both specialized exper- tise to reproduce exploits and a systematic approach to evaluating unpredictable attacks. To address this challenge, we introduce CVE-Bench, a real-world cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Our experiments show that the state-of-the-art agent framework can exploit up to 13% of the vulnerabilities.
Yuxuan Zhu 0003, Antony Kellermann, Dylan Bowman, Philip Li, Akul Gupta, Adarsh Danda, Richard Fang, Conner Jensen, Eric Ihli, Jason Benn, Jet Geronimo, Avi Dhir, Sudhit Rao, Kaicheng Yu, Twm Stone, Daniel Kang 0001
ICML14
2025 Design and Development of a Deformable Spherical Robot for Amphibious Applications*
abstract
This paper presents a deformable spherical robot with a six-strut topological structure capable of achieving multimodal locomotion in complex amphibious environments. The robot realizes isotropic rolling and asymmetric jumping through its innovative geometric-based configuration while integrating an airbag-driven module for underwater buoyancy control. Based on collision dynamics analysis, we develop a prototype of the deformable spherical robot. Experiments conducted on land, in transition zones, and underwater validate the robot’s multimodal locomotion feasibility in multi-medium environments.
Ruoyu Ren, Hao Lee, Hank Zhang, Kaicheng Yu, Siyu Gao
IROS6
2025 CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting
abstract
Vehicle-to-everything (V2X) communication plays a crucial role in autonomous driving, enabling cooperation between vehicles and infrastructure. While simulation has significantly contributed to various autonomous driving tasks, its potential for data generation and augmentation in V2X scenarios remains underexplored. In this paper, we introduce CRUISE, a comprehensive reconstruction-and-synthesis framework designed for V2X driving environments. CRUISE employs decomposed Gaussian Splatting to accurately reconstruct real-world scenes while supporting flexible editing. By decomposing dynamic traffic participants into editable Gaussian representations, CRUISE allows for seamless modification and augmentation of driving scenes. Furthermore, the framework renders images from both ego-vehicle and infrastructure views, enabling large-scale V2X dataset augmentation for training and evaluation. Our experimental results demonstrate that: 1) CRUISE reconstructs real-world V2X driving scenes with high fidelity; 2) using CRUISE improves 3D detection across ego-vehicle, infrastructure, and cooperative views, as well as cooperative 3D tracking on the V2X-Seq benchmark; and 3) CRUISE effectively generates challenging corner cases. The code will be publicly available at https://github.com/SainingZhang/CRUISE.
Haoran Xu 0003, Saining Zhang, Peishuo Li, Baijun Ye, Xiaoxue Chen, Huan-ang Gao, Jv Zheng, Ziqiao Peng, Run Miao, Jinrang Jia, Yifeng Shi, Guangqi Yi, Hang Zhao 0021, Hao Tang 0005, Hongyang Li 0001, Kaicheng Yu, Hao Zhao 0002
IROS17
2025 OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
abstract
Autonomous driving has seen remarkable advancements, largely driven by extensive real-world data collection. However, acquiring diverse and corner-case data remains costly and inefficient. Generative models have emerged as a promising solution by synthesizing realistic sensor data. However, existing approaches primarily focus on single-modality generation, leading to inefficiencies and misalignment in multimodal sensor data. To address these challenges, we propose OminiGen, which generates aligned multimodal sensor data in a unified framework. Our approach leverages a shared Bird's Eye View (BEV) space to unify multimodal features and designs a novel generalizable multimodal reconstruction method, UAE, to jointly decode LiDAR and multi-view camera data. UAE achieves multimodal sensor decoding through volume rendering, enabling accurate and flexible reconstruction. Furthermore, we incorporate a Diffusion Transformer (DiT) with a ControlNet branch to enable controllable multimodal sensor generation. Our comprehensive experiments demonstrate that OminiGen achieves desired performances in unified multimodal sensor data generation with multimodal consistency and flexible sensor adjustments.
Enhui Ma, Tianyi Yan, Xueyang Zhang, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Jiawang Bian, Kaicheng Yu, Xiaodan Liang
ACM Multimedia11
2025 Semantic-DARTS: Elevating Semantic Learning for Mobile Differentiable Architecture Search
abstract
Differentiable architecture search (DARTS) is a prevailing direction in automatic machine learning, but it may suffer from performance collapse and generalization issues. Recent efforts mitigate them by integrating regularization into architectural parameters or rule-based operations selection. These efforts primarily emphasize learning the global class-specific features through the image classification task, while overlooking the fine-grained local information during the search process. In this article, we take the first trial to observe that three semantic challenges arise from the classification-based DARTS: 1) inaccurate class-specific features; 2) partial target attention; and 3) blurred semantic regions. To tackle them in one shot, we propose Semantic-DARTS, combining the masked image modeling (MIM) paradigm with the classification task to incorporate local semantic information into the architecture search. Specifically, we design a lightweight reconstruction head that recovers the corrupted image based on the condensed latent feature, which learns both the local semantics and their relationship patch-wisely. Simultaneously, the concurrent classification head strengthens the connection between the global category of the target and the local semantics of their parts. As evidenced by our experiments, the proposed approach achieves state-of-the-art results on CIFAR-10, CIFAR-100, and ImageNet. Furthermore, the searched model is not only able to improve global class-specific features but also to capture fine-grained local representations, improving both the classification performance and the generalization ability.
Bicheng Guo, Shibo He, Miaojing Shi, Kaicheng Yu, Jiming Chen 0001, Xuemin Shen
IEEE Internet Things J.4
2025 BEVHeight++: Toward Robust Visual Centric 3D Object Detection
abstract
While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric detection methods perform poorly on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight++, to address this issue. In essence, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. By incorporating both height and depth encoding techniques, we achieve a more accurate and robust projection from 2D to BEV spaces. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. In terms of the ego-vehicle scenario, BEVHeight++ surpasses depth-only methods with increases of +2.8% NDS and +1.7% mAP on the nuScenes test set, and even higher gains of +9.3% NDS and +8.8% mAP on the nuScenes-C benchmark with object-level distortion. Consistent and substantial performance improvements are achieved across the KITTI, KITTI-360, and Waymo datasets as well.
Lei Yang 0060, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Yi Huang 0038, Xinyu Zhang 0001, Kaicheng Yu
IEEE Trans. Pattern Anal. Mach. Intell.11
2024 AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis
abstract
Neural implicit fields have been a de facto standard in novel view synthesis. Recently, there exist some methods exploring fusing multiple modalities within a single field, aiming to share implicit features from different modalities to enhance reconstruction performance. However, these modalities often exhibit misaligned behaviors: optimizing for one modality, such as LiDAR, can adversely affect another, like camera performance, and vice versa. In this work, we conduct comprehensive analyses on the multimodal implicit field of LiDAR-camera joint synthesis, revealing the underlying issue lies in the misalignment of different sensors. Furthermore, we introduce AlignMiF, a geometrically aligned multimodal implicit field with two proposed modules: Geometry-Aware Alignment (GAA) and Shared Geometry Initialization (SGI). These modules effectively align the coarse geometry across different modalities, significantly enhancing the fusion process between LiDAR and camera data. Through extensive experiments across various datasets and scenes, we demonstrate the effectiveness of our approach in facilitating better interaction between LiDAR and camera modalities within a unified neural field. Specifically, our proposed AlignMiF, achieves remarkable improvement over recent implicit fusion methods (+2.01 and +3.11 image PSNR on the KITTI-360 and Waymo datasets) and consistently surpasses single modality performance (13.8% and 14.2% reduction in LiDAR Cham-fer Distance on the respective datasets). Code release: https://github.com/tangtaogo/alignmif.
Tang Tao, Guangrun Wang, Yixing Lao, Peng Chen 0054, Liang Lin 0004, Kaicheng Yu, Xiaodan Liang
CVPR7
2024 Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt Training
abstract
The rapid advancement of deep learning models is often attributed to their ability to leverage massive training data. In contrast, such privilege has not yet fully benefited 3D deep learning, mainly due to the limited availability of large-scale 3D datasets. Merging multiple available data sources and letting them collaboratively train a single model is a potential solution. However, due to the large domain gap between 3D point cloud datasets, such mixed supervision could adversely affect the model's performance and lead to degenerated performance (i.e., negative transfer) compared to single-dataset training. In view of this challenge, we introduce Point Prompt Training (PPT), a novel framework for multi-dataset synergistic learning in the context of 3D representation learning that supports multiple pre-training paradigms. Based on this framework, we propose Prompt-driven Normalization, which adapts the model to different datasets with domain-specific prompts and Language-guided Categorical Alignment that decently unifies the multiple-dataset label spaces by leveraging the relationship between label text. Extensive experiments verify that PPT can overcome the negative transfer associated with synergistic learning and produce generalizable representations. Notably, it achieves state-of-the-art performance on each dataset using a single weight-shared model with supervised multi-dataset training. Moreover, when served as a pre-training framework, it outperforms other pre-training approaches regarding representation quality and attains remarkable state-of-the-art performance across over ten diverse downstream tasks spanning both indoor and outdoor 3D scenarios.
Xiaoyang Wu 0002, Zhuotao Tian, Xin Wen 0004, Bohao Peng, Xihui Liu, Kaicheng Yu, Hengshuang Zhao
CVPR6
2024 OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection
Hu Zhang 0005, Xin Yu 0002, Zi Huang, Kaicheng Yu
ECCV (84)7
2024 LiDAR-NeRF: Novel LiDAR View Synthesis via Neural Radiance Fields
Tang Tao, Longfei Gao, Guangrun Wang, Yixing Lao, Peng Chen 0054, Hengshuang Zhao, Dayang Hao, Xiaodan Liang, Mathieu Salzmann, Kaicheng Yu
ACM Multimedia10
2024 Tutorial: Large Language-Vision Model in Society
abstract
The tutorial "Large Vision-Language Model in the Society" aims to provide a comprehensive overview of state-of-the-art techniques and applications of large vision-language models (LVLMs), which integrate visual and textual data to transform multimedia research and applications. LVLMs are poised to revolutionize domains such as content creation, social media analysis, education, healthcare, and entertainment by enabling sophisticated content analysis, retrieval, and generation. This tutorial will cover the fundamentals of vision-language integration, state-of-the-art models, training techniques, applications, ethical considerations, and future directions. It is designed to be educational and instructive, providing an in-depth introduction rather than a cursory survey. Attendees will gain practical skills, and insights into the latest research, and engage in interactive sessions to reinforce learning. By addressing both technical and societal aspects, the tutorial will significantly benefit the multimedia community, driving innovation and progress in the field.
Kaicheng Yu, Siyuan Qi, Dongfang Liu
ACM Multimedia1
2024 LiT: Unifying LiDAR "Languages" with LiDAR Translator
abstract
LiDAR data exhibits significant domain gaps due to variations in sensors, vehicles, and driving environments, creating “language barriers” that limit the effective use of data across domains and the scalability of LiDAR perception models. To address these challenges, we introduce the LiDAR Translator (LiT), a framework that directly translates LiDAR data across domains, enabling both cross-domain adaptation and multi-domain joint learning. LiT integrates three key components: a scene modeling module for precise foreground and background reconstruction, a LiDAR modeling module that models LiDAR rays statistically and simulates ray-drop, and a fast, hardware-accelerated ray casting engine. LiT enables state-of-the-art zero-shot and unified domain detection across diverse LiDAR datasets, marking a step toward data-driven domain unification for autonomous driving systems. Source code and demos are available at: https://yxlao.github.io/lit.
Yixing Lao, Xiaoyang Wu 0002, Peng Chen 0054, Kaicheng Yu, Hengshuang Zhao
NeurIPS5
2023 BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection
abstract
While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric bird's eye view detection methods have inferior performances on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight, to address this issue. In essence, instead of predicting the pixel-wise depth, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. The code is available at https://github.com/ADLab-AutoDrive/BEVHeight.
Lei Yang 0060, Kaicheng Yu, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Xinyu Zhang 0001
CVPR2
2023 Painting 3D Nature in 2D: View Synthesis of Natural Scenes from a Single Semantic Mask
abstract
We introduce a novel approach that takes a single semantic mask as input to synthesize multi-view consistent color images of natural scenes, trained with a collection of single images from the Internet. Prior works on 3D-aware image synthesis either require multi-view supervision or learning category-level prior for specific classes of objects, which are inapplicable to natural scenes. Our key idea to solve this challenge is to use a semantic field as the intermediate representation, which is easier to reconstruct from an input semantic mask and then translated to a radiance field with the assistance of off-the-shelf semantic image synthesis models. Experiments show that our method outperforms baseline methods and produces photorealistic and multi-view consistent videos of a variety of natural scenes. The project website is https://zju3dv.github.io/paintingnature/.
Shangzhan Zhang, Sida Peng, Tianrun Chen, Linzhan Mou, Haotong Lin, Kaicheng Yu, Yiyi Liao, Xiaowei Zhou 0001
CVPR6
2023 Mendam: Multi-Expert Network with Distribution-Aware Momentum for Long-Tailed Recognition
abstract
Long-tailed data distribution (i.e., minority classes occupy most of the data, while most classes have very few samples) is a common problem in image classification. The existing deferred re-balancing methods suffer from low accuracy for tail classes due to their bias towards head classes. In this paper, we find that properly adjusting the momentum can significantly improve the performance for class-imbalanced tasks, which provides a novel perspective to solve this thorny problem. Based on this finding, we propose a new way to calculate a distribution-aware momentum (DAM) based on the data distribution to avoid bias towards head classes. Furthermore, we design a multi-expert network (MEN) to designate each expert to focus on different tail classes. We achieve new state-of-the-art on CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT and iNaturalist2018 for image classification.
Qingheng Zhang, Haibo Ye, Kaicheng Yu
ICASSP3
2022 Knowledge Distillation via the Target-aware Transformer
abstract
Knowledge distillation becomes a de facto standard to improve the performance of small neural networks. Most of the previous works propose to regress the representational features from the teacher to the student in a one-to-one spatial matching fashion. However, people tend to overlook the fact that, due to the architecture differences, the semantic information on the same spatial location usually vary. This greatly undermines the underlying assumption of the one-to-one distillation approach. To this end, we propose a novel one-to-all spatial matching knowledge distillation approach. Specifically, we allow each pixel of the teacher feature to be distilled to all spatial locations of the student features given its similarity, which is generated from a target-aware transformer. Our approach surpasses the state-of-the-art methods by a significant margin on various computer vision benchmarks, such as ImageNet, Pascal VOC and COCOStuff10k. Code is available at https://github.com/sihaoevery/TaT.
Sihao Lin, Hongwei Xie, Bing Wang 0013, Kaicheng Yu, Xiaojun Chang, Xiaodan Liang
CVPR4
2022 NAS-Bench-Suite: NAS Evaluation is (Now) Surprisingly Easy
Yash Mehta, Colin White, Arber Zela, Arjun Krishnakumar, Guri Zabergja, Shakiba Moradian, Mahmoud Safari, Kaicheng Yu, Frank Hutter
ICLR8
2022 BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework
abstract
Fusing the camera and LiDAR information has become a de-facto standard for 3D object detection tasks. Current methods rely on point clouds from the LiDAR sensor as queries to leverage the feature from the image space. However, people discovered that this underlying assumption makes the current fusion framework infeasible to produce any prediction when there is a LiDAR malfunction, regardless of minor or major. This fundamentally limits the deployment capability to realistic autonomous driving scenarios. In contrast, we propose a surprisingly simple yet novel fusion framework, dubbed BEVFusion, whose camera stream does not depend on the input of LiDAR data, thus addressing the downside of previous methods. We empirically show that our framework surpasses the state-of-the-art methods under the normal training settings. Under the robustness training settings that simulate various LiDAR malfunctions, our framework significantly surpasses the state-of-the-art methods by 15.7% to 28.9% mAP. To the best of our knowledge, we are the first to handle realistic LiDAR malfunction and can be deployed to realistic scenarios without any post-processing procedure.
Tingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia, Yongtao Wang, Zhi Tang 0001
NeurIPS3
2022 An Analysis of Super-Net Heuristics in Weight-Sharing NAS
abstract
Weight sharing promises to make neural architecture search (NAS) tractable even on commodity hardware. Existing methods in this space rely on a diverse set of heuristics to design and train the shared-weight backbone network, a.k.a. the super-net. Since heuristics substantially vary across different methods and have not been carefully studied, it is unclear to which extent they impact super-net training and hence the weight-sharing NAS algorithms. In this paper, we disentangle super-net training from the search algorithm, isolate 14 frequently-used training heuristics, and evaluate them over three benchmark search spaces. Our analysis uncovers that several commonly-used heuristics negatively impact the correlation between super-net and stand-alone performance, whereas simple, but often overlooked factors, such as proper hyper-parameter settings, are key to achieve strong performance. Equipped with this knowledge, we show that simple random search achieves competitive performance to complex state-of-the-art NAS algorithms when the super-net is properly trained.
Kaicheng Yu, René Ranftl, Mathieu Salzmann
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Landmark Regularization: Ranking Guided Super-Net Training in Neural Architecture Search
abstract
Weight sharing has become a de facto standard in neural architecture search because it enables the search to be done on commodity hardware. However, recent works have empirically shown a ranking disorder between the performance of stand-alone architectures and that of the corresponding shared-weight networks. This violates the main assumption of weight-sharing NAS algorithms, thus limiting their effectiveness. We tackle this issue by proposing a regularization term that aims to maximize the correlation between the performance rankings of the shared-weight network and that of the standalone architectures using a small set of landmark architectures. We incorporate our regularization term into three different NAS algorithms and show that it consistently improves performance across algorithms, search-spaces, and tasks.
Kaicheng Yu, René Ranftl, Mathieu Salzmann
CVPR1
2021 Pyramid Architecture Search for Real-Time Image Deblurring
abstract
Multi-scale and multi-patch deep models have been shown effective in removing blurs of dynamic scenes. However, these methods still suffer from one major obstacle: manually designing a lightweight and high-efficiency network is challenging and time-consuming. To tackle this obstacle, we propose a novel deblurring method, dubbed PyNAS (pyramid neural architecture search network), towards automatically designing hyper-parameters including the scales, patches, and standard cell operators. The proposed PyNAS adopts gradient-based search strategies and innovatively searches the hierarchy patch and scale scheme not limited to cell searching. Specifically, we introduce a hierarchical search strategy tailored to the multi-scale and multi-patch deblurring task. The strategy follows the principle that the first distinguishes between the top-level (pyramid-scales and pyramid-patches) and bottom-level variables (cell operators) and then searches multi-scale variables using the top-to-bottom principle. During the search stage, PyNAS employs an early stopping strategy to avoid the collapse and computational issues. Furthermore, we use a path-level binarization mechanism for multi-scale cell searching to save the memory consumption. Our primary contribution is a real-time deblurring algorithm (around 58 fps) for 720p images while achieves state-of-the-art deblurring performance on the GoPro and Video Deblurring datasets.
Xiaobin Hu, Wenqi Ren, Kaicheng Yu, Kaihao Zhang, Xiaochun Cao, Wei Liu 0005, Bjoern Menze
ICCV3
2020 Evaluating The Search Phase of Neural Architecture Search
Kaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Cristian Musat, Mathieu Salzmann
ICLR1
2020 An Enhanced AUV- Aided TDoA Localization Algorithm for Underwater Acoustic Sensor Networks
Kun Hao, Kaicheng Yu, Zijun Gong, Xiujuan Du, Yonglei Liu
Mob. Networks Appl.2
2019 Recurrent U-Net for Resource-Constrained Segmentation
abstract
State-of-the-art segmentation methods rely on very deep networks that are not always easy to train without very large training datasets and tend to be relatively slow to run on standard GPUs. In this paper, we introduce a novel recurrent U-Net architecture that preserves the compactness of the original U-Net [33], while substantially increasing its performance to the point where it outperforms the state of the art on several benchmarks. We will demonstrate its effectiveness for several tasks, including hand segmentation, retina vessel segmentation, and road segmentation. We also introduce a large-scale dataset for hand segmentation.
Wei Wang 0108, Kaicheng Yu, Joachim Hugonot, Pascal Fua, Mathieu Salzmann
ICCV2
2019 Overcoming Multi-model Forgetting
abstract
We identify a phenomenon, which we refer to as multi-model forgetting, that occurs when sequentially training multiple deep networks with partially-shared parameters; the performance of previously-trained models degrades as one optimizes a subsequent one, due to the overwriting of shared parameters. To overcome this, we introduce a statistically-justified weight plasticity loss that regularizes the learning of a model’s shared parameters according to their importance for the previous models, and demonstrate its effectiveness when training two models sequentially and for neural architecture search. Adding weight plasticity in neural architecture search preserves the best models to the end of the search and yields improved results in both natural language processing and computer vision tasks.
Yassine Benyahia, Kaicheng Yu, Kamil Bennani-Smires, Martin Jaggi, Anthony C. Davison, Mathieu Salzmann, Claudiu Cristian Musat
ICML2
2018 Statistically-Motivated Second-Order Pooling
Kaicheng Yu, Mathieu Salzmann
ECCV (7)1