Di Feng

dblp:180/4314 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 10 since 2021Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 3 since 2021Computer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2026 PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
abstract
Yuzhen Shi, Huanghai Liu, Yiran HU, Song Gaojie, Xu Xinran, Yubo Ma, Tianyi Tang, Li Zhang, Qingjing Chen, Feng Di, Wenbo Lv, Weiheng Wu, Kexin Yang, Sen Yang, Wei Wang, Rongyao Shi, Qiu Yuanyang, Yuemeng Qi, Zhang Jingwen, Sui Xiaoyu, Yifan Chen, Zhang Yi, An Yang, Bowen Yu, Dayiheng Liu, Junyang Lin, Weixing Shen, Bing Zhao, Charles L. A. Clarke, HU Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song, Xinran Xu, Yubo Ma, Qingjing Chen, Di Feng, Wenbo Lv, Weiheng Wu, Kexin Yang 0002, Wei Wang 0225, Rongyao Shi, Yuanyang Qiu, Yuemeng Qi, Xiaoyu Sui, Yi Zhang 0101, An Yang, Bowen Yu 0002, Dayiheng Liu, Junyang Lin, Weixing Shen, Charles L. A. Clarke, Hu Wei
ACL (1)10
2026 Implicitly inspired prediction approach for design thinking with multi-domain analogical knowledge driven by electroencephalogram data
Liting Jing, Jianglong Du, Yubo Dou, Chulin Tian, Di Feng, Shaofei Jiang
Adv. Eng. Informatics5
2025 UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital Agents
Harsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan, Ari Seff, Di Feng, Ruijia Cheng, Andres Romero Mier Y. Teran, Esteban Gomez, Abhishek Sundararajan, Forrest Huang, Amanda Swearngin, Mohana Prasad Sathya Moorthy, Jeffrey Nichols 0001, Alexander Toshev
ICCV6
2025 Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
abstract
Building a generalist model for user interface (UI) understanding is challenging due to various foundational issues, such as platform diversity, resolution variation, and data limitation. In this paper, we introduce Ferret-UI 2, a multimodal large language model (MLLM) designed for universal UI understanding across a wide range of platforms, including iPhone, Android, iPad, Webpage, and AppleTV. Building on the foundation of Ferret-UI, Ferret-UI 2 introduces three key innovations: support for multiple platform types, high-resolution perception through adaptive scaling, and advanced task training data generation powered by GPT-4o with set-of-mark visual prompting. These advancements enable Ferret-UI 2 to perform complex, user-centered interactions, making it highly versatile and adaptable for the expanding diversity of platform ecosystems. Extensive empirical experiments on referring, grounding, user-centric advanced tasks (comprising 9 subtasks $\times$ 5 platforms), GUIDE next-action prediction dataset, and GUI-World multi-platform benchmark demonstrate that Ferret-UI 2 significantly outperforms Ferret-UI, and also shows strong cross-platform transfer capabilities.
Zhangheng Li, Keen You, Haotian Zhang 0005, Di Feng, Harsh Agrawal, Xiujun Li, Mohana Prasad Sathya Moorthy, Jeffrey Nichols 0001, Yinfei Yang, Zhe Gan
ICLR4
2024 LCA-on-the-Line: Benchmarking Out of Distribution Generalization with Class Taxonomies
abstract
We tackle the challenge of predicting models' Out-of-Distribution (OOD) performance using in-distribution (ID) measurements without requiring OOD data. Existing evaluations with ``Effective robustness'', which use ID accuracy as an indicator of OOD accuracy, encounter limitations when models are trained with diverse supervision and distributions, such as class labels (*Vision Models, VMs, on ImageNet*) and textual descriptions (*Visual-Language Models, VLMs, on LAION*). VLMs often generalize better to OOD data than VMs despite having similar or lower ID performance. To improve the prediction of models' OOD performance from ID measurements, we introduce the *Lowest Common Ancestor (LCA)-on-the-Line* framework. This approach revisits the established concept of LCA distance, which measures the hierarchical distance between labels and predictions within a predefined class hierarchy, such as WordNet. We assess 75 models using ImageNet as the ID dataset and five significantly shifted OOD variants, uncovering a strong linear correlation between ID LCA distance and OOD top-1 accuracy. Our method provides a compelling alternative for understanding why VLMs tend to generalize better. Additionally, we propose a technique to construct a taxonomic hierarchy on any dataset using $K$-means clustering, demonstrating that LCA distance is robust to the constructed taxonomic hierarchy. Moreover, we demonstrate that aligning model predictions with class taxonomies, through soft labels or prompt engineering, can enhance model generalization. Open source code in our [Project Page](https://elvishelvis.github.io/papers/lca/).
Gautam Rajendrakumar Gare, Jinjin Tian, Siqi Chai, Zhiqiu Lin, Arun Balajee Vasudevan, Di Feng, Francesco Ferroni, Shu Kong
ICML7
2024 Mitigating Causal Confusion in Vector-Based Behavior Cloning for Safer Autonomous Planning
abstract
The utilization of vector-based deep learning techniques has great prospects in the realm of autonomous driving, particularly in the domains of prediction and planning tasks. However, the application of vector-based backbones for prediction and planning tasks may lead to the occurrence of causal confusion. Previous studies have explored the phenomenon of causal confusion, with a specific emphasis on the context of visual imitation learning. As for the vector-based model, we observe that the states of surrounding vehicles can be a nuisance shortcut. In our work, an off-policy approach is proposed to alleviate the issue by incorporating de-confounding supervision. Additionally, to better capture the environmental cues, such as route and traffic lights, in vectorized representation, a decoder utilizing iterative route fusion is devised. By incorporating auxiliary supervision and employing a dedicated decoder, we demonstrate the effectiveness of our methods in reducing causal confusion and improving performance in planning tasks through reactive and nonreactive closed-loop simulations on the nuPlan dataset.
Jiayu Guo 0001, Mingyue Feng, Jinsheng Dou, Di Feng, Chengjun Li, Ru Wan, Jian Pu
ICRA5
2024 FastOcc: Accelerating 3D Occupancy Prediction by Fusing the 2D Bird's-Eye View and Perspective View
abstract
In autonomous driving, 3D occupancy prediction outputs voxel-wise status and semantic labels for more comprehensive understandings of 3D scenes compared with traditional perception tasks, such as 3D object detection and bird’s-eye view (BEV) semantic segmentation. Recent researchers have extensively explored various aspects of this task, including view transformation techniques, ground-truth label generation, and elaborate network design, aiming to achieve superior performance. However, the inference speed, crucial for running on an autonomous vehicle, is neglected. To this end, a new method, dubbed FastOcc, is proposed. By carefully analyzing the network effect and latency from four parts, including the input image resolution, image backbone, view transformation, and occupancy prediction head, it is found that the occupancy prediction head holds considerable potential for accelerating the model while keeping its accuracy. Targeted at improving this component, the time-consuming 3D convolution network is replaced with a novel residual-like architecture, where features are mainly digested by a lightweight 2D BEV convolution network and compensated by integrating the 3D voxel features interpolated from the original image features. Experiments on the Occ3D-nuScenes benchmark demonstrate that our FastOcc achieves state-of-the-art results with a fast inference speed.
Wenhao Guan, Di Feng, Yuheng Du, Xiangyang Xue 0001, Jian Pu
ICRA5
2024 HP3: Hierarchical Prediction-Pretrained Planning for Unprotected Left Turn
Zhihao Ou, Yue Hua, Jinsheng Dou, Di Feng, Jian Pu
IROS5
2024 Conceptual design decision-making considering multigranularity heterogeneous evaluation semantics with uncertain beliefs
Liting Jing, Yubo Dou, Di Feng, Weiqiang Jia, Shaofei Jiang
Expert Syst. Appl.4
2023 Data-driven implicit design preference prediction model for product concept evaluation via BP neural network and EEG
Liting Jing, Chulin Tian, Shun He, Di Feng, Shaofei Jiang, Chunfu Lu
Adv. Eng. Informatics4
2023 Impatient Queuing for Intelligent Task Offloading in Multiaccess Edge Computing
abstract
Multi-access edge computing (MEC) emerges as an essential part of the upcoming Fifth Generation (5G) and future beyond-5G mobile communication systems. It adds computational power towards the edge of cellular networks, much closer to energy-constrained user devices, and therewith allows the users to offload tasks to the edge computing nodes for low-latency applications with very-limited battery consumption. However, due to the high dynamics of user demand and server load, task congestion may occur at the edge nodes resulting in long queuing delay. Such delays can significantly degrade the quality of experience (QoE) of some latency-sensitive applications, raise the risk of service outage, and cannot be efficiently resolved by conventional queue management solutions. In this article, we study a latency-outage critical scenario, where users intend to limit the risk of latency outage. We propose an impatience-based queuing strategy for such users to intelligently choose between MEC offloading and local computation, allowing them to rationally renege from the task queue. The proposed approach is demonstrated by numerical simulations to be efficient for generic service model, when a perfect queue status information is available. For the practical case where the users obtain only imperfect queue status information, we design an optimal online learning strategy to enable its application in Poisson service scenarios.
Bin Han 0004, Vincenzo Sciancalepore, Yihua Xu, Di Feng, Hans D. Schotten
IEEE Trans. Wirel. Commun.4
2022 DeepFusion: A Robust and Modular 3D Object Detector for Lidars, Cameras and Radars
abstract
We propose DeepFusion, a modular multi-modal architecture to fuse lidars, cameras and radars in different combinations for 3D object detection. Specialized feature extractors take advantage of each modality and can be exchanged easily, making the approach simple and flexible. Extracted features are transformed into bird's-eye-view as a common representation for fusion. Spatial and semantic alignment is performed prior to fusing modalities in the feature space. Finally, a detection head exploits rich multi-modal features for improved 3D detection performance. Experimental results for lidar-camera, lidar-camera-radar and camera-radar fusion show the flexibility and effectiveness of our fusion approach. In the process, we study the largely unexplored task of faraway car detection up to 225 meters, showing the benefits of our lidar-camera fusion. Furthermore, we investigate the required density of lidar points for 3D object detection and illustrate implications at the example of robustness against adverse weather conditions. Moreover, ablation studies on our camera-radar fusion highlight the importance of accurate depth estimation.
Florian Drews, Di Feng, Florian Faion, Lars Rosenbaum, Michael Ulrich, Claudius Gläser
IROS2
2022 A Review and Comparative Study on Probabilistic Object Detection in Autonomous Driving
abstract
Capturing uncertainty in object detection is indispensable for safe autonomous driving. In recent years, deep learning has become the de-facto approach for object detection, and many probabilistic object detectors have been proposed. However, there is no summary on uncertainty estimation in deep object detection, and existing methods are either built with different network architectures and uncertainty estimation methods, or evaluated on different datasets with a wide range of evaluation metrics. As a result, a comparison among methods remains challenging, as does the selection of a model that best suits a particular application. This paper aims to alleviate this problem by providing a review and comparative study on existing probabilistic object detection methods for autonomous driving applications. First, we provide an overview of practical uncertainty estimation methods in deep learning, and then systematically survey existing methods and evaluation metrics for probabilistic object detection. Next, we present a strict comparative study for probabilistic object detection based on an image detector and three public autonomous driving datasets. Finally, we present a discussion of the remaining challenges and future works. Code has been made available athttps://github.com/asharakeh/pod_compare.git.
Di Feng, Ali Harakeh, Steven Lake Waslander, Klaus Dietmayer
IEEE Trans. Intell. Transp. Syst.1
2022 Labels are Not Perfect: Inferring Spatial Uncertainty in Object Detection
abstract
The availability of many real-world driving datasets is a key reason behind the recent progress of object detection algorithms in autonomous driving. However, there exist ambiguity or even failures in object labels due to error-prone annotation process or sensor observation noise. Current public object detection datasets only provide deterministic object labels without considering their inherent uncertainty, as does the common training process or evaluation metrics for object detectors. As a result, an in-depth evaluation among different object detection methods remains challenging, and the training process of object detectors is sub-optimal, especially in probabilistic object detection. In this work, we infer the uncertainty in bounding box labels from LiDAR point clouds based on a generative model, and define a new representation of the probabilistic bounding box through a spatial uncertainty distribution. Comprehensive experiments show that the proposed model reflects complex environmental noises in LiDAR perception and the label quality. Furthermore, we propose Jaccard IoU (JIoU) as a new evaluation metric that extends IoU by incorporating label uncertainty. We conduct an in-depth comparison among several LiDAR-based object detectors using the JIoU metric. Finally, we incorporate the proposed label uncertainty in a loss function to train a probabilistic object detector and to improve its detection accuracy. We verify our proposed methods on two public datasets (KITTI, Waymo), as well as on simulation data. Code is released athttps://github.com/ZiningWang/Inferring-Spatial-Uncertainty-in-Object-Detection.
Di Feng, Yiyang Zhou, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer, Masayoshi Tomizuka
IEEE Trans. Intell. Transp. Syst.1
2021 A Simple and Efficient Multi-task Network for 3D Object Detection and Road Understanding
abstract
Detecting dynamic objects and predicting static road information such as drivable areas and ground heights are crucial for safe autonomous driving. Previous works studied each perception task separately, and lacked a collective quantitative analysis. In this work, we show that it is possible to perform all perception tasks via a simple and efficient multi-task network. Our proposed network, LidarMTL, takes raw LiDAR point cloud as inputs, and predicts six perception outputs for 3D object detection and road understanding. The network is based on an encoder-decoder architecture with 3D sparse convolution and deconvolution operations. Extensive experiments verify the proposed method with competitive accuracies compared to state-of-the-art object detectors and other task-specific networks. LidarMTL is also leveraged for online localization. Code and pre-trained model have been made available at https://github.com/frankfengdi/LidarMTL.
Di Feng, Yiyang Zhou, Chenfeng Xu, Masayoshi Tomizuka
IROS1
2021 Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
abstract
Recent advancements in perception for autonomous driving are driven by deep learning. In order to achieve robust and accurate scene understanding, autonomous vehicles are usually equipped with different sensors (e.g. cameras, LiDARs, Radars), and multiple sensing modalities can be fused to exploit their complementary properties. In this context, many methods have been proposed for deep multi-modal perception problems. However, there is no general guideline for network architecture design, and questions of “what to fuse”, “when to fuse”, and “how to fuse” remain open. This review paper attempts to systematically summarize methodologies and discuss challenges for deep multi-modal object detection and semantic segmentation in autonomous driving. To this end, we first provide an overview of on-board sensors on test vehicles, open datasets, and background information for object detection and semantic segmentation in autonomous driving research. We then summarize the fusion methodologies and discuss challenges and open questions. In the appendix, we provide tables that summarize topics and methods. We also provide an interactive online platform to navigate each reference: https://boschresearch.github.io/multimodalperception/.
Di Feng, Christian Haase-Schütz, Lars Rosenbaum, Heinz Hertlein, Claudius Gläser, Fabian Timm, Werner Wiesbeck, Klaus Dietmayer
IEEE Trans. Intell. Transp. Syst.1
2020 Inferring Spatial Uncertainty in Object Detection
abstract
The availability of real-world datasets is the prerequisite for developing object detection methods for autonomous driving. While ambiguity exists in object labels due to error-prone annotation process or sensor observation noises, current object detection datasets only provide deterministic annotations without considering their uncertainty. This precludes an in-depth evaluation among different object detection methods, especially for those that explicitly model predictive probability. In this work, we propose a generative model to estimate bounding box label uncertainties from LiDAR point clouds, and define a new representation of the probabilistic bounding box through spatial distribution. Comprehensive experiments show that the proposed model represents uncertainties commonly seen in driving scenarios. Based on the spatial distribution, we further propose an extension of IoU, called the Jaccard IoU (JIoU), as a new evaluation metric that incorporates label uncertainty. Experiments on the KITTI and the Waymo Open Datasets show that JIoU is superior to IoU when evaluating probabilistic object detectors.
Di Feng, Yiyang Zhou, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer, Masayoshi Tomizuka
IROS2
2020 Leveraging Uncertainties for Deep Multi-modal Object Detection in Autonomous Driving
abstract
This work presents a probabilistic deep neural network that combines LiDAR point clouds and RGB camera images for robust, accurate 3D object detection. We explicitly model uncertainties in the classification and regression tasks, and leverage uncertainties to train the fusion network via a sampling mechanism. We validate our method on three datasets with challenging real-world driving scenarios. Experimental results show that the predicted uncertainties reflect complex environmental uncertainty like difficulties of a human expert to label objects. The results also show that our method consistently improves the Average Precision by up to 7% compared to the baseline method. When sensors are temporally misaligned, the sampling method improves the Average Precision by up to 20%, showing its high robustness against noisy sensor inputs.
Di Feng, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer
IV1
2020 Multiservice-Based Network Slicing Orchestration With Impatient Tenants
abstract
The combination of recent emerging technologies such as network function virtualization (NFV) and network programmability (SDN) gave birth to the novel Network Slicing paradigm. 5G networks consist of multi-tenant infrastructures capable of offering leased network “slices” to new customers (e.g., vertical industries) enabling a new telecom business model: Slice-as-a-Service (SlaaS). However, as the service demand gets increasingly dense, slice requests congestion may occur leading to undesired waiting periods. This may turn into impatient tenant behaviors that increase potential loss of the business attractiveness to customers. In this paper, we aim to: 1) study the slicing admission control problem by means of a multi-queuing system for heterogeneous tenant requests; 2) derive its statistical behavior model; 3) find out the rational strategy of impatient tenants waiting in queue-based slice admission control systems; 4) prove mathematically and empirically the benefits of allowing infrastructure providers to share its information with the upcoming tenants; and 5) provide a utility model for network slices admission optimization. Our results analyze the capability of the proposed SlaaS system to be approximately Markovian and evaluate its performance as compared to a baseline solution.
Bin Han 0004, Vincenzo Sciancalepore, Xavier Pérez Costa, Di Feng, Hans D. Schotten
IEEE Trans. Wirel. Commun.4
2020 Study on the adaptability of augmented reality smartglasses for astigmatism based on holographic waveguide grating
abstract
Augmented reality (AR) smartglasses are considered as the next generation of smart devices to replace mobile phones, and are widely concerned. But at present, AR smartglasses are usually designed according to the human normal eyes. In order to experience AR smartglasses perfectly, abnormal eye users must first wear diopters. For people with astigmatism to use AR smartglasses without wearing a diopter lens, a cylindrical lens waveguide grating is designed in this study based on the principle of holographic waveguide grating. First, a cylindrical lens waveguide substrate is constructed for external light deflection to satisfy the users' normal viewing of the real world. Further, a variable period grating structure is established based on the cylindrical lens waveguide substrate to normally emit the light from the virtual world in the optical machine to the human eyes. Finally, the structural parameters of grating are optimized to improve the diffraction efficiency. The results show that the structure of cylindrical lens waveguide grating allows people with astigmatism to wear AR smartglasses directly. The total light utilization rate reaches 90% with excellent imaging uniformity. The brightness difference is less than 0.92% and the vertical field of view is 10°. This research serves as a guide for AR product designs for people with long/short sightedness and promotes the development of such products.
Shenze Wang, Kaikai Du, Ningfang Song, Dongfeng Zhao, Di Feng, Zhengqian Tu
Virtual Real. Intell. Hardw.5
2019 A Utility-Driven Multi-Queue Admission Control Solution for Network Slicing
abstract
The combination of recent emerging technologies such as network function virtualization (NFV) and network programmability (SDN) gave birth to the Network Slicing revolution. 5G networks consist of multi-tenant infrastructures capable of offering leased network “slices” to new customers (e.g., vertical industries) enabling a new telecom business model: Slice-as-a-Service (SlaaS). In this paper, we aim i) to study the slicing admission control problem by means of a multi-queuing system for heterogeneous tenant requests, ii) to derive its statistical behavior model, and iii) to provide a utility-based admission control optimization. Our results analyze the capability of the proposed SlaaS system to be approximately Markovian and evaluate its performance as compared to legacy solutions.
Bin Han 0004, Vincenzo Sciancalepore, Di Feng, Xavier Pérez Costa, Hans D. Schotten
INFOCOM3
2019 Leveraging Heteroscedastic Aleatoric Uncertainties for Robust Real-Time LiDAR 3D Object Detection
abstract
We present a robust real-time LiDAR 3D object detector that leverages heteroscedastic aleatoric uncertainties to significantly improve its detection performance. A multi-loss function is designed to incorporate uncertainty estimations predicted by auxiliary output layers. Using our proposed method, the network ignores to train from noisy samples, and focuses more on informative ones. We validate our method on the KITTI object detection benchmark. Our method surpasses the baseline method which does not explicitly estimate uncertainties by up to nearly 9% in terms of Average Precision (AP). It also produces state-of-the-art results compared to other methods, while running with an inference time of only 72ms. In addition, we conduct extensive experiments to understand how aleatoric uncertainties behave. Extracting aleatoric uncertainties brings almost no additional computation cost during the deployment, making our method highly desirable for autonomous driving applications.
Di Feng, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer
IV1
2019 Deep Active Learning for Efficient Training of a LiDAR 3D Object Detector
abstract
Training a deep object detector for autonomous driving requires a huge amount of labeled data. While recording data via on-board sensors such as camera or LiDAR is relatively easy, annotating data is very tedious and time-consuming, especially when dealing with 3D LiDAR points or radar data. Active learning has the potential to minimize human annotation efforts while maximizing the object detector's performance. In this work, we propose an active learning method to train a LiDAR 3D object detector with the least amount of labeled training data necessary. The detector leverages 2D region proposals generated from the RGB images to reduce the search space of objects and speed up the learning process. Experiments show that our proposed method works under different uncertainty estimations and query functions, and can save up to 60% of the labeling efforts while reaching the same network performance.
Di Feng, Lars Rosenbaum, Atsuto Maki, Klaus Dietmayer
IV1
2016 Structure Filling and Matching for Three-Dimensional Reconstruction of Buildings From Single High-Resolution SAR Image
abstract
In this letter, a structure filling and matching method is proposed for the 3-D reconstruction of buildings, particularly those with combinational structures, from a single high-resolution synthetic aperture radar (SAR) image. This method consists of two stages. First, the structure model of the investigated building is constructed by filling the building structure elements, i.e., walls and roofs, into an initialized empty scene. Subsequently, a series of simulated SAR images is generated by this structure model and matched with the real SAR image of the building. By finding the maximal mutual information between the simulated and real images, the optimal geometric parameters of the model are obtained, which make the output more accurate. Compared with other methods, this method can reconstruct different kinds of practical buildings and retrieve their exact geometric parameters, using only a single SAR image. The performance of this method is evaluated on real TerraSAR-X images to show its effectiveness.
Di Feng, Weidong Chen 0010
IEEE Geosci. Remote. Sens. Lett.1