Lile Cai

dblp:140/1442 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0001-8783-0186ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 STFAR: Test-time adaptive object detection through self-training and feature alignment regularization
Nanqing Liu, Yongyi Su, Lile Cai, Heng-Chao Li 0001, Kui Jia, Tianrui Li 0001, Xun Xu 0002, Chuan-Sheng Foo
Expert Syst. Appl.5
2025 Evidential Learning-based Certainty Estimation for Robust Dense Feature Matching
abstract
Dense feature matching methods aim to estimate a dense correspondence field between images. Inaccurate correspondence can occur due to the presence of unmatchable region, necessitating the need for certainty measurement. This is typically addressed by training a binary classifier to decide whether each predicted correspondence is reliable. However, deep neural network-based classifiers can be vulnerable to image corruptions or perturbations, making it difficult to obtain reliable matching pairs in corrupted scenario. In this work, we propose an evidential deep learning framework to enhance the robustness of dense matching against corruptions. We modify the certainty prediction branch in dense matching models to generate appropriate belief masses and compute the certainty score by taking expectation over the resulting Dirichlet distribution. We evaluate our method on a wide range of benchmarks and show that our method leads to improved robustness against common corruptions and adversarial attacks, achieving up to 10.1\% improvement under severe corruptions.
Lile Cai, Chuan-Sheng Foo, Xun Xu 0002, Zaiwang Gu, Jun Cheng 0003, Xulei Yang
ICLR1
2025 Exploring Active Learning for Label-Efficient Training of Semantic Neural Radiance Field
abstract
Neural Radiance Field (NeRF) models are implicit neural scene representation methods that offer unprecedented capabilities in novel view synthesis. Semantically-aware NeRFs not only capture the shape and radiance of a scene, but also encode semantic information of the scene. The training of semantically-aware NeRFs typically requires pixel-level class labels, which can be prohibitively expensive to collect. In this work, we explore active learning as a potential solution to alleviate the annotation burden. We investigate various design choices for active learning of semantically-aware NeRF, including selection granularity and selection strategies. We further propose a novel active learning strategy that takes into account 3D geometric constraints in sample selection. Our experiments demonstrate that active learning can effectively reduce the annotation cost of training semantically-aware NeRF, achieving more than 2× reduction in annotation cost compared to random sampling.
Yuzhe Zhu, Lile Cai, Kangkang Lu 0001, Fayao Liu, Xulei Yang
ICME2
2025 Uncertainty Aware Interest Point Detection and Description
abstract
Interest point detection and description play an important role in many visual tasks, including image registration, pose estimation, 3D reconstruction, and more. State-of-the-art interest point detection techniques are based on deep neural networks (NNs), which are prone to produce overconfident predictions. However, calibrated and ro-bust uncertainty measurement is crucial when deploying deep NN models in safety critical applications. In this work, we propose a novel Uncertainty-Aware interest Point (UAPoint) detection method to address this problem. Our method leverages evidential learning to learn both aleatoric and epistemic uncertainty. We further propose a constrained sampling scheme to construct more efficient training pairs for the descriptor decoder. We evaluate our method on a wide range of benchmarks and show that our method achieves state-of-the-art performance. Code will be released in https://github.com/JingboZeng/UAPoint.
Jingbo Zeng, Zaiwang Gu, Weide Liu, Lile Cai, Jun Cheng 0003
WACV4
2024 Box-Level Class-Balanced Sampling For Active Object Detection
abstract
Training deep object detectors demands expensive bounding box annotation. Active learning (AL) is a promising technique to alleviate the annotation burden. Performing AL at box-level for object detection, i.e., selecting the most informative boxes to label and supplementing the sparsely-labelled image with pseudo labels, has been shown to be more cost-effective than selecting and labelling the entire image. In box-level AL for object detection, we observe that models at early stage can only perform well on majority classes, making the pseudo labels severely class-imbalanced. We propose a class-balanced sampling strategy to select more objects from minority classes for labelling, so as to make the final training data, i.e., ground truth labels obtained by AL and pseudo labels, more class-balanced to train a better model. We also propose a task-aware soft pseudo labelling strategy to increase the accuracy of pseudo labels. We evaluate our method on public benchmarking datasets and show that our method achieves state-of-the-art performance.
Jingyi Liao, Xun Xu 0002, Chuan-Sheng Foo, Lile Cai
ICIP4
2024 The Initialization Factor: Understanding its Impact on Active Learning for Analog Circuit Design
abstract
Active learning, which aims to enhance modeling efficiency, precision, and cost effectiveness through selective labeling, is emerging as a promising strategy for analog circuit modeling. However, analog circuits are constrained by strict functional and technological limitations, resulting in scarcity of data for modeling, and additional data acquisition involves expensive and time-consuming simulations. For efficient and effective active learning for analog circuit modeling, our research analyzes data-driven initial sampling techniques which lays the foundation for the active learning process. Our experiments reveal that these initialization strategies expedite the learning process, decrease the demand for extensive simulations, and produces more accurate models. Furthermore, the results demonstrate that active learning techniques, which uniformly sample the design space, tend to benefit from distance-based initialization technique.
Sezin Kircali Ata, Zhi-Hui Kong, Anusha James, Lile Cai, Kiat Seng Yeo, Khin Mi Mi Aung, Chuan-Sheng Foo, Ashish James
ISCAS4
2024 Revisiting pretraining for semi-supervised learning in the low-label regime
Xun Xu 0002, Jingyi Liao, Lile Cai, Kangkang Lu 0001, Wanyue Zhang, Yasin Yazici, Chuan-Sheng Foo
Neurocomputing3
2024 Exploring Diversity-Based Active Learning for 3D Object Detection in Autonomous Driving
abstract
3D object detection has recently received much attention due to its great potential in autonomous vehicle (AV). The success of deep learning based object detectors relies on the availability of large-scale annotated datasets, which is time-consuming and expensive to compile, especially for 3D bounding box annotation. In this work, we investigate diversity-based active learning (AL) as a potential solution to alleviate the annotation burden. Given limited annotation budget, only the most informative frames and objects are automatically selected for human to annotate. Technically, we take the advantage of the multimodal information provided in an AV dataset, and propose a novel acquisition function that enforces spatial and temporal diversity in the selected samples. We benchmark the proposed method against other AL strategies under realistic annotation cost measurements, where the realistic costs for annotating a frame and a 3D bounding box are both taken into consideration. We demonstrate the effectiveness of the proposed method on the nuScenes dataset and show that it outperforms existing AL strategies significantly.
Jinpeng Lin, Zhihao Liang 0002, Shengheng Deng, Lile Cai, Tao Jiang 0014, Tianrui Li 0001, Kui Jia, Xun Xu 0002
IEEE Trans. Intell. Transp. Syst.4
2022 Exploring Active Learning for Semiconductor Defect Segmentation
abstract
The development of X-Ray microscopy (XRM) technology has enabled non-destructive inspection of semiconductor structures for defect identification. Deep learning is widely used as the state-of-the-art approach to perform visual analysis tasks. However, deep learning based models require large amount of annotated data to train. This can be time-consuming and expensive to obtain especially for dense prediction tasks like semantic segmentation. In this work, we explore active learning (AL) as a potential solution to alleviate the annotation burden. We identify two unique challenges when applying AL on semiconductor XRM scans: large domain shift and severe class-imbalance. To address these challenges, we propose to perform contrastive pretraining on the unlabelled data to obtain the initialization weights for each AL cycle, and a rareness-aware acquisition function that favors the selection of samples containing rare classes. We evaluate our method on a semiconductor dataset that is compiled from XRM scans of high bandwidth memory structures composed of logic and memory dies, and demonstrate that our method achieves state-of-the-art performance.
Lile Cai, Ramanpreet Singh Pahwa, Xun Xu 0002, Jie Wang 0042, Richard Chang 0002, Lining Zhang, Chuan-Sheng Foo
ICIP1
2021 Revisiting Superpixels for Active Learning in Semantic Segmentation With Realistic Annotation Costs
abstract
State-of-the-art methods for semantic segmentation are based on deep neural networks that are known to be data-hungry. Region-based active learning has shown to be a promising method for reducing data annotation costs. A key design choice for region-based AL is whether to use regularly-shaped regions (e.g., rectangles) or irregularly-shaped region (e.g., superpixels). In this work, we address this question under realistic, click-based measurement of annotation costs. In particular, we revisit the use of super-pixels and demonstrate that the inappropriate choice of cost measure (e.g., the percentage of labeled pixels), may cause the effectiveness of the superpixel-based approach to be under-estimated. We benchmark the superpixel-based approach against the traditional "rectangle+polygon"-based approach with annotation cost measured in clicks, and show that the former outperforms on both Cityscapes and PASCAL VOC. We further propose a class-balanced acquisition function to boost the performance of the superpixel-based approach and demonstrate its effectiveness on the evaluation datasets. Our results strongly argue for the use of superpixel-based AL for semantic segmentation and highlight the importance of using realistic annotation costs in evaluating such methods.
Lile Cai, Xun Xu 0002, Jun Hao Liew, Chuan-Sheng Foo
CVPR1
2021 Exploring Spatial Diversity for Region-Based Active Learning
abstract
State-of-the-art methods for semantic segmentation are based on deep neural networks trained on large-scale labeled datasets. Acquiring such datasets would incur large annotation costs, especially for dense pixel-level prediction tasks like semantic segmentation. We consider region-based active learning as a strategy to reduce annotation costs while maintaining high performance. In this setting, batches of informative image regions instead of entire images are selected for labeling. Importantly, we propose that enforcing local spatial diversity is beneficial for active learning in this case, and to incorporate spatial diversity along with the traditional active selection criterion, e.g., data sample uncertainty, in a unified optimization framework for region-based active learning. We apply this framework to the Cityscapes and PASCAL VOC datasets and demonstrate that the inclusion of spatial diversity effectively improves the performance of uncertainty-based and feature diversity-based active learning methods. Our framework achieves 95% performance of fully supervised methods with only 5 - 9% of the labeled pixels, outperforming all state-of-the-art region-based active learning methods for semantic segmentation.
Lile Cai, Xun Xu 0002, Lining Zhang, Chuan-Sheng Foo
IEEE Trans. Image Process.1
2019 MaxpoolNMS: Getting Rid of NMS Bottlenecks in Two-Stage Object Detectors
abstract
Modern convolutional object detectors have improved the detection accuracy significantly, which in turn inspired the development of dedicated hardware accelerators to achieve real-time performance by exploiting inherent parallelism in the algorithm. Non-maximum suppression (NMS) is an indispensable operation in object detection. In stark contrast to most operations, the commonly-adopted GreedyNMS algorithm does not foster parallelism, which can be a major performance bottleneck. In this paper, we introduce MaxpoolNMS, a parallelizable alternative to the NMS algorithm, which is based on max-pooling classification score maps. By employing a novel multi-scale multi-channel max-pooling strategy, our method is 20x faster than GreedyNMS while simultaneously achieves comparable accuracy, when quantified across various benchmarking datasets, i.e., MS COCO, KITTI and PASCAL VOC. Furthermore, our method is better suited for hardware-based acceleration than GreedyNMS.
Lile Cai, Zhe Wang 0019, Jie Lin 0001, Chuan-Sheng Foo, Mohamed M. Sabry, Vijay Chandrasekhar 0001
CVPR1
2019 TEA-DNN: the Quest for Time-Energy-Accuracy Co-optimized Deep Neural Networks
abstract
Embedded deep learning platforms have witnessed two simultaneous improvements. First, the accuracy of convolutional neural networks (CNNs) has been significantly improved through the use of automated neural-architecture search (NAS) algorithms to determine CNN structure. Second, there has been increasing interest in developing hardware accelerators for CNNs that provide improved inference performance and energy consumption compared to GPUs. Such embedded deep learning platforms differ in the amount of compute resources and memory-access bandwidth, which would affect performance and energy consumption of CNNs. It is therefore critical to consider the available hardware resources in the network architecture search. To this end, we introduce TEA-DNN, a NAS algorithm targeting multi-objective optimization of execution time, energy consumption, and classification accuracy of CNN workloads on embedded architectures. TEA-DNN leverages energy and execution time measurements on embedded hardware when exploring the Pareto-optimal curves across accuracy, execution time, and energy consumption and does not require additional effort to model the underlying hardware. We apply TEA-DNN for image classification on actual embedded platforms (NVIDIA Jetson TX2 and Intel Movidius Neural Compute Stick). We highlight the Pareto-optimal operating points that emphasize the necessity to explicitly consider hardware characteristics in the search process. To the best of our knowledge, this is the most comprehensive study of Pareto-optimal models across a range of hardware platforms using actual measurements on hardware to obtain objective values.
Lile Cai, Anne-Maelle Barneche, Arthur Herbout, Chuan-Sheng Foo, Jie Lin 0001, Vijay Chandrasekhar 0001, Mohamed M. Sabry
ISLPED1
2017 Anomaly detection in thermal images using deep neural networks
abstract
Infrared thermography has become an effective tool in electrical preventive maintenance program due to its high precision and the capability of performing non-contact diagnostic. Anomalies in a thermal image is typically detected by comparing the temperatures of the equipment with reference temperatures. Manual detection is time-consuming and unreliable, making it unable to meet the excessive demand for condition monitoring in industrial applications. In this paper, we propose an automatic method to detect thermal anomalies based on deep neural networks (DNNs). The DNN model is trained to learn the statistical regularities of normal thermal images, and anomalies are detected based on pixel-wise comparison between the learned reference temperatures and the actual temperatures. We test our method on a variety of electrical equipment and the experimental results demonstrated the effectiveness of the proposed method.
Lile Cai
ICIP1
2017 A two-level clustering approach for multidimensional transfer function specification in volume visualization
Lile Cai, Binh P. Nguyen, Chee-Kong Chui, Sim Heng Ong
Vis. Comput.1
2015 Rule-Enhanced Transfer Function Generation for Medical Volume Visualization
abstract
Abstract In volume visualization, transfer functions are used to classify the volumetric data and assign optical properties to the voxels. In general, transfer functions are generated in a transfer function space, which is the feature space constructed by data values and properties derived from the data. If volumetric objects have the same or overlapping data values, it would be difficult to separate them in the transfer function space. In this paper, we present a rule‐enhanced transfer function design method that allows important structures of the volume to be more effectively separated and highlighted. We define a set of rules based on the local frequency distribution of volume attributes. A rule‐selection method based on a genetic algorithm is proposed to learn the set of rules that can distinguish the user‐specified target tissue from other tissues. In the rendering stage, voxels satisfying these rules are rendered with higher opacities in order to highlight the target tissue. The proposed method was tested on various volumetric datasets to enhance the visualization of important structures that are difficult to be visualized by traditional transfer function design methods. The results demonstrate the effectiveness of the proposed method.
Lile Cai, Binh P. Nguyen, Chee-Kong Chui, Sim Heng Ong
Comput. Graph. Forum1