EDBT 2026 Demo / reviewers in the wild / expert
Fen Fang
dblp:79/9756
· DBLP profile ↗
31ranked-venue papers
16as first author
18since 2021 · last 2026
0000-0002-3834-4795ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 13 since 2021Artificial intelligence and machine learning · 15 · 7 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Illumination-Aware Restoration of Metalens-Captured Images: A New Dataset and a Strong BaselineabstractMetalenses offer compelling advantages such as lightweight and ultra-thin design, making them promising alternatives to conventional lenses. However, their widespread adoption is hindered by image quality degradation caused by chromatic and angular aberrations. To mitigate this, restoration processes are often necessary to recover high-quality RGB images from metalens-captured inputs. While recent deep learning-based restoration methods show promise, they typically (1) blur or distort peripheral regions, or (2) fail entirely under unseen illumination conditions. To advance metalens image restoration, we introduce IlluMeta---the first and largest real-world, illumination-aware metalens image dataset—captured across diverse lighting environments. In addition, we propose a novel end-to-end restoration framework that directs attention to challenging regions and adaptively adjusts to varying illuminations via reinforcement learning. Experiments show that our method can be applied in a plug-and-play manner to enhance existing models, significantly improving image restoration quality, especially under unseen lighting conditions, paving the way for broader real-world deployment of metalens technologies. Fen Fang, Xinan Liang, Muli Yang, Jinghong Zheng 0001, Tobias Wilhelm W. Mass, Ying Sun 0001, Xulei Yang, Xuewu Xu, Zhengguo Li |
AAAI | 1 |
| 2026 | Next-Generation Metalens Vision System: Powered by AI and Applied to AIabstractMetalenses have been widely recognized as a key building block of next-generation optical systems, offering unprecedented advantages in compactness, lightweight design, and scalable manufacturing compared to traditional refractive optics. Despite this promise, practical use is limited by optical aberrations, blur, and illumination sensitivity, which degrade both visual quality and machine perception. In this demonstration, we present an end-to-end metalens vision system—from hardware sensing with a custom-built RGB metalens camera, to physics-informed imaging and real-time restoration, and finally to downstream vision applications such as object detection and depth estimation. By integrating spatially-aware attention enhancement and reinforcement learning-based illumination control into a real-time system, our solution transforms degraded raw captures into high-fidelity images that are both visually interpretable and functionally reliable for machine vision. This AI-powered pipeline highlights metalenses as a cornerstone for next-generation imaging, where advances in optics and machine intelligence jointly drive the future of visual perception. Fen Fang, Muli Yang, Henan Wang, Xinan Liang, Tobias Wilhelm W. Mass, Xuewu Xu, Xulei Yang, Zhengguo Li |
AAAI | 1 |
| 2026 | Editing Is a Bargaining Game: Balanced Knowledge Editing in Large Language ModelsabstractLarge Language Models (LLMs) are prone to generating incorrect or outdated information, thereby necessitating efficient and precise mechanisms for knowledge updates. Existing knowledge editing approaches, however, often encounter conflicts between two competing objectives: maintaining existing knowledge (preservation) and incorporating new information (editing). During gradient-based optimization, these conflicting objectives can lead to imbalanced update directions, where one gradient dominates, ultimately resulting in suboptimal learning dynamics. To address this challenge, we propose a balanced knowledge editing framework inspired by Nash bargaining theory. Our method guides the optimization process toward a Pareto stationary point, ensuring an equilibrium solution wherein any deviation from the final state would degrade the overall performance with respect to both objectives. This guarantees optimality in preserving prior knowledge while integrating new information. We empirically validate the effectiveness of our approach across a range of evaluation metrics on standard benchmark datasets. Extensive experiments show that our method consistently outperforms state-of-the-art techniques, achieving a superior balance between knowledge preservation and update accuracy. Jiexi Yan, Muli Yang, Fen Fang, Cheng Deng 0002 |
AAAI | 4 |
| 2026 | Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If CalibratedabstractDespite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training. Specifically, models tend to overfit to superficial artifacts that do not generalize well across different generation methods, leading to a misaligned decision threshold when faced with test-time distribution shift. To address this, we propose a theoretically grounded post-hoc calibration framework based on Bayesian decision theory. In particular, we introduce a learnable scalar correction to the model’s logits, optimized on a small validation set from the target distribution while keeping the backbone frozen. This parametric adjustment compensates for distributional shift in model output, realigning the decision boundary even without requiring ground-truth labels. Experiments on challenging benchmarks show that our approach significantly improves robustness without retraining, offering a lightweight and principled solution for reliable and adaptive AI-generated image detection in the open world. Muli Yang, Gabriel James Goenawan, Henan Wang, Huaiyuan Qin, Yanhua Yang, Fen Fang, Ying Sun 0001, Joo-Hwee Lim, Hongyuan Zhu 0002 |
AAAI | 7 |
| 2026 | From Language to Driving: A Dual-Loop SLM-Enhanced Framework for Multi-Planner Scheduling via a Domain-Specific LanguageabstractJiawei Liu, Xun Gong, Muli Yang, Xingrui Yu, Fen Fang, Xulei Yang, Ivor Tsang, Yunfeng hu, Hong Chen, Qing Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xun Gong 0007, Muli Yang, Xingrui Yu, Fen Fang, Xulei Yang, Ivor W. Tsang, Yunfeng Hu 0003, Hong Chen 0003, Qing Guo 0005 |
ACL (1) | 5 |
| 2026 | Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective DiffusionabstractProcedure planning in instructional videos entails predicting an action sequence that transitions a given start state to a desired goal state. This task is particularly challenging due to two key sources of uncertainty: limited visual observations and an enormous decision space. The former results in multiple plausible plan variations due to missing intermediate visual states, while the latter complicates prediction by requiring selection from a large set of potential actions. Unlike prior work that addresses these issues implicitly, we propose an explicit solution. To mitigate the first challenge, we employ image generation models to synthesize diverse intermediate visual states using various text prompts, followed by a prompt selection module integrated within a diffusion model. To tackle the second challenge, we introduce a task-selective diffusion model that applies a task-specific mask to constrain the action space. As the effectiveness of this mask depends on accurate task classification, we further enhance visual representation by leveraging pre-trained vision-language models to generate action-aware, text-enriched multimodal embeddings. Extensive experiments on three benchmark datasets validate the superior performance of our proposed approach. Fen Fang, Muli Yang, Min Wu 0008, Yanhua Yang, Qianli Xu, Joo-Hwee Lim, Xulei Yang, Hongyuan Zhu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | See, Predict, Plan: Diffusion for Procedure Planning in Robotic Surgical Videos
Ziyuan Zhao, Fen Fang, Xulei Yang, Qianli Xu, Cuntai Guan, Shaohua Kevin Zhou |
MICCAI (6) | 2 |
| 2024 | Localizing discriminative regions for fine-grained visual recognition: One could be better than many
Fen Fang, Yun Liu 0011, Qianli Xu |
Neurocomputing | 1 |
| 2024 | Enhancing Representation Learning With Spatial Transformation and Early Convolution for Reinforcement Learning-Based Small Object DetectionabstractAlthough object detection has achieved significant progress in the past decade, detecting small objects is still far from satisfactory due to the high variability of object scales and complex backgrounds. The common way to enhance small object detection is to use high-resolution (HR) images. However, this method incurs huge computational resources which grow squarely with the resolution of images. To achieve both accuracy and efficiency, we propose a novel reinforcement learning framework that employs an efficient policy network consisting of a Spatial Transformation Network to enhance the state representation learning and a Transformer model with early convolution to improve feature extraction. Our method has two main steps: (1) coarse location query (CLQ), where an RL agent is trained to predict the locations of small objects on low-resolution (LR) (down-sampled version of HR) images; (2) context-sensitive object detection where HR image patches are used to detect objects on the selected coarse locations and LR image patches on background areas (containing no small objects). In this way, we can obtain high detection performance on small objects while avoiding unnecessary computation on background areas. The proposed method has been tested and benchmarked on various datasets. On the Caltech Pedestrians Detection and Web Pedestrians datasets, the proposed method improves the detection accuracy by 2%, while reducing the number of processed pixels. On the Vision meets Drone object detection dataset and the Oil and Gas Storage Tank dataset, the proposed method outperforms the state-of-the-art (SotA) methods. On MS COCO mini-val set, our method outperforms SotA methods on small object detection, while also achieving comparable performance on medium and large objects. Fen Fang, Wenyu Liang, Qianli Xu, Joo-Hwee Lim |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Data Augmentation Using Corner CutMix and an Auxiliary Self-Supervised LossabstractDeep convolutional neural networks (CNNs) have achieved remarkable success in computer vision tasks, but their training is susceptible to overfitting when the training sample size is insufficient. In this paper, we introduce Corner CutMix, a novel data augmentation technique for CNN training. During training, Corner CutMix randomly selects a region from one of four corner areas in an image and replaces it with a randomly chosen region from a distractor image. Additionally, we design an auxiliary self-supervised loss function to learn the position of the selected corner region, thereby improving the transferability and generalizability of the learned representation. Corner CutMix is easy to implement, adding little computational overhead, and can be combined with other augmentation methods such as random cropping, color distortion, and flipping. Our extensive classification task experiments in self-supervised learning on public datasets (e.g., CIFAR10, CIFAR100, and STL10) demonstrate the effectiveness of Corner CutMix, which consistently outperforms strong baselines such as CutOut and CutMix. Fen Fang, Nhat M. Hoang, Qianli Xu, Joo-Hwee Lim |
ICIP | 1 |
| 2022 | Improving Generalization of Reinforcement Learning Using a Bilinear Policy NetworkabstractIn deep reinforcement learning (DRL), the agent is usually trained on seen environments by optimizing a policy network. However, it is difficult to be generalized to unseen environments properly, even when the environmental variations are insignificant. This is partly because the policy network cannot effectively learn the representation of visual difference that is subtle among highly similar states in the environments. Because a bilinear structured model containing two feature extractors allows pairwise feature interactions in a translation-ally invariant manner which makes it particularly useful for subtle difference recognition among highly similar states, in this work, a bilinear policy network is employed to enhance representation learning, and thus to improve generalization of the DRL. The proposed bilinear policy network is tested on various DRL task, including a control task on path planning for active object detection, and Grid World, an AI game task. The test results show that the generalization of DRL can be improved by the proposed network. Fen Fang, Wenyu Liang, Yan Wu 0002, Qianli Xu, Joo-Hwee Lim |
ICIP | 1 |
| 2022 | Hierarchical Defect Detection Based On Reinforcement LearningabstractIn this paper, we propose a novel reinforcement learning (RL) based method for defects detection in high-resolution (HR) images. e.g. cracks and scratches on the surfaces of buildings, constructions, and products. Our innovation leverages RL to explore challenging images in progressive manner, using pre-trained deep learning (DL) detection as feedback mechanism. First, The DL model is pre-trained on low resolution (LR) images with relatively high defect background ratio (DBR). The RL agent is trained by optimizing a policy network according to feedback of DL model on selected regions of HR images with fairly low DBR to coarsely predict defective region by executing two actions: defective region selection and region refinement. Then, the selected defective regions are evaluated using the DL model to generate final defect region which will be mapped back to the HR images. Experimental results on HR crack and scratch images indicate that our method is able to achieve state-of-the-art performance with 0.976 and 0.965 F1-score respectively. Fen Fang, Qianli Xu, Joo-Hwee Lim |
ICIP | 1 |
| 2022 | Image Understanding With Reinforcement Learning: Auto-Tuning Image Attributes and Model Parameters for Object Detection and SegmentationabstractModels for image semantics understanding, such as deep learning (DL) models and mathematical models, are often trained on specific dataset or configured with specific parameters. When deploying such models on new tasks in a different test environment, it requires considerable effort to re-train the model or extensive expertise to tune the parameters. In this paper, we propose a smart reinforcement learning (RL) agent that could learn to tune parameters automatically to enhance model performance. The learning process is formulated as a generic control task for parameter adjustment, and applied to two use scenarios: (1) image attributes tuning to improve object detection performance on fixed DL model, and (2) parameter tuning of the mathematical model (Level Set) for image segmentation. We design a novel dynamic threshold mechanism in a multi-branch RL agent to effectively tune parameters of image qualities (for object detection) and Level Set models (for object segmentation). We conduct experiments on Pascal-VOC testing set, MS COCO validation set and a proprietary dataset of industrial components, where we achieve substantial improvement on object detection accuracy. We also perform experiments on the automatic parameter tuning of Level Set models. Results show that our method facilitates considerable performance improvement on public datasets compared with baseline method. Fen Fang, Qianli Xu, Ying Sun 0001, Joo-Hwee Lim |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | TAILOR: Teaching with Active and Incremental Learning for Object RegistrationabstractWhen deploying a robot to a new task, one often has to train it to detect novel objects, which is time-consuming and labor- intensive. We present TAILOR - a method and system for ob- ject registration with active and incremental learning. When instructed by a human teacher to register an object, TAILOR is able to automatically select viewpoints to capture informa- tive images by actively exploring viewpoints, and employs a fast incremental learning algorithm to learn new objects without potential forgetting of previously learned objects. We demonstrate the effectiveness of our method with a KUKA robot to learn novel objects used in a real-world gearbox as- sembly task through natural interactions. Qianli Xu, Nicolas Gauthier, Wenyu Liang, Fen Fang, Hui Li Tan, Ying Sun 0001, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
AAAI | 4 |
| 2021 | Enhancing Multi-Step Action Prediction for Active Object DetectionabstractActive vision for robots is one promising solution to open world visual detection problems. A fundamental issue is view planning, i.e., predicting next best views to capture images of interest to reduce uncertainty. While multi-step action in a reinforcement learning (RL) setup can boost the efficiency of view planning, existing methods suffer from unstable detection outcome when the Q-values of multiple branches of action advantages (i.e., action range and action type) are combined naively. To tackle this issue, we propose a novel mechanism to disentangle action range from action type through a two-stage training strategy on a deep Q-network. It combines well-crafted loss functions with respect to action range and action type to enforce separated training of these two branches. We evaluate our method on two public datasets and show that it facilitates substantial gain in view planning efficiency, while enhancing detection accuracy. Fen Fang, Qianli Xu, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim |
ICIP | 1 |
| 2021 | Towards Efficient Multiview Object Detection with Adaptive Action PredictionabstractActive vision is a desirable perceptual feature for robots. Existing approaches usually make strong assumptions about the task and environment, thus are less robust and efficient. This study proposes an adaptive view planning approach to boost the efficiency and robustness of active object detection. We formulate the multi-object detection task as an active multiview object detection problem given the initial location of the objects. Next, we propose a novel adaptive action prediction (A2P) method built on a deep Q-learning network with a dueling architecture. The A2P method is able to perform view planning based on visual information of multiple objects; and adjust action ranges according to the task status. Evaluated on the AVD dataset, A2P leads to 21.9% increase in detection accuracy in unfamiliar environments, while improving efficiency by 22.7%. On the T-LESS dataset, multi-object detection boosts efficiency by more than 30% while achieving equivalent detection accuracy. Qianli Xu, Fen Fang, Nicolas Gauthier, Wenyu Liang, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
ICRA | 2 |
| 2021 | Predicting Event Memorability from Contextual Visual SemanticsabstractEpisodic event memory is a key component of human cognition. Predicting event memorability,i.e., to what extent an event is recalled, is a tough challenge in memory research and has profound implications for artificial intelligence. In this study, we investigate factors that affect event memorability according to a cued recall process. Specifically, we explore whether event memorability is contingent on the event context, as well as the intrinsic visual attributes of image cues. We design a novel experiment protocol and conduct a large-scale experiment with 47 elder subjects over 3 months. Subjects’ memory of life events is tested in a cued recall process. Using advanced visual analytics methods, we build a first-of-its-kind event memorability dataset (called R3) with rich information about event context and visual semantic features. Furthermore, we propose a contextual event memory network (CEMNet) that tackles multi-modal input to predict item-wise event memorability, which outperforms competitive benchmarks. The findings inform deeper understanding of episodic event memory, and open up a new avenue for prediction of human episodic memory. Source code is available at https://github.com/ffzzy840304/Predicting-Event-Memorability. Qianli Xu, Fen Fang, Ana Garcia del Molino, Vigneshwaran Subbaraju, Joo-Hwee Lim |
NeurIPS | 2 |
| 2021 | Lifelog Image Retrieval Based on Semantic Relevance MappingabstractLifelog analytics is an emerging research area with technologies embracing the latest advances in machine learning, wearable computing, and data analytics. However, state-of-the-art technologies are still inadequate to distill voluminous multimodal lifelog data into high quality insights. In this article, we propose a novel semantic relevance mapping ( SRM ) method to tackle the problem of lifelog information access. We formulate lifelog image retrieval as a series of mapping processes where a semantic gap exists for relating basic semantic attributes with high-level query topics. The SRM serves both as a formalism to construct a trainable model to bridge the semantic gap and an algorithm to implement the training process on real-world lifelog data. Based on the SRM, we propose a computational framework of lifelog analytics to support various applications of lifelog information access, such as image retrieval, summarization, and insight visualization. Systematic evaluations are performed on three challenging benchmarking tasks to show the effectiveness of our method. Qianli Xu, Ana Garcia del Molino, Jie Lin 0001, Fen Fang, Vigneshwaran Subbaraju, Liyuan Li, Joo-Hwee Lim |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Task-Oriented Multi-Modal Question Answering For Collaborative ApplicationsabstractCobots that can work in human workspaces and adapt to human need to understand and respond to human’s inquiry and instruction. In this paper, we propose new question answering (QA) task and dataset for human-robot collaboration on task-oriented operation, i.e., task-oriented collaborative QA (TCQA). Differing from conventional video QA for answering questions about what happened in video clips constrained by scripts and subtitles, TC-QA aims to share common ground for task-oriented operation through question answering. We propose an open-end (OE) format of answer with text reply, image with annotated related objects, and video with operation duration to guide operation execution. Designed for grounding, the TC-QA dataset comprises query videos and questions to seek acknowledgement, correction, attention to task-related objects, and information on objects or operation. Due to the flexibility of real-world task with limited training sample, we propose and evaluate a baseline method based on a hybrid approach. The hybrid approach employs deep learning methods for object detection, hand detection and gesture recognition, and symbolic reasoning to ground question on observation for providing the answer. Our experiments show that the hybrid method is effective for the TC-QA task. Hui Li Tan, Mei Chee Leong, Qianli Xu, Liyuan Li, Fen Fang, Nicolas Gauthier, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 5 |
| 2020 | Active Image Sampling on Canonical Views for Novel Object DetectionabstractTo alleviate the costly data annotation problem in deep learning-based object detection, we leverage the canonical view model for active sample selection to improve the effectiveness of learning. Inspired by the view-approximation model, we hypothesize that visual features learned from canonical views denote better representations of objects, thus boosting the effectiveness of object learning. We validate the hypothesis empirically in the context of robot learning for novel object detection. Based on this, we propose a novel on-line viewpoint exploration (OLIVE) method that (1) defines goodness-of-view by combining informativeness of visual features and consistency of model-based object detection, and (2) systematically explores and selects viewpoints to boost learning efficiency. Furthermore, we train a legacy Faster R-CNN model with a data augmentation method while leveraging data samples generated by the OLIVE pipeline. We test our method on the T-LESS dataset and show that the proposed method outperforms competitive benchmarking methods, especially when the samples are few. Qianli Xu, Fen Fang, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim |
ICIP | 2 |
| 2020 | Detecting Objects with High Object Region PercentageabstractObject shape is a subtle but important factor for object detection. It has been observed that the object-region-percentage (ORP) can be utilized to improve detection accuracy for elongated objects, which have much lower ORPs than other types of objects. In this paper, we propose an approach to improve the detection performance for objects with high ORPs. Our method consists of three steps. First, we adjust the ground truth bounding boxes of high-ORP objects to an optimal range. Second, we train an object detector, Faster R-CNN, based on adjusted bounding boxes to achieve high recall. Finally, we train a DCNN to learn the adjustment ratios towards four directions and adjust detected bounding boxes of objects to get better localization for higher precision. We evaluate the effectiveness of our method on 12 high-ORP objects in COCO and 8 objects in a proprietary gearbox dataset. The experimental results show that our method can achieve state-of-the-art performance on these objects while costing less resources in training and inference stages. Fen Fang, Qianli Xu, Liyuan Li, Joo-Hwee Lim |
ICPR | 1 |
| 2020 | A novel hybrid approach for crack detection
Fen Fang, Liyuan Li, Hongyuan Zhu 0002, Joo-Hwee Lim |
Pattern Recognit. | 1 |
| 2020 | Combining Faster R-CNN and Model-Driven Clustering for Elongated Object DetectionabstractWhile analyzing the performance of state-of-the-art R-CNN based generic object detectors, we find that the detection performance for objects with low object-region-percentages (ORPs) of the bounding boxes are much lower than the overall average. Elongated objects are examples. To address the problem of low ORPs for elongated object detection, we propose a hybrid approach which employs a Faster R-CNN to achieve robust detections of object parts, and a novel model-driven clustering algorithm to group the related partial detections and suppress false detections. First, we train a Faster R-CNN with partial region proposals of suitable and stable ORPs. Next, we introduce a deep CNN (DCNN) for orientation classification on the partial detections. Then, on the outputs of the Faster R-CNN and DCNN, the algorithm of adaptive model-driven clustering first initializes a model of an elongated object with a data-driven process on local partial detections, and refines the model iteratively by model-driven clustering and data-driven model updating. By exploiting Faster R-CNN to produce robust partial detections and model-driven clustering to form a global representation, our method is able to generate a tight oriented bounding box for elongated object detection. We evaluate the effectiveness of our approach on two typical elongated objects in the COCO dataset, and other typical elongated objects, including rigid objects (pens, screwdrivers and wrenches) and non-rigid objects (cracks). Experimental results show that, compared with the state-of-the-art approaches, our method achieves a large margin of improvements for both detection and localization of elongated objects in images. Fen Fang, Liyuan Li, Hongyuan Zhu 0002, Joo-Hwee Lim |
IEEE Trans. Image Process. | 1 |
| 2019 | Towards Real-Time Crack Detection Using a Deep Neural Network With a Bayesian Fusion AlgorithmabstractSurface cracks can represent very small and thin objects in images. With irregular shapes and sizes, and non-fixed texture patterns, the detection of cracks can be a challenging problem in computer vision. Prior work has been undertaken on detecting cracks for images using a sliding window mode. However, such methods can be time consuming, and result in high false alarms. To help address this problem, a new crack detection and segmentation method is proposed in this paper. Specifically, our method includes three main features: (1) a Faster R-CNN model to detect crack patches in images; (2) the use of a Bayesian fusion algorithm to suppress false alarms based on detected patch orientation; and (3) image processing functions to obtain final segmentation masks, such as for Gaussian blur, erosion, etc. Experimental results show that our method can achieve high detection accuracy on sampled images in real-time. Fen Fang, Liyuan Li, Mark D. Rice, Joo-Hwee Lim |
ICIP | 1 |
| 2019 | An Adaptive Fitting Approach for the Visual Detection and Counting of Small Circular Objects in Manufacturing ApplicationsabstractDetecting, localizing and counting small circular objects in machine parts is an important task in many applications for manufacturing. Existing methods of circle detection face difficulties due to the high-curvature and limited edge points of circles. As a result, in this paper we propose a novel two-stage circle detection method, which integrates bottom-up coarse detection and top-down circle fitting. First, a circle detector combining low-level feature descriptors and a linear SVM is developed. This is used to scan an input image in a sliding window mode to detect small circles with coarse estimates of locations and scales. Next, a hierarchical Bayesian model performs a top-down adaptive circle fitting, with the ability to achieve a maximum a posteriori probability to fit circles to local image features. The evaluation of our approach with manufacturing images has demonstrated to be efficient in detecting small circles in machine parts. Liyuan Li, Fen Fang, Mark D. Rice, Jamie Ng, Wei Xiong 0001, Joo-Hwee Lim |
ICIP | 3 |
| 2018 | Image-based Parking Place Identification for Regulating Shared Bicycle ParkingabstractWe propose a novel method and system to prevent indiscriminate parking of dockless shared bicycles using location-based geo-fencing and image-based parking place identification. The geo-fencing is used to define the approximate regions for different types of bicycle parking regulations. The parking place identification uses a method based on deep Convolutional Neural Network (DCNN) to automatically identify designated bicycle parking places from photos captured by the cyclist using a mobile phone. Combining these two modalities, the parking of shared bicycles can be restricted in designated zones in various environments. Experiments are conducted using photos taken from the designated parking places with different parking indications at various locations. We evaluate the performance of the image-based parking place identification and use heatmaps to analyze potential features that are exploit by the DCNN models. The method achieves high performance on the testing dataset; and the features used for parking place identification are largely consistent with human perceptions. Shudong Xie, Qianli Xu, Fen Fang, Liyuan Li |
ICARCV | 4 |
| 2015 | Identification of faces in line drawings by edge decomposition
Fen Fang, Yong Tsui Lee, Mei Chee Leong |
Pattern Recognit. | 1 |
| 2014 | Efficient decomposition of line drawings of connected manifolds without face identification
Fen Fang, Yong Tsui Lee |
Comput. Aided Des. | 1 |
| 2013 | A Search-and-Validate Method for Face Identification from Single Line DrawingsabstractSeveral studies have been made in finding the faces of an object depicted in a line drawing, but the problem has not been completely solved. Although existing methods can find the correct faces in most cases, there is no mechanism to ascertain that they are indeed correct, leaving the human user to do so. This paper uses a two-stage approach--find potential faces, then validate their correctness--to ensure that only correct faces are delivered ultimately. The face finding itself uses a double breadth-first search algorithm, which yields the shortest path, to find the potential faces. The basic premise is that the smallest faces found are more likely the correct ones. They serve as the "seed" potential faces, from which the algorithm proceeds to search for more faces. If the potential faces found satisfy the validation rules, then they are accepted as correct. Otherwise, the wrong potential faces are identified and removed, and new ones found in their place. The validation process is then repeated. The algorithm is fast and reliable, can deal with planar-faced manifold and nonmanifold objects, and can deliver the different results when a drawing has multiple interpretations. Our extensive tests show that the method can deal with most cases efficiently, including those that previous methods cannot solve. Mei Chee Leong, Yong Tsui Lee, Fen Fang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | A new hybrid method for 3D object recovery from 2D drawings and its validation against the cubic corner method and the optimisation-based method
Yong Tsui Lee, Fen Fang |
Comput. Aided Des. | 2 |
| 2011 | 3D reconstruction of polyhedral objects from single parallel projections using cubic corner
Yong Tsui Lee, Fen Fang |
Comput. Aided Des. | 2 |