EDBT 2026 Demo / reviewers in the wild / expert
Jun Cheng 0003
dblp:78/5816-3
· DBLP profile ↗
89ranked-venue papers
19as first author
45since 2021 · last 2026
0000-0003-1786-6188ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 12 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 30 · 3 first-author · 21 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Heterogeneous Complementary DistillationabstractKnowledge distillation (KD) transfers the ``dark knowledge'' from a complex teacher model to a compact student model. However, heterogeneous architecture distillation, such as Vision Transformer (ViT) to ResNet18, faces challenges due to differences in spatial feature representations. Traditional KD methods are mostly designed for homogeneous architectures and hence struggle to effectively address the disparity. Although heterogeneous KD approaches have been developed recently to solve these issues, they often incur high computational costs and complex designs, or overly rely on logit alignment, which limits their ability to leverage the complementary features. To overcome these limitations, we propose Heterogeneous Complementary Distillation (HCD), a simple yet effective framework that integrates complementary teacher and student features to align representations in shared logits. These logits are decomposed and constrained to facilitate diverse knowledge transfer to the student. Specifically, HCD processes the student’s intermediate features through convolutional projector and adaptive pooling, concatenates them with teacher's feature from the penultimate layer and then maps them via the Complementary Feature Mapper (CFM) module, comprising fully connected layer, to produce shared logits. We further introduce Sub-logit Decoupled Distillation (SDD) that partitions the shared logits into n sub-logits, which are fused with teacher's logits to rectify classification. To ensure sub-logit diversity and reduce redundant knowledge transfer, we propose an Orthogonality Loss (OL). By preserving student-specific strengths and leveraging teacher knowledge, HCD enhances robustness and generalization in students. Extensive experiments on the CIFAR-100, fine-grained (e.g., CUB200, Aircraft) and ImageNet-1K datasets demonstrate that HCD outperforms state-of-the-art KD methods, establishing it as an effective solution for heterogeneous KD. Liuchi Xu, Lu Wang 0001, Lisheng Xu, Jun Cheng 0003 |
AAAI | 5 |
| 2026 | BL-UDA: Towards Unsupervised Domain-Adaptive Surgical Instrument Segmentation with Source Box LabelsabstractRecent advances in unsupervised domain adaptation (UDA) by adapting the model from one domain to another unseen domain have shown considerable promise in improving surgical instrument segmentation performance across domains. However, existing UDA methods primarily rely on pixel-wise labels, which are always difficult to collect due to the labor-intensive annotation process. In this work, we aim to relax the dependence on pixel-level supervision and investigate a challenging UDA setting - source box annotations, where weak supervision and domain shifts coexist. To achieve this, we introduce a novel unsupervised domain adaptation framework, BL-UDA, which leverages bounding box annotations for surgical instrument segmentation across domains. By utilizing the Segment Anything Model (SAM) for pseudo label generation from box annotations, our method effectively bridges object-level and pixel-level domain adaptation. The proposed BL-UDA framework comprises a teacher-student network with entropy minimization for object detection and an entropy-based label selection strategy for generating box prompts to SAM, facilitating pixel-level domain adaptation. Extensive experiments on the EndoVis 2017 and 2018 datasets demonstrate the superiority of BL-UDA over existing UDA methods, significantly mitigating domain shifts and addressing weak supervision challenges with minimal annotation requirements. Ziyuan Zhao, Yifang Yin, Yichen Zhang 0002, Xulei Yang, Jun Cheng 0003, Roger Zimmermann, Cuntai Guan, Shaohua Kevin Zhou |
ICMR | 6 |
| 2026 | Evidential Robust Feature Learning for Generalized Few-Shot Segmentation
Weide Liu, Xiaoyang Zhong, Lu Wang 0001, Chunbo Lang, Yuming Fang 0001, Jun Cheng 0003, Xulei Yang, Gong Cheng 0003 |
Int. J. Comput. Vis. | 6 |
| 2026 | Curvilinear structure-preserving unpaired cross-domain medical image translation
Yi Zhou 0024, Xudong Jiang 0001, Li Chen 0011, Leopold Schmetterer, Bingyao Tan, Jun Cheng 0003 |
Neurocomputing | 7 |
| 2026 | SRNet: Self-supervised structure regularization for stereo matching
Jun Cheng 0003, Zaiwang Gu, Weide Liu, Jiayuan Fan 0001, Zhengguo Li, Chuan-Sheng Foo |
Neurocomputing | 1 |
| 2026 | Integrating SAM Supervision for 3D Weakly Supervised Point Cloud SegmentationabstractCurrent methods for 3D semantic segmentation propose training models with limited annotations to address the difficulty of annotating large, irregular, and unordered 3D point cloud data. They usually focus on the 3D domain only, without leveraging the complementary nature of 2D and 3D data. Besides, some methods extend original labels or generate pseudo labels to guide the training, but they often fail to fully use these labels or address the noise within them. Meanwhile, the emergence of comprehensive and adaptable foundation models has offered effective solutions for segmenting 2D data. Leveraging this advancement, we present a novel approach that maximizes the utility of sparsely available 3D annotations by incorporating segmentation masks generated by 2D foundation models. We further propagate the 2D segmentation masks into the 3D space by establishing geometric correspondences between 3D scenes and 2D views. We extend the highly sparse annotations to encompass the areas delineated by 3D masks, thereby substantially augmenting the pool of available labels. Furthermore, we apply confidence- and uncertainty-based consistency regularization on augmentations of the 3D point cloud and select the reliable pseudo labels, which are further spread on the 3D masks to generate more labels. This innovative strategy bridges the gap between limited 3D annotations and the powerful capabilities of 2D foundation models, ultimately improving the performance of 3D weakly supervised segmentation. Lechun You, Weide Liu, Xulei Yang, Jun Cheng 0003, Wei Zhou 0021, Bharadwaj Veeravalli, Guosheng Lin |
IEEE Trans. Image Process. | 5 |
| 2025 | Debiased Distillation for Consistency RegularizationabstractKnowledge distillation transfers "dark knowledge" from a large teacher model to a smaller student model, yielding a highly efficient network. To improve network's generalization ability, existing works use a larger temperature coefficient for knowledge distillation. Nevertheless, these methods may lower the target category's confidence and lead to ambiguous recognition of similar samples. To mitigate this issue, some studies introduce intra-batch distillation to reduce prediction discrepancy. However, these methods overlook the inconsistency between background information and the target category, which may increase prediction bias due to noise disturbance. Additionally, label imbalance from random sampling and batch size can undermine network generalization reliability. To tackle these challenges, we propose a simple yet effective Intra-class Knowledge Distillation (IKD) method that facilitates knowledge sharing within the same class to ensure consistent predictions. First, we initialize the matrix and the vector to store logits and class counts provided by the teacher, respectively. Then, in the first epoch, we calculate the sum of logits and sample counts per class and perform KD to prevent knowledge omission. Finally, in subsequent training, we update the matrix to obtain the average logits and compute the KL divergence between the student's output and the updated matrix according to the label index. This process ensures intra-class consistency and improves the student's performance. Furthermore, this method theoretically reduces prediction bias by ensuring intra-class consistency. Extensive experiments on the CIFAR-100, ImageNet-1K, and Tiny-ImageNet datasets validate the superiority of IKD. Lu Wang 0001, Liuchi Xu, Zhenhua Huang 0001, Jun Cheng 0003 |
AAAI | 5 |
| 2025 | Rectification-specific Supervision and Constrained Estimator for Online Stereo RectificationabstractOnline stereo rectification is critical for autonomous vehicles and robots in dynamic environments, where factors such as vibration, temperature fluctuations, and mechanical stress can affect rectification accuracy and severely degrade downstream stereo depth estimation. Current dominant approaches for online stereo rectification involve estimating relative camera poses in real time to derive rectification homographies. However, they do not directly optimize for rectification constraints. Additionally, the general-purpose correspondence matchers used in these methods are not trained for rectification, while training of these matchers typically requires ground-truth correspondences which are not available in stereo rectification datasets. To address these limitations, we propose a matching-based stereo rectification framework that is directly optimized for rectification and does not require ground-truth correspondence annotations for training. We assume intrinsics are known as they are generally available on modern devices and are relatively stable. Our framework incorporates a rectification-constrained estimator and applies multi-level, rectification-specific supervision that trains the matcher network for rectification without relying on ground-truth correspondences. Additionally, we create a new rectification dataset with ground-truth optical flow annotations, eliminating bias from evaluation metrics used in prior work that relied on pretrained keypoint matching or optical flow models. Extensive experiments show that our approach outperforms both state-of-the-art matching-based and matching-free methods in vertical flow metric by 10.7% on the Carla-Flowguided dataset and 21.3% on the Semi-Truck Highway dataset, offering superior rectification accuracy. Kim-Hui Yap, Weide Liu, Xulei Yang, Jun Cheng 0003 |
CVPR | 5 |
| 2025 | Local Dense Logit Relations for Enhanced Knowledge DistillationabstractState-of-the-art logit distillation methods exhibit versatility, simplicity, and efficiency. Despite the advances, existing studies have yet to delve thoroughly into fine-grained relationships within logit knowledge. In this paper, we propose Local Dense Relational Logit Distillation (LDRLD), a novel method that captures inter-class relationships through recursively decoupling and recombining logit information, thereby providing more detailed and clearer insights for student learning. To further optimize the performance, we introduce an Adaptive Decay Weight (ADW) strategy, which can dynamically adjust the weights for critical category pairs using Inverse Rank Weighting (IRW) and Exponential Rank Decay (ERD). Specifically, IRW assigns weights inversely proportional to the rank differences between pairs, while ERD adaptively controls weight decay based on total ranking scores of category pairs. Furthermore, after the recursive decoupling, we distill the remaining non-target knowledge to ensure knowledge completeness and enhance performance. Ultimately, our method improves the student's performance by transferring fine-grained knowledge and emphasizing the most critical relationships. Extensive experiments on datasets such as CIFAR-100, ImageNet-1K, and Tiny-ImageNet demonstrate that our method compares favorably with state-of-the-art logit-based distillation approaches. The code will be made publicly available. Liuchi Xu, Jinshuai Liu, Lu Wang 0001, Lisheng Xu, Jun Cheng 0003 |
ICCV | 6 |
| 2025 | Evidential Learning-based Certainty Estimation for Robust Dense Feature MatchingabstractDense feature matching methods aim to estimate a dense correspondence field between images. Inaccurate correspondence can occur due to the presence of unmatchable region, necessitating the need for certainty measurement. This is typically addressed by training a binary classifier to decide whether each predicted correspondence is reliable. However, deep neural network-based classifiers can be vulnerable to image corruptions or perturbations, making it difficult to obtain reliable matching pairs in corrupted scenario. In this work, we propose an evidential deep learning framework to enhance the robustness of dense matching against corruptions. We modify the certainty prediction branch in dense matching models to generate appropriate belief masses and compute the certainty score by taking expectation over the resulting Dirichlet distribution. We evaluate our method on a wide range of benchmarks and show that our method leads to improved robustness against common corruptions and adversarial attacks, achieving up to 10.1\% improvement under severe corruptions. Lile Cai, Chuan-Sheng Foo, Xun Xu 0002, Zaiwang Gu, Jun Cheng 0003, Xulei Yang |
ICLR | 5 |
| 2025 | LiftFeat: 3D Geometry-Aware Local Feature MatchingabstractRobust and efficient local feature matching plays a crucial role in applications such as SLAM and visual localization for robotics. Despite great progress, it is still very challenging to extract robust and discriminative visual features in scenarios with drastic lighting changes, low texture areas, or repetitive patterns. In this paper, we propose a new lightweight network called LiftFeat, which lifts the robustness of raw descriptor by aggregating 3D geometric feature. Specifically, we first adopt a pre-trained monocular depth estimation model to generate pseudo surface normal label, supervising the extraction of 3D geometric feature in terms of predicted surface normal. We then design a 3D geometry-aware feature lifting module to fuse surface normal feature with raw 2D descriptor feature. Integrating such 3D geometric feature enhances the discriminative ability of 2D feature description in extreme conditions. Extensive experimental results on relative pose estimation, homography estimation, and visual localization tasks, demonstrate that our LiftFeat outperforms some lightweight state-of-the-art methods. Code will be released at: https://github.com/lyp-deeplearning/LiftFeat. Yepeng Liu 0002, Wenpeng Lai, Yuxuan Xiong, Jinchi Zhu, Jun Cheng 0003, Yongchao Xu |
ICRA | 6 |
| 2025 | Dual Correlation-Aware Mamba for Microvascular Obstruction Identification in Non-contrast Cine Cardiac Magnetic Resonance
Yige Yan, Jun Cheng 0003, Xulei Yang, Shuang Leng, Ru-San Tan, Liang Zhong 0001, Jagath C. Rajapakse |
MICCAI (1) | 2 |
| 2025 | Improving OCTA Imaging Through Cross-Domain Adaptation: A Noise-Guided Framework Using Intralipid-Enhanced Rat Data
Bingyu Yang, Bingyao Tan, Zaiwang Gu, Leopold Schmetterer, Huiqi Li, Jun Cheng 0003 |
MICCAI (7) | 6 |
| 2025 | Spatiotemporal-Sensitive Network for Microvascular Obstruction Segmentation from Cine Cardiac Magnetic Resonance
Yang Yu 0079, Christopher Kok 0001, Jun Cheng 0003, Shuang Leng, Ru-San Tan, Liang Zhong 0001, Xulei Yang |
MICCAI (16) | 4 |
| 2025 | Robust Incomplete-Modality Alignment for Ophthalmic Disease Grading and Diagnosis via Labeled Optimal Transport
Qinkai Yu, Jianyang Xie, Yitian Zhao, Cheng Chen 0013, Jun Cheng 0003, Lu Liu 0001, Yalin Zheng, Yanda Meng |
MICCAI (15) | 7 |
| 2025 | Intra- and Cross-View Enhancement for OCTA Imaging
Jingbo Zeng, Bingyao Tan, Zaiwang Gu, Shenghua Gao, Leopold Schmetterer, Jun Cheng 0003 |
MICCAI (13) | 6 |
| 2025 | Uncertainty Aware Interest Point Detection and DescriptionabstractInterest point detection and description play an important role in many visual tasks, including image registration, pose estimation, 3D reconstruction, and more. State-of-the-art interest point detection techniques are based on deep neural networks (NNs), which are prone to produce overconfident predictions. However, calibrated and ro-bust uncertainty measurement is crucial when deploying deep NN models in safety critical applications. In this work, we propose a novel Uncertainty-Aware interest Point (UAPoint) detection method to address this problem. Our method leverages evidential learning to learn both aleatoric and epistemic uncertainty. We further propose a constrained sampling scheme to construct more efficient training pairs for the descriptor decoder. We evaluate our method on a wide range of benchmarks and show that our method achieves state-of-the-art performance. Code will be released in https://github.com/JingboZeng/UAPoint. Jingbo Zeng, Zaiwang Gu, Weide Liu, Lile Cai, Jun Cheng 0003 |
WACV | 5 |
| 2025 | Improving multi-modal brain tumor segmentation via pre-training and knowledge distillation based post-training
Weide Liu, Jingwen Hou, Xiaoyang Zhong, Huijing Zhan, Jun Cheng 0003, Yuming Fang 0001, Guanghui Yue 0001 |
Neurocomputing | 5 |
| 2025 | Physically-guided open vocabulary segmentation with weighted patched alignment loss
Weide Liu, Jieming Lou, Wei Zhou 0021, Jun Cheng 0003, Xulei Yang |
Neurocomputing | 5 |
| 2025 | Multimodal multitask similarity learning for vision language model on radiological images and reports
Yang Yu 0079, Weide Liu, Ivan Ho Mien, Pavitra Krishnaswamy, Xulei Yang, Jun Cheng 0003 |
Neurocomputing | 7 |
| 2025 | MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image MatchingabstractMany keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality image matching, they often struggle with multimodal data because the descriptors trained on single-modality data tend to lack robustness against the non-linear variations present in multimodal data. Extending such methods to multimodal image matching often requires well-aligned multimodal data to learn modality-invariant descriptors. However, acquiring such data is often costly and impractical in many real-world scenarios. To address this challenge, we propose a modality-invariant feature learning network (MIFNet) to compute modality-invariant features for keypoint descriptions in multimodal image matching using only single-modality training data. Specifically, we propose a novel latent feature aggregation module and a cumulative hybrid aggregation module to enhance the base keypoint descriptors trained on single-modality data by leveraging pre-trained features from Stable Diffusion models. We validate our method with recent keypoint detection and description methods in three multimodal retinal image datasets (CF-FA, CF-OCT, EMA-OCTA) and two remote sensing datasets (Optical-SAR and Optical-NIR). Extensive experiments demonstrate that the proposed MIFNet is able to learn modality-invariant feature for multimodal image matching without accessing the targeted modality and has good zero-shot generalization ability. The code will be released at https://github.com/lyp-deeplearning/MIFNet. Yepeng Liu 0002, Zhichao Sun 0004, Baosheng Yu, Yitian Zhao, Bo Du 0001, Yongchao Xu, Jun Cheng 0003 |
IEEE Trans. Image Process. | 7 |
| 2025 | Masked Vascular Structure Segmentation and Completion in Retinal ImagesabstractEarly retinal vascular changes in diseases such as diabetic retinopathy often occur at a microscopic level. Accurate evaluation of retinal vascular networks at a micro-level could significantly improve our understanding of angiopathology and potentially aid ophthalmologists in disease assessment and management. Multiple angiogram-related retinal imaging modalities, including fundus, optical coherence tomography angiography, and fluorescence angiography, project continuous, inter-connected retinal microvascular networks into imaging domains. However, extracting the microvascular network, which includes arterioles, venules, and capillaries, is challenging due to the limited contrast and resolution. As a result, the vascular network often appears as fragmented segments. In this paper, we propose a backbone-agnostic Masked Vascular Structure Segmentation and Completion (MaskVSC) method to reconstruct the retinal vascular network. MaskVSC simulates missing sections of blood vessels and uses this simulation to train the model to predict the missing parts and their connections. This approach simulates highly heterogeneous forms of vessel breaks and mitigates the need for massive data labeling. Accordingly, we introduce a connectivity loss function that penalizes interruptions in the vascular network. Our findings show that masking 40% of the segments yields optimal performance in reconstructing the interconnected vascular network. We test our method on three different types of retinal images across five separate datasets. The results demonstrate that MaskVSC outperforms state-of-the-art methods in maintaining vascular network completeness and segmentation accuracy. Furthermore, MaskVSC has been introduced to different segmentation backbones and has successfully improved performance. The code and 2PFM data are available at: https://github.com/Zhouyi-Zura/MaskVSC. Yi Zhou 0024, Thiara Sana Ahmed, Meng Wang 0038, Eric A. Newman, Leopold Schmetterer, Huazhu Fu, Jun Cheng 0003, Bingyao Tan |
IEEE Trans. Medical Imaging | 7 |
| 2024 | SuperJunction: Learning-Based Junction Detection for Retinal Image RegistrationabstractKeypoints-based approaches have shown to be promising for retinal image registration, which superimpose two or more images from different views based on keypoint detection and description. However, existing approaches suffer from ineffective keypoint detector and descriptor training. Meanwhile, the non-linear mapping from 3D retinal structure to 2D images is often neglected. In this paper, we propose a novel learning-based junction detection approach for retinal image registration, which enhances both the keypoint detector and descriptor training. To improve the keypoint detection, it uses a multi-task vessel detection to regularize the model training, which helps to learn more representative features and reduce the risk of over-fitting. To achieve effective training for keypoints description, a new constrained negative sampling approach is proposed to compute the descriptor loss. Moreover, we also consider the non-linearity between retinal images from different views during matching. Experimental results on FIRE dataset show that our method achieves mean area under curve of 0.850, which is 12.6% higher than 0.755 by the state-of-the-art method. All the codes are available at https://github.com/samjcheng/SuperJunction. Zaiwang Gu, Weide Liu, Wee Siong Ng, Weimin Huang 0002, Jun Cheng 0003 |
AAAI | 7 |
| 2024 | Progressive Retinal Image Registration via Global and Local Deformable TransformationsabstractRetinal image registration plays an important role in the ophthalmological diagnosis process. Since there exist variances in viewing angles and anatomical structures across different retinal images, keypoint-based approaches become the mainstream methods for retinal image registration thanks to their robustness and low latency. These methods typically assume the retinal surfaces are planar, and adopt feature matching to obtain the homography matrix that represents the global transformation between images. Yet, such a planar hypothesis inevitably introduces registration errors since retinal surface is approximately curved. This limitation is more prominent when registering image pairs with significant differences in viewing angles. To address this problem, we propose a hybrid registration framework called HybridRetina, which progressively registers retinal images with global and local deformable transformations. For that, we use a keypoint detector and a deformation network called GAMorph to estimate the global transformation and local deformable transformation, respectively. Specifically, we integrate multi-level pixel relation knowledge to guide the training of GAMorph. Additionally, we utilize an edge attention module that includes the geometric priors of the images, ensuring the deformation field focuses more on the vascular regions of clinical interest. Experiments on two widely-used datasets, FIRE and FLoRI21, show that our proposed HybridRetina significantly outperforms some state-of-the-art methods. The code is available at https://github.com/lyp-deeplearning/awesome-retinal-registration. Yepeng Liu 0002, Baosheng Yu, Yuliang Gu, Bo Du 0001, Yongchao Xu, Jun Cheng 0003 |
BIBM | 7 |
| 2024 | Learning Intra-View and Cross-View Geometric Knowledge for Stereo MatchingabstractGeometric knowledge has been shown to be beneficial for the stereo matching task. However, prior attempts to in-tegrate geometric insights into stereo matching algorithms have largely focused on geometric knowledge from single images while crucial cross-view factors such as occlusion and matching uniqueness have been overlooked. To address this gap, we propose a novel Intra-view and Cross-view Geometric knowledge learning Network (ICGNet), specifically crafted to assimilate both intra-view and cross-view geo-metric knowledge. ICGNet harnesses the power of interest points to serve as a channel for intra-view geometric understanding. Simultaneously, it employs the correspon-dences among these points to capture cross-view geometric relationships. This dual incorporation empowers the proposed ICGNet to leverage both intra-view and cross-view geometric knowledge in its learning process, substantially improving its ability to estimate disparities. Our extensive experiments demonstrate the superiority of the ICGNet over contemporary leading models. The code will be available at https://github.com/DFSDDDDDl199/ICGNet. Weide Liu, Zaiwang Gu, Xulei Yang, Jun Cheng 0003 |
CVPR | 5 |
| 2024 | Latent Degradation Representation Constraint for Single Image DerainingabstractSince rain shows a variety of shapes and directions, learning the degradation representation is extremely challenging for single image deraining. Existing methods mainly propose to designing complicated modules to implicitly learn latent degradation representation from rainy images. However, it is hard to decouple the content-independent degradation representation due to the lack of explicit constraint, resulting in over- or under-enhancement problems. To tackle this issue, we propose a novel Latent Degradation Representation Constraint Network (LDRCNet) that consists of the Direction-Aware Encoder (DAEncoder), Deraining Network, and Multi-Scale Interaction Block (MSIBlock). Specifically, the DAEncoder is proposed to extract latent degradation representation adaptively by first using the deformable convolutions to exploit the direction property of rain streaks. Next, a constraint loss is introduced to explicitly constraint the degradation representation learning during training. Last, we propose an MSIBlock to fuse with the learned degradation representation and decoder features of the deraining network for adaptive information interaction to remove various complicated rainy patterns and reconstruct image details. Experimental results on five synthetic and four real datasets demonstrate that our method achieves state-of-the-art performance. The source code is available at https://github.com/Madeline-hyh/LDRCNet. Long Peng 0003, Lu Wang 0001, Jun Cheng 0003 |
ICASSP | 4 |
| 2024 | Coarse-Grained Mask Regularization for Microvascular Obstruction Identification from Non-contrast Cardiac Magnetic Resonance
Yige Yan, Jun Cheng 0003, Xulei Yang, Zaiwang Gu, Shuang Leng, Ru-San Tan, Liang Zhong 0001, Jagath C. Rajapakse |
MICCAI (1) | 2 |
| 2024 | On-the-fly Point Feature Representation for Point Clouds AnalysisabstractPoint cloud analysis is challenging due to its unique characteristics of unorderness, sparsity and irregularity. Prior works attempt to capture local relationships by convolution operations or attention mechanisms, exploiting geometric information from coordinates implicitly. These methods, however, are insufficient to describe the explicit local geometry, e.g., curvature and orientation. In this paper, we propose On-the-fly Point Feature Representation (OPFR), which captures abundant geometric information explicitly through Curve Feature Generator module. This is inspired by Point Feature Histogram (PFH) from computer vision community. However, the utilization of vanilla PFH encounters great difficulties when applied to large datasets and dense point clouds, as it demands considerable time for feature generation. In contrast, we introduce the Local Reference Constructor module, which approximates the local coordinate systems based on triangle sets. Owing to this, our OPFR only requires extra 1.56ms for inference (65X faster than vanilla PFH) and 0.012M more parameters, and it can serve as a versatile plug-and-play module for various backbones, particularly MLP-based and Transformer-based backbones examined in this study. Additionally, we introduce the novel Hierarchical Sampling module aimed at enhancing the quality of triangle sets, thereby ensuring robustness of the obtained geometric features. Our proposed method improves overall accuracy (OA) on ModelNet40 from 90.7% to 94.5% (+3.8%) for classification, and OA on S3DIS Area-5 from 86.4% to 90.0% (+3.6%) for semantic segmentation, respectively, building upon PointNet++ backbone. When integrated with Point Transformer backbone, we achieve state-of-the-art results on both tasks: 94.8% OA on ModelNet40 and 91.7% OA on S3DIS Area-5. Jiangyi Wang, Zhongyao Cheng, Na Zhao 0004, Jun Cheng 0003, Xulei Yang |
ACM Multimedia | 4 |
| 2024 | KA-Seg: Improving LiDAR Point Cloud
Kaining Cui, Lu Wang 0001, Jun Cheng 0003 |
PRCV (10) | 4 |
| 2024 | VPFNET: A Scale-Adaptive Voxel Point Fusion Network for Semantic Segmentation of Point Clouds
Kaining Cui, Lu Wang 0001, Zhenfei Liu, Bingxin Yu, Jun Cheng 0003 |
PRCV (10) | 7 |
| 2024 | ASPVNet: Attention Based Sparse Point-Voxel Network for 3D Object Detection
Bingxin Yu, Lu Wang 0001, Jun Cheng 0003 |
PRCV (10) | 5 |
| 2024 | Harmonizing Base and Novel Classes: A Class-Contrastive Approach for Generalized Few-Shot Segmentation
Weide Liu, Yuming Fang 0001, Chuan-Sheng Foo, Jun Cheng 0003, Guosheng Lin |
Int. J. Comput. Vis. | 6 |
| 2024 | Structure and Intensity Unbiased Translation for 2D Medical Image SegmentationabstractData distribution gaps often pose significant challenges to the use of deep segmentation models. However, retraining models for each distribution is expensive and time-consuming. In clinical contexts, device-embedded algorithms and networks, typically unretrainable and unaccessable post-manufacture, exacerbate this issue. Generative translation methods offer a solution to mitigate the gap by transferring data across domains. However, existing methods mainly focus on intensity distributions while ignoring the gaps due to structure disparities. In this paper, we formulate a new image-to-image translation task to reduce structural gaps. We propose a simple, yet powerful Structure-Unbiased Adversarial (SUA) network which accounts for both intensity and structural differences between the training and test sets for segmentation. It consists of a spatial transformation block followed by an intensity distribution rendering module. The spatial transformation block is proposed to reduce the structural gaps between the two images. The intensity distribution rendering module then renders the deformed structure to an image with the target intensity distribution. Experimental results show that the proposed SUA method has the capability to transfer both intensity distribution and structural content between multiple pairs of datasets and is superior to prior arts in closing the gaps for improving segmentation. Tianyang Miller, Shaoming Zheng, Jun Cheng 0003, Xi Jia, Joseph Bartlett, Xinxing Cheng, Zhaowen Qiu, Huazhu Fu, Jiang Liu 0001, Ales Leonardis, Jinming Duan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | FedCov: Enhanced Trustworthy Federated Learning for Machine RUL Prediction With Continuous-to-Discrete ConversionabstractNumerous approaches have been proposed for predicting machine remaining useful life (RUL), which helps prevent unnecessary downtime and reduces the maintenance cost in industrial systems. Most existing RUL methods rely on centralized learning and require large-scale datasets with manual labels, which are infeasible to collect. As a decentralized learning paradigm, federated learning (FL) has recently been integrated into these approaches, which aims to utilize the distributed data from local users for model training, while preserving their data privacy. However, data heterogeneity in the industry poses a critical challenge for FL, leading to model drifting issue and degraded global model performance. A straightforward method to tackle this problem is to estimate the data distribution of clients. However, it is difficult to apply this method to a regression task, since the prediction space is continuous, which increases the difficulty in estimating the data distribution. Motivated by the digital-to-analog converter in electronics, we propose a novel approach called FedCov, which involves a converter module that transforms continuous RUL values into discrete categories. Subsequently, a generator is trained to aggregate user information based on the discrete label distribution, and it is broadcasted to users as a data enhancement tool to address the data heterogeneity problem. Furthermore, to improve the performance and reliability of our FedCov, an uncertainty estimation module is proposed, which utilizes the confidence level of model predictions to adjust the training direction. Extensive experiments are conducted on theC-MAPSSbenchmark, which demonstrates that our proposed FedCov effectively solves the model drift issues and improves the performances on the RUL task, achieving state-of-the-arts performances. Yuming Fang 0001, Weide Liu, Ruibing Jin, Jun Cheng 0003, Zhenghua Chen |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Target-Guided Diffusion Models for Unpaired Cross-Modality Medical Image TranslationabstractIn a clinical setting, the acquisition of certain medical image modality is often unavailable due to various considerations such as cost, radiation, etc. Therefore, unpaired cross-modality translation techniques, which involve training on the unpaired data and synthesizing the target modality with the guidance of the acquired source modality, are of great interest. Previous methods for synthesizing target medical images are to establish one-shot mapping through generative adversarial networks (GANs). As promising alternatives to GANs, diffusion models have recently received wide interests in generative tasks. In this paper, we propose a target-guided diffusion model (TGDM) for unpaired cross-modality medical image translation. For training, to encourage our diffusion model to learn more visual concepts, we adopted a perception prioritized weight scheme (P2W) to the training objectives. For sampling, a pre-trained classifier is adopted in the reverse process to relieve modality-specific remnants from source data. Experiments on both brain MRI-CT and prostate MRI-US datasets demonstrate that the proposed method achieves a visually realistic result that mimics a vivid anatomical section of the target organ. In addition, we have also conducted a subjective assessment based on the synthesized samples to further validate the clinical value of TGDM. Yimin Luo, Qinyu Yang, Ziyi Liu 0010, Zenglin Shi, Weimin Huang 0002, Guoyan Zheng, Jun Cheng 0003 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | ELFNet: Evidential Local-global Fusion for Stereo MatchingabstractAlthough existing stereo matching models have achieved continuous improvement, they often face issues related to trustworthiness due to the absence of uncertainty estimation. Additionally, effectively leveraging multi-scale and multi-view knowledge of stereo pairs remains unexplored. In this paper, we introduce the Evidential Local-global Fusion (ELF) framework for stereo matching, which endows both uncertainty estimation and confidence-aware fusion with trustworthy heads. Instead of predicting the disparity map alone, our model estimates an evidential-based disparity considering both aleatoric and epistemic uncertainties. With the normal inverse-Gamma distribution as a bridge, the proposed framework realizes intra evidential fusion of multi-level predictions and inter evidential fusion between cost-volume-based and transformer-based stereo matching. Extensive experimental results show that the proposed framework exploits multi-view information effectively and achieves state-of-the-art overall performance both on accuracy and cross-domain generalization. The codes are available at https://github.com/jimmy19991222/ELFNet. Jieming Lou, Weide Liu, Fayao Liu, Jun Cheng 0003 |
ICCV | 5 |
| 2023 | RFDNet: Real-Time 3D Object Detection Via Range Feature DecorationabstractHigh-performance real-time 3D object detection is crucial in autonomous driving perception systems. Voxel-or point-based 3D object detectors are highly accurate but inefficient and difficult to deploy, while other methods use 2D projection views to improve efficiency, but information loss usually degrades performance. To balance effectiveness and efficiency, we propose a scheme called RFDNet that uses range features to decorate points. Specifically, RFDNet adaptively aggregates point features projected to independent grids and nearby regions via Dilated Grid Feature Encoding (DGFE) to generate a range view, which can handle occlusion and multi-frame inputs while the established geometric correlation between grid with surrounding space weakens the effects of scale distortion. We also propose a Soft Box Regression (SBR) strategy that supervises 3D box regression on a more extensive range than conventional methods to enhance model robustness. In addition, RFDNet benefits from our designed Semantic-assisted Ground-truth Sample (SA-GTS) data augmentation, which additionally considers collisions and spatial distributions of objects. Experiments on the nuScenes benchmark show that RFDNet outperforms all LiDAR-only non-ensemble 3D object detectors and runs at high speed of 20 FPS, achieving a better effectiveness-efficiency trade-off. Code is available at https://github.com/wy17646051/RFDNet. Hongda Chang, Lu Wang 0001, Jun Cheng 0003 |
IROS | 3 |
| 2023 | CLC-Net: Contextual and local collaborative network for lesion segmentation in diabetic retinopathy images
Yuqi Fang, Sen Yang 0006, Delong Zhu 0001, Jing Zhang 0051, Jun Zhang 0018, Jun Cheng 0003, Raymond Kai-Yu Tong, Xiao Han 0011 |
Neurocomputing | 8 |
| 2023 | Perceptual Quality Assessment of Enhanced Colonoscopy Images: A Benchmark Dataset and an Objective MethodabstractIn colonoscopy, the captured images are usually with low-quality appearance, such as non-uniform illumination, low contrast, etc., due to the specialized imaging environment, which may provide poor visual feedback and bring challenges to subsequent disease analysis. Many low-light image enhancement (LIE) algorithms have recently proposed to improve the perceptual quality. However, how to fairly evaluate the quality of enhanced colonoscopy images (ECIs) generated by different LIE algorithms remains a rarely-mentioned and challenging problem. In this study, we carry out a pioneering investigation on perceptual quality assessment of ECIs. Firstly, considering the lack of specific datasets, we collect 300 low-light images with diverse contents during the real-world colonoscopy and conduct rigorous subjective studies to compare the performance of 8 popular LIE methods, resulting in a benchmark dataset (named ECIQAD) for ECIs. Secondly, in view of the distinctive distortion characteristics of ECIs, we propose an effective no-reference Enhanced Colonoscopy Image Quality (ECIQ) method to automatically evaluate the perceptual quality of ECIs via analysis of brightness, contrast, colorfulness, naturalness, and noise. Extensive experiments on ECIQAD demonstrate the superiority of our proposed ECIQ method over 14 mainstream no-reference image quality assessment methods. Guanghui Yue 0001, Tianwei Zhou, Jingwen Hou, Weide Liu, Long Xu 0001, Tianfu Wang 0001, Jun Cheng 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2023 | CPP-Net: Context-Aware Polygon Proposal Network for Nucleus SegmentationabstractNucleus segmentation is a challenging task due to the crowded distribution and blurry boundaries of nuclei. Recent approaches represent nuclei by means of polygons to differentiate between touching and overlapping nuclei and have accordingly achieved promising performance. Each polygon is represented by a set of centroid-to-boundary distances, which are in turn predicted by features of the centroid pixel for a single nucleus. However, using the centroid pixel alone does not provide sufficient contextual information for robust prediction and thus degrades the segmentation accuracy. To handle this problem, we propose a Context-aware Polygon Proposal Network (CPP-Net) for nucleus segmentation. First, we sample a point set rather than one single pixel within each cell for distance prediction. This strategy substantially enhances contextual information and thereby improves the robustness of the prediction. Second, we propose a Confidence-based Weighting Module, which adaptively fuses the predictions from the sampled point set. Third, we introduce a novel Shape-Aware Perceptual (SAP) loss that constrains the shape of the predicted polygons. Here, the SAP loss is based on an additional network that is pre-trained by means of mapping the centroid probability map and the pixel-to-boundary distance maps to a different nucleus representation. Extensive experiments justify the effectiveness of each component in the proposed CPP-Net. Finally, CPP-Net is found to achieve state-of-the-art performance on three publicly available databases, namely DSB2018, BBBC06, and PanNuke. Code of this paper is available at https://github.com/csccsccsccsc/cpp-net. Shengcong Chen, Changxing Ding, Minfeng Liu, Jun Cheng 0003, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2022 | Proxy-Bridged Image Reconstruction Network for Anomaly Detection in Medical ImagesabstractAnomaly detection in medical images refers to the identification of abnormal images with only normal images in the training set. Most existing methods solve this problem with a self-reconstruction framework, which tends to learn an identity mapping and reduces the sensitivity to anomalies. To mitigate this problem, in this paper, we propose a novel Proxy-bridged Image Reconstruction Network (ProxyAno) for anomaly detection in medical images. Specifically, we use an intermediate proxy to bridge the input image and the reconstructed image. We study different proxy types, and we find that the superpixel-image (SI) is the best one. We set all pixels' intensities within each superpixel as their average intensity, and denote this image as SI. The proposed ProxyAno consists of two modules, a Proxy Extraction Module and an Image Reconstruction Module. In the Proxy Extraction Module, a memory is introduced to memorize the feature correspondence for normal image to its corresponding SI, while the memorized correspondence does not apply to the abnormal images, which leads to the information loss for abnormal image and facilitates the anomaly detection. In the Image Reconstruction Module, we map an SI to its reconstructed image. Further, we crop a patch from the image and paste it on the normal SI to mimic the anomalies, and enforce the network to reconstruct the normal image even with the pseudo abnormal SI. In this way, our network enlarges the reconstruction error for anomalies. Extensive experiments on brain MR images, retinal OCT images and retinal fundus images verify the effectiveness of our method for both image-level and pixel-level anomaly detection. Kang Zhou 0001, Jing Li 0117, Weixin Luo, Jianlong Yang, Huazhu Fu, Jun Cheng 0003, Jiang Liu 0001, Shenghua Gao |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Memorizing Structure-Texture Correspondence for Image Anomaly DetectionabstractThis work focuses on image anomaly detection by leveraging only normal images in the training phase. Most previous methods tackle anomaly detection by reconstructing the input images with an autoencoder (AE)-based model, and an underlying assumption is that the reconstruction errors for the normal images are small, and those for the abnormal images are large. However, these AE-based methods, sometimes, even reconstruct the anomalies well; consequently, they are less sensitive to anomalies. To conquer this issue, we propose to reconstruct the image by leveraging the structure-texture correspondence. Specifically, we observe that, usually, for normal images, the texture can be inferred from its corresponding structure (e.g., the blood vessels in the fundus image and the structured anatomy in optical coherence tomography image), while it is hard to infer the texture from a destroyed structure for the abnormal images. Therefore, a structure-texture correspondence memory (STCM) module is proposed to reconstruct image texture from its structure, where a memory mechanism is used to characterize the mapping from the normal structure to its corresponding normal texture. As the correspondence between destroyed structure and texture cannot be characterized by the memory, the abnormal images would have a larger reconstruction error, facilitating anomaly detection. In this work, we utilize two kinds of complementary structures (i.e., the semantic structure with human-labeled category information and the low-level structure with abundant details), which are extracted by two structure extractors. The reconstructions from the two kinds of structures are fused together by a learned attention weight to get the final reconstructed image. We further feed the reconstructed image into the two aforementioned structure extractors to extract structures. On the one hand, constraining the consistency between the structures extracted from the original input and that from the reconstructed image would regularize the network training; on the other hand, the error between the structures extracted from the original input and that from the reconstructed image can also be used as a supplement measurement to identify the anomaly. Extensive experiments validate the effectiveness of our method for image anomaly detection on both industrial inspection images and medical images. Kang Zhou 0001, Jing Li 0117, Jianlong Yang, Jun Cheng 0003, Wen Liu 0003, Weixin Luo, Jiang Liu 0001, Shenghua Gao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | MOS: A Low Latency and Lightweight Framework for Face Detection, Landmark Localization, and Head Pose Estimation
Yepeng Liu 0002, Zaiwang Gu, Shenghua Gao, Yusheng Zeng, Jun Cheng 0003 |
BMVC | 6 |
| 2021 | CS2-Net: Deep learning segmentation of curvilinear structures in medical imaging
Lei Mou, Yitian Zhao, Huazhu Fu, Yonghuai Liu, Jun Cheng 0003, Yalin Zheng, Pan Su 0001, Jianlong Yang, Li Chen 0011, Alejandro F. Frangi, Masahiro Akiba, Jiang Liu 0001 |
Medical Image Anal. | 5 |
| 2021 | Structure and Illumination Constrained GAN for Medical Image EnhancementabstractThe development of medical imaging techniques has greatly supported clinical decision making. However, poor imaging quality, such as non-uniform illumination or imbalanced intensity, brings challenges for automated screening, analysis and diagnosis of diseases. Previously, bi-directional GANs (e.g., CycleGAN), have been proposed to improve the quality of input images without the requirement of paired images. However, these methods focus on global appearance, without imposing constraints on structure or illumination, which are essential features for medical image interpretation. In this paper, we propose a novel and versatile bi-directional GAN, named Structure and illumination constrained GAN (StillGAN), for medical image quality enhancement. Our StillGAN treats low- and high-quality images as two distinct domains, and introduces local structure and illumination constraints for learning both overall characteristics and local details. Extensive experiments on three medical image datasets (e.g., corneal confocal microscopy, retinal color fundus and endoscopy images) demonstrate that our method performs better than both conventional methods and other deep learning-based methods. In addition, we have investigated the impact of the proposed method on different medical image analysis and clinical tasks such as nerve segmentation, tortuosity grading, fovea localization and disease classification. Yuhui Ma, Jiang Liu 0001, Yonghuai Liu, Huazhu Fu, Jun Cheng 0003, Yufei Wu 0013, Jiong Zhang 0004, Yitian Zhao |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Encoding Structure-Texture Relation with P-Net for Anomaly Detection in Retinal Images
Kang Zhou 0001, Jianlong Yang, Jun Cheng 0003, Wen Liu 0003, Weixin Luo, Zaiwang Gu, Jiang Liu 0001, Shenghua Gao |
ECCV (20) | 4 |
| 2020 | Cycle Structure and Illumination Constrained GAN for Medical Image Enhancement
Yuhui Ma, Yonghuai Liu, Jun Cheng 0003, Yalin Zheng, Morteza Ghahremani, Honghan Chen, Jiang Liu 0001, Yitian Zhao |
MICCAI (2) | 3 |
| 2020 | HARaaS: HAR as a service using wifi signal in IoT-enabled edge computing: poster abstractabstractHuman activity recognition (HAR) is an important component in context awareness IoT applications such smart home, smart building etc. With the proliferation of WiFi-integrated devices, researchers exploit WiFi signals to recognize various human activities. In this work, we introduce a HAR as a Service (HARaaS) model for activity recognition services applied in IoT areas. HARaaS proposes a novel edge computing model in the concept of the Sensing as a Service (S2aaS) architecture to offer accurate and real-time activities recognition services with good energy efficiency. HARaaS distributes the resource-hungry computing workload i.e. training recognition model to edge terminals, and exploits the built-in intelligence of IoT devices. A WiFi-based activity recognition service is designed following the HARaaS architecture, and the lightweight machine learning and deep learning model are incorporated in the service for accurate activity recognition. Experiments are conducted and demonstrate the service achieves an activity recognition accuracy of 95% with extremely low latency and high energy efficiency. Jin Zhang 0013, Bo Wei 0003, Jun Cheng 0003 |
SenSys | 3 |
| 2020 | Speckle reduction of OCT via super resolution reconstruction and its application on retinal layer segmentation
Qifeng Yan, Bang Chen, Jun Cheng 0003, Jianlong Yang, Jiang Liu 0001, Yitian Zhao |
Artif. Intell. Medicine | 4 |
| 2020 | Glove-type massage feature acquisition by human organ simulators: Feasibility and comparison
Jiangting Liu, Jun Cheng 0003 |
Signal Process. Image Commun. | 4 |
| 2020 | Guest Editorial Ophthalmic Image Analysis and InformaticsabstractThe papers in this special son focus on This special issue invited contributions reporting on methodological breakthroughs in artificial intelligence in ophthalmology, and systems and insights that make use of large-scale datasets linking across multiple imaging modalities, image phenotyping, and imaging omics. These papers presented in this special issue introduce the latest advances in the field of ophthalmic image analysis and informatics, which enable and drive the research, development, and application of key technologies into ocular healthcare. Jun Cheng 0003, Huazhu Fu, Delia Cabrera DeBuc, Jie Tian 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Correction to "Noise Adaptation Generative Adversarial Network for Medical Image Analysis"abstractIn the above article[1],Tables II,III, andVandFig. 6are incorrect. The correct images are provided below: Tianyang Miller, Jun Cheng 0003, Huazhu Fu, Zaiwang Gu, Kang Zhou 0001, Shenghua Gao, Ru Zheng, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Dense Dilated Network With Probability Regularized Walk for Vessel DetectionabstractThe detection of retinal vessel is of great importance in the diagnosis and treatment of many ocular diseases. Many methods have been proposed for vessel detection. However, most of the algorithms neglect the connectivity of the vessels, which plays an important role in the diagnosis. In this paper, we propose a novel method for retinal vessel detection. The proposed method includes a dense dilated network to get an initial detection of the vessels and a probability regularized walk algorithm to address the fracture issue in the initial detection. The dense dilated network integrates newly proposed dense dilated feature extraction blocks into an encoder-decoder structure to extract and accumulate features at different scales. A multi-scale Dice loss function is adopted to train the network. To improve the connectivity of the segmented vessels, we also introduce a probability regularized walk algorithm to connect the broken vessels. The proposed method has been applied on three public data sets: DRIVE, STARE and CHASE_DB1. The results show that the proposed method outperforms the state-of-the-art methods in accuracy, sensitivity, specificity and also area under receiver operating characteristic curve. Lei Mou, Li Chen 0011, Jun Cheng 0003, Zaiwang Gu, Yitian Zhao, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Noise Adaptation Generative Adversarial Network for Medical Image AnalysisabstractMachine learning has been widely used in medical image analysis under an assumption that the training and test data are under the same feature distributions. However, medical images from difference devices or the same device with different parameter settings are often contaminated with different amount and types of noises, which violate the above assumption. Therefore, the models trained using data from one device or setting often fail to work for that from another. Moreover, it is very expensive and tedious to label data and re-train models for all different devices or settings. To overcome this noise adaptation issue, it is necessary to leverage on the models trained with data from one device or setting for new data. In this paper, we reformulate this noise adaptation task as an image-to-image translation task such that the noise patterns from the test data are modified to be similar to those from the training data while the contents of the data are unchanged. In this paper, we propose a novel Noise Adaptation Generative Adversarial Network (NAGAN), which contains a generator and two discriminators. The generator aims to map the data from source domain to target domain. Among the two discriminators, one discriminator enforces the generated images to have the same noise patterns as those from the target domain, and the second discriminator enforces the content to be preserved in the generated images. We apply the proposed NAGAN on both optical coherence tomography (OCT) images and ultrasound images. Results show that the method is able to translate the noise style. In addition, we also evaluate our proposed method with segmentation task in OCT and classification task in ultrasound. The experimental results show that the proposed NAGAN improves the analysis outcome. Tianyang Miller, Jun Cheng 0003, Huazhu Fu, Zaiwang Gu, Kang Zhou 0001, Shenghua Gao, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Topology Reconstruction of Tree-Like Structure in Images via Structural Similarity Measure and Dominant Set ClusteringabstractThe reconstruction and analysis of tree-like topological structures in the biomedical images is crucial for biologists and surgeons to understand biomedical conditions and plan surgical procedures. The underlying tree-structure topology reveals how different curvilinear components are anatomically connected to each other. Existing automated topology reconstruction methods have great difficulty in identifying the connectivity when two or more curvilinear components cross or bifurcate, due to their projection ambiguity, imaging noise and low contrast. In this paper, we propose a novel curvilinear structural similarity measure to guide a dominant-set clustering approach to address this indispensable issue. The novel similarity measure takes into account both intensity and geometric properties in representing the curvilinear structure locally and globally, and group curvilinear objects at crossover points into different connected branches by dominant-set clustering. The proposed method is applicable to different imaging modalities, and quantitative and qualitative results on retinal vessel, plant root, and neuronal network datasets show that our methodology is capable of advancing the current state-of-the-art techniques. Jianyang Xie, Yitian Zhao, Yonghuai Liu, Pan Su 0001, Yifan Zhao 0001, Jun Cheng 0003, Yalin Zheng, Jiang Liu 0001 |
CVPR | 6 |
| 2019 | A Deep Step Pattern Representation for Multimodal Retinal Image RegistrationabstractThis paper presents a novel feature-based method that is built upon a convolutional neural network (CNN) to learn the deep representation for multimodal retinal image registration. We coined the algorithm deep step patterns, in short DeepSPa. Most existing deep learning based methods require a set of manually labeled training data with known corresponding spatial transformations, which limits the size of training datasets. By contrast, our method is fully automatic and scale well to different image modalities with no human intervention. We generate feature classes from simple step patterns within patches of connecting edges formed by vascular junctions in multiple retinal imaging modalities. We leverage CNN to learn and optimize the input patches to be used for image registration. Spatial transformations are estimated based on the output possibility of the fully connected layer of CNN for a pair of images. One of the key advantages of the proposed algorithm is its robustness to non-linear intensity changes, which widely exist on retinal images due to the difference of acquisition modalities. We validate our algorithm on extensive challenging datasets comprising poor quality multimodal retinal images which are adversely affected by pathologies (diseases), speckle noise and low resolutions. The experimental results demonstrate the robustness and accuracy over state-of-the-art multimodal image registration algorithms. Jimmy Addison Lee, Peng Liu 0049, Jun Cheng 0003, Huazhu Fu |
ICCV | 3 |
| 2019 | Ki-GAN: Knowledge Infusion Generative Adversarial Network for Photoacoustic Image Reconstruction In Vivo
Hengrong Lan, Kang Zhou 0001, Jun Cheng 0003, Jiang Liu 0001, Shenghua Gao, Fei Gao 0010 |
MICCAI (1) | 4 |
| 2019 | CS-Net: Channel and Spatial Attention Network for Curvilinear Structure Segmentation
Lei Mou, Yitian Zhao, Li Chen 0011, Jun Cheng 0003, Zaiwang Gu, Huaying Hao, Yalin Zheng, Alejandro F. Frangi, Jiang Liu 0001 |
MICCAI (1) | 4 |
| 2019 | SkrGAN: Sketching-Rendering Unconditional Generative Adversarial Networks for Medical Image Synthesis
Tianyang Miller, Huazhu Fu, Yitian Zhao, Jun Cheng 0003, Mengjie Guo, Zaiwang Gu, Shenghua Gao, Jiang Liu 0001 |
MICCAI (4) | 4 |
| 2019 | CE-Net: Context Encoder Network for 2D Medical Image SegmentationabstractMedical image segmentation is an important step in medical image analysis. With the rapid development of a convolutional neural network in image processing, deep learning has been used for medical image segmentation, such as optic disc segmentation, blood vessel detection, lung segmentation, cell segmentation, and so on. Previously, U-net based approaches have been proposed. However, the consecutive pooling and strided convolutional operations led to the loss of some spatial information. In this paper, we propose a context encoder network (CE-Net) to capture more high-level information and preserve spatial information for 2D medical image segmentation. CE-Net mainly contains three major components: a feature encoder module, a context extractor, and a feature decoder module. We use the pretrained ResNet block as the fixed feature extractor. The context extractor module is formed by a newly proposed dense atrous convolution block and a residual multi-kernel pooling block. We applied the proposed CE-Net to different 2D medical image segmentation tasks. Comprehensive results show that the proposed method outperforms the original U-Net method and other state-of-the-art methods for optic disc segmentation, vessel detection, lung segmentation, cell contour segmentation, and retinal optical coherence tomography layer segmentation. Zaiwang Gu, Jun Cheng 0003, Huazhu Fu, Kang Zhou 0001, Huaying Hao, Yitian Zhao, Tianyang Miller, Shenghua Gao, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Combining Multiple Deep Features for Glaucoma ClassificationabstractGlaucoma is one of the leading cause of blindness. Although there is still no cure, early detection can prevent serious vision loss. Therefore automated glaucoma detection/classification is an important issue. In the past decade, segmentation based approach such as those based on cup-to-disc-ratio are popular, but single indicator limit its performance. Recently, convolutional neural network based image classification approaches that can use more image cues achieve good performance. In this paper, we propose a new glaucoma classification by combining multiple features extracted by different convolutional neural networks. Its effectiveness is clearly demonstrated on the publicly available Origa [1] dataset. It achieves an area under the receiver operating characteristic curve of 0.8483, which better than the 0.838 given by on manual marked cup-to-disc-ratio. To our knowledge, it is the first approach surpass human in glaucoma classification. Annan Li, Yunhong Wang 0001, Jun Cheng 0003, Jiang Liu 0001 |
ICASSP | 3 |
| 2018 | Retinal Artery and Vein Classification via Dominant Sets Clustering-Based Vascular Topology Estimation
Yitian Zhao, Jianyang Xie, Pan Su 0001, Yalin Zheng, Yonghuai Liu, Jun Cheng 0003, Jiang Liu 0001 |
MICCAI (2) | 6 |
| 2018 | A Robust Outliers' Elimination Scheme for Multimodal Retina Image Registration Using Constrained Affine TransformationabstractThis paper proposes a robust outliers' elimination scheme for the registration of multimodal retina images. Our proposed constrained affine transformation Least Trimmed Squares (CAT-LTS) method has been designed to deal with image registration problems where the putatively matched feature points has a very large fraction of wrong matches. The constrained affine transformation allows all combinations of transformations such as scaling, rotation, translation and shear but disallows reflection. We use the Scale-Invariant Feature Transform (SIFT) feature points and Partial Intensity Invariant Feature Descriptors (PIIFD) to obtain the putatively matched feature points. We show that our proposed scheme when applied to the application of registering color fundus to enface optical coherence tomography (OCT) images significantly outperforms other outliers' elimination methods, namely the m-estimator sample and concensus (MSAC) and Random sample consensus (RANSAC) methods. Ee Ping Ong, Jun Cheng 0003, Damon Wing Kee Wong, Hwei Yee Teo, Leonard W. L. Yip |
TENCON | 2 |
| 2018 | Learning supervised descent directions for optic disc segmentation
Annan Li, Zhiheng Niu, Jun Cheng 0003, Fengshou Yin, Damon Wing Kee Wong, Shuicheng Yan, Jiang Liu 0001 |
Neurocomputing | 3 |
| 2018 | Sparse Range-Constrained Learning and Its Application for Medical Image GradingabstractSparse learning has been shown to be effective in solving many real-world problems. Finding sparse representations is a fundamentally important topic in many fields of science including signal processing, computer vision, genome study, and medical imaging. One important issue in applying sparse representation is to find the basis to represent the data, especially in computer vision and medical imaging where the data are not necessary incoherent. In medical imaging, clinicians often grade the severity or measure the risk score of a disease based on images. This process is referred to as medical image grading. Manual grading of the disease severity or risk score is often used. However, it is tedious, subjective, and expensive. Sparse learning has been used for automatic grading of medical images for different diseases. In the grading, we usually begin with one step to find a sparse representation of the testing image using a set of reference images or atoms from the dictionary. Then in the second step, the selected atoms are used as references to compute the grades of the testing images. Since the two steps are conducted sequentially, the objective function in the first step is not necessarily optimized for the second step. In this paper, we propose a novel sparse range-constrained learning (SRCL) algorithm for medical image grading. Different from most of existing sparse learning algorithms, SRCL integrates the objective of finding a sparse representation and that of grading the image into one function. It aims to find a sparse representation of the testing image based on atoms that are most similar in both the data or feature representation and the medical grading scores. We apply the new proposed SRCL to two different applications, namely, cup-to-disc ratio (CDR) computation from retinal fundus images and cataract grading from slit-lamp lens images. Experimental results show that the proposed method is able to improve the accuracy in CDR computation and cataract grading. Jun Cheng 0003 |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Structure-Preserving Guided Retinal Image Filtering and Its Application for Optic Disk AnalysisabstractRetinal fundus photographs have been used in the diagnosis of many ocular diseases such as glaucoma, pathological myopia, age-related macular degeneration, and diabetic retinopathy. With the development of computer science, computer aided diagnosis has been developed to process and analyze the retinal images automatically. One of the challenges in the analysis is that the quality of the retinal image is often degraded. For example, a cataract in human lens will attenuate the retinal image, just as a cloudy camera lens which reduces the quality of a photograph. It often obscures the details in the retinal images and posts challenges in retinal image processing and analyzing tasks. In this paper, we approximate the degradation of the retinal images as a combination of human-lens attenuation and scattering. A novel structure-preserving guided retinal image filtering (SGRIF) is then proposed to restore images based on the attenuation and scattering model. The proposed SGRIF consists of a step of global structure transferring and a step of global edge-preserving smoothing. Our results show that the proposed SGRIF method is able to improve the contrast of retinal images, measured by histogram flatness measure, histogram spread, and variability of local luminosity. In addition, we further explored the benefits of SGRIF for subsequent retinal image processing and analyzing tasks. In the two applications of deep learning-based optic cup segmentation and sparse learning-based cup-to-disk ratio (CDR) computation, our results show that we are able to achieve more accurate optic cup segmentation and CDR measurements from images processed by SGRIF. Jun Cheng 0003, Zhengguo Li, Zaiwang Gu, Huazhu Fu, Damon Wing Kee Wong, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Joint Optic Disc and Cup Segmentation Based on Multi-Label Deep Network and Polar TransformationabstractGlaucoma is a chronic eye disease that leads to irreversible vision loss. The cup to disc ratio (CDR) plays an important role in the screening and diagnosis of glaucoma. Thus, the accurate and automatic segmentation of optic disc (OD) and optic cup (OC) from fundus images is a fundamental task. Most existing methods segment them separately, and rely on hand-crafted visual feature from fundus images. In this paper, we propose a deep learning architecture, named M-Net, which solves the OD and OC segmentation jointly in a one-stage multi-label system. The proposed M-Net mainly consists of multi-scale input layer, U-shape convolutional network, side-output layer, and multi-label loss function. The multi-scale input layer constructs an image pyramid to achieve multiple level receptive field sizes. The U-shape convolutional network is employed as the main body network structure to learn the rich hierarchical representation, while the side-output layer acts as an early classifier that produces a companion local prediction map for different scale layers. Finally, a multi-label loss function is proposed to generate the final segmentation map. For improving the segmentation performance further, we also introduce the polar transformation, which provides the representation of the original image in the polar coordinate system. The experiments show that our M-Net system achieves state-of-the-art OD and OC segmentation result on ORIGA data set. Simultaneously, the proposed method also obtains the satisfactory glaucoma screening performances with calculated CDR value on both ORIGA and SCES datasets. Huazhu Fu, Jun Cheng 0003, Yanwu Xu 0001, Damon Wing Kee Wong, Jiang Liu 0001, Xiaochun Cao |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Disc-Aware Ensemble Network for Glaucoma Screening From Fundus ImageabstractGlaucoma is a chronic eye disease that leads to irreversible vision loss. Most of the existing automatic screening methods first segment the main structure and subsequently calculate the clinical measurement for the detection and screening of glaucoma. However, these measurement-based methods rely heavily on the segmentation accuracy and ignore various visual features. In this paper, we introduce a deep learning technique to gain additional image-relevant information and screen glaucoma from the fundus image directly. Specifically, a novel disc-aware ensemble network for automatic glaucoma screening is proposed, which integrates the deep hierarchical context of the global fundus image and the local optic disc region. Four deep streams on different levels and modules are, respectively, considered as global image stream, segmentation-guided network, local disc region stream, and disc polar transformation stream. Finally, the output probabilities of different streams are fused as the final screening result. The experiments on two glaucoma data sets (SCES and new SINDI data sets) show that our method outperforms other state-of-the-art algorithms. Huazhu Fu, Jun Cheng 0003, Yanwu Xu 0001, Changqing Zhang 0002, Damon Wing Kee Wong, Jiang Liu 0001, Xiaochun Cao |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Auto-flag the baseline for Mingantu Ultrawide Spectral Radioheliograph with LSTMabstractMingantu Ultrawide Spectral Radioheliograph (MUSER) is an aperture synthesis telescope consisting of a group of small antennas to image the Sun. Each two antennas form a baseline contributing a Fourier sampling point for each time of imaging. Both amplitude and phase of a baseline form a time sequence in a time interval. Normally, amplitude/phase should vary linearly in a short time interval. However, many reasons would result in errors of baselines, damaging image quality. This work makes the first attempt to auto-flag baselines so as to delete bad baselines and improve image quality. Inspired by the big success of long short-term memory (LSTM) for time series analysis, we believe that LSTM could accomplish auto-flagging (a binary classification task) of baselines even better by treating them as time sequences. Thus, LSTM is employed to learn the representation of the phase of each baseline for classification, where the interaction and connection within a phase sequence are explored, benefiting classification. The experimental results demonstrate that LSTM can well capture the characteristics of the phase of a baseline, and thus achieves better classification. Jun Cheng 0003, Long Xu 0001, Xuexin Yu, Linjie Chen, Wei Wang 0141, Yihua Yan |
VCIP | 1 |
| 2016 | Speckle Reduction in 3D Optical Coherence Tomography of Retina by A-Scan ReconstructionabstractOptical coherence tomography (OCT) is a micrometer-scale, cross-sectional imaging modality for biological tissue. It has been widely used for retinal imaging in ophthalmology. Speckle noise is problematic in OCT. A raw OCT image/volume usually has very poor image quality due to speckle noise, which often obscures the retinal structures. Overlapping scan is often used for speckle reduction in a 2D line-scan. However, it leads to an increase of the data acquisition time. Therefore, it is unpractical in 3D scan as it requires a much longer data acquisition time. In this paper, we propose a new method for speckle reduction in 3D OCT. The proposed method models each A -scan as the sum of underlying clean A -scan and noise. Based on the assumption that neighboring A -scans are highly similar in the retina, the method reconstructs each A -scan from its neighboring scans. In the method, the neighboring A -scans are aligned/registered to the A -scan to be reconstructed and form a matrix together. Then low rank matrix completion using bilateral random projection is utilized to iteratively estimate the noise and recover the underlying clean A -scan. The proposed method is evaluated through the mean square error, peak signal to noise ratio and the mean structure similarity index using high quality line-scan images as reference. Experimental results show that the proposed method performs better than other methods. In addition, the subsequent retinal layer segmentation also shows that the proposed method makes the automatic retinal layer segmentation more accurate. The technology can be embedded into current OCT machines to enhance the image quality for visualization and subsequent analysis such as retinal layer segmentation. Jun Cheng 0003, Dacheng Tao, Ying Quan, Damon Wing Kee Wong, Chui Ming Gemmy Cheung, Masahiro Akiba, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2015 | A low-dimensional step pattern analysis algorithm with application to multimodal retinal image registrationabstractExisting feature descriptor-based methods on retinal image registration are mainly based on scale-invariant feature transform (SIFT) or partial intensity invariant feature descriptor (PIIFD). While these descriptors are often being exploited, they do not work very well upon unhealthy multimodal images with severe diseases. Additionally, the descriptors demand high dimensionality to adequately represent the features of interest. The higher the dimensionality, the greater the consumption of resources (e.g. memory space). To this end, this paper introduces a novel registration algorithm coined low-dimensional step pattern analysis (LoSPA), tailored to achieve low dimensionality while providing sufficient distinctiveness to effectively align unhealthy multimodal image pairs. The algorithm locates hypotheses of robust corner features based on connecting edges from the edge maps, mainly formed by vascular junctions. This method is insensitive to intensity changes, and produces uniformly distributed features and high repeatability across the image domain. The algorithm continues with describing the corner features in a rotation invariant manner using step patterns. These customized step patterns are robust to non-linear intensity changes, which are well-suited for multimodal retinal image registration. Apart from its low dimensionality, the LoSPA algorithm achieves about two-fold higher success rate in multimodal registration on the dataset of severe retinal diseases when compared to the top score among state-of-the-art algorithms. Jimmy Addison Lee, Jun Cheng 0003, Beng Hai Lee, Ee Ping Ong, Guozhen Xu, Damon Wing Kee Wong, Jiang Liu 0001, Augustinus Laude, Tock Han Lim |
CVPR | 2 |
| 2015 | Registration of Color and OCT Fundus Images Using Low-dimensional Step Pattern Analysis
Jimmy Addison Lee, Jun Cheng 0003, Guozhen Xu, Ee Ping Ong, Beng Hai Lee, Damon Wing Kee Wong, Jiang Liu 0001 |
MICCAI (2) | 2 |
| 2015 | A Robust Outlier Elimination Approach for Multimodal Retina Image Registration
Ee Ping Ong, Jimmy Addison Lee, Jun Cheng 0003, Guozhen Xu, Beng Hai Lee, Augustinus Laude, Stephen Teoh, Tock Han Lim, Damon Wing Kee Wong, Jiang Liu 0001 |
MICCAI (2) | 3 |
| 2014 | Speckle Reduction in Optical Coherence Tomography by Image Registration and Matrix Completion
Jun Cheng 0003, Lixin Duan, Damon Wing Kee Wong, Dacheng Tao, Masahiro Akiba, Jiang Liu 0001 |
MICCAI (1) | 1 |
| 2013 | Superpixel Classification Based Optic Cup Segmentation
Jun Cheng 0003, Jiang Liu 0001, Dacheng Tao, Fengshou Yin, Damon Wing Kee Wong, Yanwu Xu 0001, Tien Yin Wong |
MICCAI (3) | 1 |
| 2013 | Research and applications: Automatic glaucoma diagnosis through medical imaging informaticsabstractBACKGROUND: Computer-aided diagnosis for screening utilizes computer-based analytical methodologies to process patient information. Glaucoma is the leading irreversible cause of blindness. Due to the lack of an effective and standard screening practice, more than 50% of the cases are undiagnosed, which prevents the early treatment of the disease. OBJECTIVE: To design an automatic glaucoma diagnosis architecture automatic glaucoma diagnosis through medical imaging informatics (AGLAIA-MII) that combines patient personal data, medical retinal fundus image, and patient's genome information for screening. MATERIALS AND METHODS: 2258 cases from a population study were used to evaluate the screening software. These cases were attributed with patient personal data, retinal images and quality controlled genome data. Utilizing the multiple kernel learning-based classifier, AGLAIA-MII, combined patient personal data, major image features, and important genome single nucleotide polymorphism (SNP) features. RESULTS AND DISCUSSION: Receiver operating characteristic curves were plotted to compare AGLAIA-MII's performance with classifiers using patient personal data, images, and genome SNP separately. AGLAIA-MII was able to achieve an area under curve value of 0.866, better than 0.551, 0.722 and 0.810 by the individual personal data, image and genome information components, respectively. AGLAIA-MII also demonstrated a substantial improvement over the current glaucoma screening approach based on intraocular pressure. CONCLUSIONS: AGLAIA-MII demonstrates for the first time the capability of integrating patients' personal data, medical retinal image and genome information for automatic glaucoma diagnosis and screening in a large dataset from a population study. It paves the way for a holistic approach for automatic objective glaucoma diagnosis and screening. Jiang Liu 0001, Zhuo Zhang 0001, Damon Wing Kee Wong, Yanwu Xu 0001, Fengshou Yin, Jun Cheng 0003, Ngan Meng Tan, Chee Keong Kwoh 0001, Dong Xu 0001, Tin Aung, Tien Yin Wong |
J. Am. Medical Informatics Assoc. | 6 |
| 2013 | Superpixel Classification Based Optic Disc and Optic Cup Segmentation for Glaucoma ScreeningabstractGlaucoma is a chronic eye disease that leads to vision loss. As it cannot be cured, detecting the disease in time is important. Current tests using intraocular pressure (IOP) are not sensitive enough for population based glaucoma screening. Optic nerve head assessment in retinal fundus images is both more promising and superior. This paper proposes optic disc and optic cup segmentation using superpixel classification for glaucoma screening. In optic disc segmentation, histograms, and center surround statistics are used to classify each superpixel as disc or non-disc. A self-assessment reliability score is computed to evaluate the quality of the automated optic disc segmentation. For optic cup segmentation, in addition to the histograms and center surround statistics, the location information is also included into the feature space to boost the performance. The proposed segmentation methods have been evaluated in a database of 650 images with optic disc and optic cup boundaries manually marked by trained professionals. Experimental results show an average overlapping error of 9.5% and 24.1% in optic disc and optic cup segmentation, respectively. The results also show an increase in overlapping error as the reliability score is reduced, which justifies the effectiveness of the self-assessment. The segmented optic disc and optic cup are then used to compute the cup to disc ratio for glaucoma screening. Our proposed method achieves areas under curve of 0.800 and 0.822 in two data sets, which is higher than other methods. The methods can be used for segmentation and glaucoma screening. The self-assessment will be used as an indicator of cases with large errors and enhance the clinical deployment of the automatic segmentation and screening. Jun Cheng 0003, Jiang Liu 0001, Yanwu Xu 0001, Fengshou Yin, Damon Wing Kee Wong, Ngan Meng Tan, Dacheng Tao, Ching Yu Cheng, Tin Aung, Tien Yin Wong |
IEEE Trans. Medical Imaging | 1 |
| 2012 | Superpixel Classification Based Optic Disc Segmentation
Jun Cheng 0003, Jiang Liu 0001, Yanwu Xu 0001, Fengshou Yin, Damon Wing Kee Wong, Ngan Meng Tan, Ching Yu Cheng, Tien Yin Wong |
ACCV (2) | 1 |
| 2012 | Early age-related macular degeneration detection by focal biologically inspired featureabstractAge-related macular degeneration (AMD) is a leading cause of vision loss. The presence of drusen are often associated to AMD. Drusen are tiny yellowish-white extracellular buildup present around the macular region of the retina. Clinically, ophthalmologists examine the area around the macula to determine the presence and severity of drusen. However, manual identification and recognition of drusen is subjective, time consuming and expensive. To reduce manual workload and facilitate large-scale early AMD screening, it is essential to detect drusen automatically. In this paper, we propose to use biologically inspired features (BIF) for the purpose of AMD detection. The optic disc and macula are detected to determine a focal region around macula for feature extraction. The extracted features are then classified using support vector machines (SVM). Our experimental results, tested on 350 images, demonstrate that the biologically inspired features from the focal region is effective for drusen detection with a sensitivity of 86.3% and specificity of 91.9%. The results of our proposed approach can be used to reduce workload of ophthalmologists and diagnosis cost. Jun Cheng 0003, Damon Wing Kee Wong, Xiangang Cheng, Jiang Liu 0001, Ngan Meng Tan, Mayuri Bhargava, Chui Ming Gemmy Cheung, Tien Yin Wong |
ICIP | 1 |
| 2012 | Peripapillary atrophy detection by biologically inspired feature
Jun Cheng 0003, Jiang Liu 0001, Damon Wing Kee Wong, Ngan Meng Tan, Carol Yim-lui Cheung, Mani Baskaran, Tien Yin Wong, Seang Mei Saw |
ICPR | 1 |
| 2012 | Efficient optic cup localization based on superpixel classification for glaucoma diagnosis in digital fundus images
Yanwu Xu 0001, Jiang Liu 0001, Jun Cheng 0003, Fengshou Yin, Ngan Meng Tan, Damon Wing Kee Wong, Ching Yu Cheng, Tien Yin Wong |
ICPR | 3 |
| 2012 | Peripapillary Atrophy Detection by Sparse Biologically Inspired Feature ManifoldabstractPeripapillary atrophy (PPA) is an atrophy of pre-existing retina tissue. Because of its association with eye diseases such as myopia and glaucoma, PPA is an important indicator for diagnosis of these diseases. Experienced ophthalmologists are able to determine the presence of PPA using visual information from the retinal images. However, it is tedious, time consuming and subjective to examine all images especially in a screening program. This paper presents biologically inspired feature (BIF) for the automatic detection of PPA. BIF mimics the process of cortex for visual perception. In the proposed method, a focal region is segmented from the retinal image and the BIF is extracted. As BIF is an intrinsically low dimensional feature embedded in a high dimensional space, it is not suitable to measure the similarity between two BIFs directly based on the Euclidean distance. Therefore, it is necessary to obtain a suitable mapping to reduce the dimensionality. In this paper, we explore sparse transfer learning to transfer the label information from ophthalmologists to the sample distribution knowledge contained in all samples. Selective pair-wise discriminant analysis is used to define two strategies of sparse transfer learning: negative and positive sparse transfer learning. Experimental results show that negative sparse transfer learning is superior to the positive one for this task. The proposed BIF based approach achieves an accuracy of more than 90% in detecting PPA, much better than previous methods. It can be used to save the workload of ophthalmologists and thus reduce the diagnosis costs. Jun Cheng 0003, Dacheng Tao, Jiang Liu 0001, Damon Wing Kee Wong, Ngan Meng Tan, Tien Yin Wong, Seang Mei Saw |
IEEE Trans. Medical Imaging | 1 |
| 2011 | Focal Biologically Inspired Feature for Glaucoma Type Classification
Jun Cheng 0003, Dacheng Tao, Jiang Liu 0001, Damon Wing Kee Wong, Beng Hai Lee, Mani Baskaran, Tien Yin Wong, Tin Aung |
MICCAI (3) | 1 |
| 2011 | Sliding Window and Regression Based Cup Detection in Digital Fundus Images for Glaucoma Diagnosis
Yanwu Xu 0001, Dong Xu 0001, Stephen Lin 0001, Jiang Liu 0001, Jun Cheng 0003, Carol Yim-lui Cheung, Tin Aung, Tien Yin Wong |
MICCAI (3) | 5 |
| 2009 | Steganalysis of halftone image using inverse halftoning
Jun Cheng 0003, Alex Chichung Kot |
Signal Process. | 1 |
| 2007 | Steganalysis of Binary Cartoon Image using Distortion MeasureabstractWe present a steganalysis technique for data hiding in binary cartoon images. Due to the perturbation from the embedding, the contours of a stego binary cartoon image are distorted. When calculating the distortion of the image based on a distortion measure, the distortion score between a stego image and its de-noised version should be different from that between an original image and its de-noised version. Different binary image distortion measures are used to calculate the distortion scores, which are used as the features for classification. The sequential floating forward search (SFFS) method is used to search for the combination of the features that yields the best classification results. Jun Cheng 0003, Alex Chichung Kot, Susanto Rahardja |
ICASSP (2) | 1 |
| 2007 | Objective Distortion Measure for Binary Text Image Based on Edge Line Segment SimilarityabstractThis paper proposes a new approach to measure the distortion introduced by changing individual edge pixels in binary text images. The approach considers not only how many pixels are changed but also where the pixels are changed and how the flipping affects the overall shape formed by the edge line. Similarities between the edge line segments in the original and distorted image are compared to measure the distortion. Subjective testing shows that the new distortion measure correlates well with human visual perception. Jun Cheng 0003, Alex Chichung Kot |
IEEE Trans. Image Process. | 1 |
| 2005 | Steganalysis of binary text imagesabstractWe present in this paper a technique for the steganalysis of electronic text documents. The proposed method averages the similar patterns together in a binary text image to estimate the original patterns. The proposed method can detect the existence of a secret message hidden by any boundary flipping techniques as well as estimate the message length and locate the flipped pixels. Jun Cheng 0003, Alex Chichung Kot, Jun Liu 0069 |
ICASSP (4) | 1 |
| 2005 | Detection of data hiding in binary text imagesabstractWe present in this paper a technique for the steganalysis of electronic binary text images. The proposed method utilizes the similarity between same characters or symbols. The proposed method can detect the existence of a secret message hidden by the embedding algorithms, which hide information by flipping centers of L-shape patterns (COL). Jun Cheng 0003, Alex Chichung Kot, Jun Liu 0069 |
ICIP (3) | 1 |