EDBT 2026 Demo / reviewers in the wild / expert
Dongyun Lin
dblp:85/10305
· DBLP profile ↗
28ranked-venue papers
13as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PEVA-Net: Prompt-enhanced view aggregation network for zero/few-shot multi-view 3D shape recognition
Dongyun Lin, Aiyuan Guo, Shangbo Mao |
Neurocomputing | 1 |
| 2025 | Multi-modality integrated class incremental learning networks for 3D object recognition
Dongyun Lin, Xiao Zhang 0053, Erli Meng, Zhiping Lin 0001, Huiping Zhuang |
Knowl. Based Syst. | 2 |
| 2024 | SCA-PVNet: Self-and-cross attention based aggregation of point cloud and multi-view for 3D object retrieval
Dongyun Lin, Aiyuan Guo, Shangbo Mao |
Knowl. Based Syst. | 1 |
| 2024 | Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question AnsweringabstractThe main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing graph-based methods for VideoQA usually ignore keywords in questions and employ a simple graph to aggregate features without considering relative relations between objects, which may lead to inferior performance. In this paper, we propose a Keyword-aware Relative Spatio-Temporal (KRST) graph network for VideoQA. First, to make question features aware of keywords, we employ an attention mechanism to assign high weights to keywords during question encoding. The keyword-aware question features are then used to guide video graph construction. Second, because relations are relative, we integrate the relative relation modeling to better capture the spatio-temporal dynamics among object nodes. Moreover, we disentangle the spatio-temporal reasoning into an object-level spatial graph and a frame-level temporal graph, which reduces the impact of spatial and temporal relation reasoning on each other. Extensive experiments on the TGIF-QA, MSVD-QA and MSRVTT-QA datasets demonstrate the superiority of our KRST over multiple state-of-the-art methods. Hehe Fan, Dongyun Lin, Ying Sun 0001, Mohan Kankanhalli, Joo-Hwee Lim |
IEEE Trans. Multim. | 3 |
| 2023 | Multi-Range View Aggregation Network With Vision Transformer Feature Fusion for 3D Object RetrievalabstractView-based methods have achieved state-of-the-art performance in 3D object retrieval. However, view-based methods still encounter two major challenges. The first is how to leverage the inter-view correlation to enhance view-level visual features. The second is how to effectively fuse view-level features into a discriminative global descriptor. Towards these two challenges, we propose a multi-range view aggregation network (MRVANet) with a vision transformer based feature fusion scheme for 3D object retrieval. Unlike the existing methods which only consider aggregating neighboring or adjacent views which could bring in redundant information, we propose a multi-range view aggregation module to enhance individual view representations through view aggregation beyond only neighboring views but also incorporate the views at different ranges. Furthermore, to generate the global descriptor from view-level features, we propose to employ the multi-head self-attention mechanism introduced by vision transformer to fuse the view-level features. Extensive experiments conducted on three public datasets including ModelNet40, ShapeNet Core55 and MCB-A demonstrate the superiority of the proposed network over the state-of-the-art methods in 3D object retrieval. Dongyun Lin, Shitala Prasad, Aiyuan Guo, Yanpeng Cao |
IEEE Trans. Multim. | 1 |
| 2022 | On the Use of Component Structural Characteristics for Voxel Segmentation in Semicon 3D ImagesabstractDetecting defects buried inside chips is critical for failure analysis in semiconductor manufacturing. In this paper, we perform 3D voxel segmentation on 2.5D semicon chips to locate and identify defects that may be present in them. We integrate tree based Ensemble method with the Cascaded Anisotropic Convolutional Neural Networks to employ component structural characteristics of semicon 3D object in voxel segmentation process. We fabricate custom 2.5D chips purposely creating defective regions by using a specific fabrication and assembling process. Thereafter, use commercial 3D XRM tools for 3D imaging of these chips. We perform accurate 3D Object localization for each 3D x-ray scan by using a slice and fuse approach. Then, we perform voxel segmentation on logic die (integral component of semicon chip) to detect Cu-pillar, solder, and void regions (if any). The results show that we achieve state-of-the-art voxel segmentation dice scores for all three sub-components. Tin Lay Nwe, Ramanpreet Singh Pahwa, Richard Chang 0002, Oo Zaw Min, Jie Wang 0042, Dongyun Lin, Shitala Prasad, Sheng Dong |
ICASSP | 7 |
| 2022 | MLSA-UNet: End-to-End Multi-Level Spatial Attention Guided UNet for Industrial Defect SegmentationabstractDefect segmentation from 2D images plays a critical role in industrial product quality assessment. In practice, it is common that there are sufficient normal (defect-free) images but a very limited number of anomalous (defective) images. The existing works proposed several UNet variants (e.g., CAM-UNet) by incorporating normal images into the training process to improve the defect segmentation performance. In this paper, we propose Multi-Level Spatial Attention UNet (MLSA-UNet) to address the industrial defect segmentation task. MLSA-UNet is trained in an end-to-end manner to simultaneously classify normal/anomalous images and segment out defective regions from anomalous images. The classification process is conducted by Spatial Attention Learning Module (SALM) to generate multi-level spatial attention maps which are exploited by Spatial Attention Guided Decoding Module (SADM) to provide the guidance in the decoding process of UNet. Extensive experiments on MVTec AD dataset demonstrate the superiority of the proposed MLSA-UNet over multiple state-of-the-art UNet variants on defect segmentation. Dongyun Lin, Shitala Prasad, Aiyuan Guo |
ICIP | 1 |
| 2022 | Masked Face Recognition via Self-Attention Based Local Consistency RegularizationabstractWith the COVID-19 pandemic, one critical measure against infection is wearing masks. This measure poses a huge challenge to the existing face recognition systems by introducing heavy occlusions. In this paper, we propose an effective masked face recognition system. To alleviate the challenge of mask occlusion, we first exploit RetinaFace to achieve robust masked face detection and alignment. Secondly, we propose a deep CNN network for masked face recognition trained by minimizing ArcFace loss together with a local consistency regularization (LCR) loss. This facilitates the network to simultaneously learn globally discriminative face representations of different identities together with locally consistent representations between the non-occluded faces and their counterparts wearing synthesized facial masks. The experiments on the masked LFW dataset demonstrate that the proposed system can produce superior masked face recognition performance over multiple state-of-the-art methods. The proposed method is implemented in a portable Jetson Nano device which can achieve real-time masked face recognition. Dongyun Lin, Shitala Prasad, Aiyuan Guo |
ICIP | 1 |
| 2022 | Implicit Shape Biased Few-Shot Learning for 3D Object GeneralizationabstractThe current state-of-the-art (SOTA) methods validate the role of shape in object categorization, however, except few, most of them neglect object shape information. Motivated by low-shot learning and increasing synthetic data in vision tasks, we investigated how image-based embedding generalization can be improved by the data itself. We propose a new data augmentation approach for low-shot object generalization regime based on image-only. The proposed method learns a discriminative embedding space using SIFT shape points for 3D objects, such that it’s easier to map images and point clouds into one. Numerous experiments show that the proposed approach is superior to the existing low-shot SOTA methods. Shitala Prasad, Dongyun Lin, Aiyuan Guo |
ICIP | 3 |
| 2022 | Multi-view 3D object retrieval leveraging the aggregation of view and instance attentive features
Dongyun Lin, Shitala Prasad, Tin Lay Nwe, Sheng Dong, Aiyuan Guo |
Knowl. Based Syst. | 1 |
| 2022 | A Progressive Multi-View Learning Approach for Multi-Loss Optimization in 3D Object Recognitionabstract3D object recognition is a well studied 2D multi-view object classification task that achieves high accuracy if the object textures are distinctive. However, if objects are texture-less and are only differentiable by their shapes but at certain viewpoints. Thus, the problem is still very challenging. Furthermore, the existing methods are mostly based on supervised learning with lots of images per object which are difficult to collect and label them for training. In this letter, we introduced a multi-loss view invariant stochastic prototype embedding to minimize and improve the recognition accuracy of novel objects at different viewpoints by using a progressive multi-view learning approach. An extensive experimental results show that the proposed method outperforms the state-of-the-art methods on different types datasets and also on different backbones. Shitala Prasad, Dongyun Lin, Sheng Dong, Tin Lay Nwe |
IEEE Signal Process. Lett. | 3 |
| 2021 | Action Relational Graph for Weakly-Supervised Temporal Action LocalizationabstractThe task of weakly-supervised temporal action localization (WTAL) is to recognize plentiful unstructured actions in untrimmed videos with only video-level class labels. As various actions may occur in an untrimmed video, it is desirable to capture the correlation among different actions to effectively identify the target actions. In this paper, we propose a novel Action Relational Graph Network (ARG-Net) to model the correlation between action labels. Specifically, we build a co-occurrence graph using Graph Convolutional Network (GCN), where the graph nodes and edges are represented by word embedding of action labels and relations between two labels, respectively. Then we apply the GCNs to project the action label embeddings into a set of correlated action classifiers which are multiplied with the learned video representations for video-level classification. To facilitate discriminative video representation learning, we employ the attention mechanism to model the probability of a frame containing action instances. A new Action Normalization Loss (ANL) is proposed to further alleviate the confusion from irrelevant background frames (i.e., frames containing no actions). Experimental results on THUMOS14 and ActivityNet1.2 datasets demonstrate that our ARG-Net outperforms the state-of-the-art methods. Ying Sun 0001, Dongyun Lin, Joo-Hwee Lim |
ICIP | 3 |
| 2021 | Cam-Guided U-Net With Adversarial Regularization For Defect SegmentationabstractDefect segmentation is critical in real-wold industrial product quality assessment. There are usually a huge number of normal (defect-free) images but a very limited number of annotated anomalous images. This poses huge challenges to exploiting Fully-Convolutional Networks (FCN), e.g., UNet, as they require sufficient anomalous images with defect annotations during training. To further leverage the information from normal data, a novel CAM-guided U-Net with adversarial regularization (CAM-UNet-AR) is proposed. We first modify the existing CAM-UNet to incorporate the CAMs for both normal and anomalous classes and fine-tune the segmentation network using a combined loss which jointly considers pixel-wise classification, foreground segmentation and boundary segmentation. Secondly, an auxiliary adversarial regularization module (ARM) is proposed to facilitate the segmentation network to encode the “normal components” from training images into consistent representations. Extensive experiments on MVTec AD dataset show the superiority of our proposed network over multiple state-of-the-art U-Net variants. Dongyun Lin, Shitala Prasad, Tin Lay Nwe, Sheng Dong, Oo Zaw Min |
ICIP | 1 |
| 2021 | Few-Shot Defect Segmentation Leveraging Abundant Defect-Free Training Samples Through Normal Background Regularization And Crop-And-Paste OperationabstractIn industrial quality assessment, it is challenging to conduct automated and accurate defect segmentation under the condition that abundant defect-free images but very limited anomalous images are available. This paper tackles the challenging few-shot defect segmentation task under such condition. We propose two regularization techniques via incorporating abundant defect-free images into the training of an encoder-decoder segmentation network. We first propose a Normal Background Regularization (NBR) loss which is jointly minimized with the segmentation loss, enhancing the encoder network to produce discriminative representations for normal regions. Secondly, we crop/paste defective regions to the randomly selected normal images for data augmentation and propose a weighted binary cross-entropy loss to enhance the training by emphasizing more realistic crop-and-pasted augmented images based on feature-level similarity comparison. Extensive experiments on MVTec AD and MTSD datasets demonstrate the superiority of the proposed method over the competing methods under few-shot settings. Dongyun Lin, Yanpeng Cao |
ICME | 1 |
| 2021 | maskedFaceNet: A Progressive Semi-Supervised Masked Face DetectorabstractTo reduce the risk of infecting or being infected by the recent COVID-19 virus, wearing mask is enforced or recommended by many countries. AI based system for automatically detecting whether individuals are wearing face mask becomes an urgent requirement in high risk facilities and crowded public places. Due to lacking of existing masked face datasets and the urgent low-cost application requirement, we propose a progressive semi-supervised learning method – called maskedFaceNet to minimize the efforts on data annotation and letting deep models to learn by using less annotated training data. With this method, the detection accuracy is further improved progressively while adapting to various application scenarios. Experimental results show that our maskedFaceNet is more efficient and accurate compared to other methods. Furthermore, we also contribute two masked face datasets for benchmarking and for the benefit of future research. Shitala Prasad, Dongyun Lin, Sheng Dong |
WACV | 3 |
| 2021 | CAM-guided Multi-Path Decoding U-Net with Triplet Feature Regularization for Defect Detection and Segmentation
Dongyun Lin, Shitala Prasad, Tin Lay Nwe, Sheng Dong, Oo Zaw Min |
Knowl. Based Syst. | 1 |
| 2020 | CAM-UNET: Class Activation MAP Guided UNET with Feedback Refinement for Defect SegmentationabstractThis paper tackles the task of defect segmentation by exploiting sufficient normal (defect-free) training images and limited annotated anomalous images. We propose a class activation map guided UNet (CAM-UNet) with feedback refinement mechanism for accurate defect segmentation. We first modify and pretrain the encoder of a VGG-16 backboned UNet to classify normal and anomalous training images. Then, for each of the anomalous training images, a CAM is generated as the prior segmentation information. Based on the CAM, we propose a feedback refinement process to train two decoder networks to progressively improve the segmentation output. Extensive experiments conducted on MVTEC AD dataset show that the proposed method significantly outperforms multiple benchmarking UNet methods in terms of mean IOU. Dongyun Lin, Shitala Prasad, Tin Lay Nwe, Sheng Dong, Oo Zaw Min |
ICIP | 1 |
| 2020 | Improving 3D Brain Tumor Segmentation With Predict-Refine Mechanism Using Saliency And Feature MapsabstractThis paper demonstrates the use of 3D Anisotropic Convolutional Neural Network (CNN) with predict-refine mechanism for 3D brain tumor segmentation. We propose two networks that utilize multi-scale feedback and saliency maps respectively to segment three critical regions involved in automated brain tumor segmentation. The proposed networks are formulated to predict feature maps at different resolutions during the prediction phase. These networks perform refinement process using the saliency or feature maps as feedback information for the refinement process. The recurrent architecture allows the network to automatically rectify errors in saliency map of the previous prediction phase resulting in more reliable final predictions. Our experimental results on the BraTS2017 dataset demonstrate the superior performance of our proposed predict-refine architecture than current state of the art approaches improving results by up to 8% without any additional increase in the 1.9M model parameters. Tin Lay Nwe, Oo Zaw Min, Saisubramaniam Gopalakrishnan, Dongyun Lin, Shitala Prasad, Sheng Dong, Ramanpreet Singh Pahwa |
ICIP | 4 |
| 2020 | Rethinking of Deep Models Parameters with Respect to Data DistributionabstractThe performance of deep learning models are driven by various parameters but to tune all of them every time, for every given dataset, is a heuristic practice. In this paper, unlike the common practice of decaying the learning rate, we propose a step-wise training strategy where the learning rate and the batch size are tuned based on the dataset size. Here, the given dataset size is progressively increased during the training to boost the network performance without saturating the learning curve, which is seen after certain epochs. We conducted extensive experiments on multiple networks and datasets to validate the proposed training strategy. The experimental results proves our hypothesis that the learning rate, the batch size and the data size are interrelated and can improve the network accuracy if an optimal progressive step-wise training strategy is applied. The proposed strategy also reduces the overall training cost compared to the baseline approach. Shitala Prasad, Dongyun Lin, Sheng Dong, Oo Zaw Min |
ICPR | 2 |
| 2020 | Global dissipativity of delayed discrete-time inertial neural networks
Dongyun Lin, Weiyao Lan |
Neurocomputing | 2 |
| 2020 | Passivity Analysis of Non-autonomous Discrete-Time Inertial Neural Networks with Time-Varying Delays
Dongyun Lin |
Neural Process. Lett. | 2 |
| 2020 | RefineU-Net: Improved U-Net with progressive global feedbacks and residual attention guided local refinement for medical image segmentation
Dongyun Lin, Tin Lay Nwe, Sheng Dong, Oo Zaw Min |
Pattern Recognit. Lett. | 1 |
| 2019 | Discriminative Features for Incremental Learning ClassifierabstractAn important problem in artificial intelligence is to develop an efficient system that can adapt to new knowledge in an incremental manner without forgetting previously learned knowledge. Although Convolutional Neural Networks (CNNs) are good at learning strong classifier and discriminative features, CNNs can not perform well in incremental classifier learning due to the catastrophic forgetting problem in the retraining process. In this paper, we propose a novel yet extremely simple approach to enhance the discriminative property of features for incremental classifier learning. We build a network for the universal feature space in which a group of image classes have intra-class compactness and inter-class separability. And, we model each incremental class to have a maximum margin from the rest of the models in universal space. Experiments are conducted on CIFAR-100 dataset and IMage Database for Context Aware Advertisement (IMDB-CAA) we collected. The results demonstrate the superiority of our approach, improving performance on CIFAR-100 dataset over state-of-the-art incremental learning systems. Furthermore, experiments on few-short incremental learning setting show very promising performance although we use only 4% of training samples on CIFAR-100 dataset. Tin Lay Nwe, Balaji Nataraj, Shudong Xie, Dongyun Lin, Sheng Dong |
ICIP | 5 |
| 2018 | Linear Active Disturbance Rejection Control For Double Integrator: Separation Diagram And Frequency PropertiesabstractThis paper investigates to control the double integrator with linear active disturbance rejection control and reports two results. First, there is a separation principle which means the dynamics of the state feedback and the error of extended state observer are completely decoupled. Second, with bandwidth method in which the observer bandwidth equals to the controller bandwidth, the closed-loop always has a phase margin of 31.89 degrees and a gain crossover frequency equals to the observer bandwidth, which can be set arbitrary. Huiyu Jin, Jingchao Song, Song Zeng, Dongyun Lin, Weiyao Lan |
ICARCV | 4 |
| 2018 | Deep CNNs for microscopic image classification by exploiting transfer learning and feature concatenationabstractDeep convolutional neural networks (CNNs) have become one of the state-of-the-art methods for image classification in various domains. For biomedical image classification where the number of training images is generally limited, transfer learning using CNNs is often applied. Such technique extracts generic image features from nature image datasets and these features can be directly adopted for feature extraction in smaller datasets. In this paper, we propose a novel deep neural network architecture based on transfer learning for microscopic image classification. In our proposed network, we concatenate the features extracted from three pretrained deep CNNs. The concatenated features are then used to train two fully-connected layers to perform classification. In the experiments on both the 2D-Hela and the PAP-smear datasets, our proposed network architecture produces significant performance gains comparing to the neural network structure that uses only features extracted from single CNN and several traditional classification methods. Long D. Nguyen, Dongyun Lin, Zhiping Lin 0001, Jiuwen Cao |
ISCAS | 2 |
| 2017 | LLC encoded BoW features and softmax regression for microscopic image classificationabstractThis paper proposes a method based on the bag-of-words (BoW) and the softmax regression for microscopic image classification. Essentially, the locality-constrained linear coding (LLC) is adopted for local feature encoding. Compared with the traditionally adopted vector quantization (VQ) in the BoW framework, the LLC encodes local structures of microscopic images with lower quantization errors and generates a sparse image representation. This enables the use of linear classifiers with low computational complexity. A softmax regression classifier is then adopted to address the multi-categorical classification task where the confidence of categorical prediction is quantified by posterior probabilities. Compared with other linear classifiers (such as the linear SVM) which only assign labels to images, such probabilistic outputs provide extra quantitative information to analyze misclassified images. Our experiments on the 2D-Hela and the PAP smear data sets show significant performance improvement of the proposed method comparing with competing methods using different features and classifiers under the BoW framework. Dongyun Lin, Zhiping Lin 0001, Lei Sun 0006, Kar-Ann Toh, Jiuwen Cao |
ISCAS | 1 |
| 2017 | Automatic endosomal structure detection and localization in fluorescence microscopic imagesabstractThis paper proposes a modified spatially-constrained similarity measure (mSCSM) method for endosomal structure detection and localization under the bag-of-words (BoW) framework. To our best knowledge, the proposed mSCSM is the first method for fully automatic detection and localization of complex subcellular compartments like endosomes. Essentially, a new similarity score and a novel two-stage output control scheme are proposed for localization by extracting discriminative information within a group of query images. Compared with the original SCSM which is formulated for instance localization, the proposed mSCSM can address category based localization problems. The preliminary experimental results show the proposed mSCSM can correctly detect and localize 79.17% of the existing endosomal structures in the microscopic images of human myeloid endothelial cells. Dongyun Lin, Zhiping Lin 0001, Ramraj Velmurugan, Raimund J. Ober |
ISCAS | 1 |
| 2016 | Reaction-diffusion based level set method with local entropy thresholding for melasma image segmentationabstractThis paper proposes a new method for melasma pigmentary area segmentation utilizing re action-diffusion based level set model (RDLSM) together with local entropy thresholding. In the adopted level set model, a diffusion term is used to regularize the level set function while a reaction term with anticipated sign property is used to force the zero level set towards desired locations. Then local entropy thresholding is applied to address the over-segmentation issue of RDLSM and to extract desired boundaries with higher overall local entropy. As a result, the melasma pigmentary areas and the normal skin areas can be better identified. Experimental results show that the proposed method performs well for melasma image segmentation, especially for cases with severe non-uniform illumination distribution. Yunfeng Liang, Dongyun Lin, Zhiping Lin 0001, Steven Tien Guan Thng, Emily Yiping Gan, Evelyn Yuxin Tay |
ICARCV | 3 |