EDBT 2026 Demo / reviewers in the wild / expert
Xiaoqing Zhang 0001
dblp:22/7627-1 · also Xiao-Qing Zhang 0001
· DBLP profile ↗
29ranked-venue papers
12as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive wavelet filters as practical texture amplifiers for early Parkinson's disease screening from retinal pathology perspectiveabstractParkinson’s disease (PD) is a prevalent neurodegenerative disorder globally. The eye’s retina is an extension of the brain, and clinical evidence has suggested the great potential of retinal pathology as surrogate biomarkers for early PD diagnosis. In particular, recent studies have shown that texture features extracted from retinal layers based on optical coherence tomography (OCT) images are strongly associated with PD-related retinal pathology. Additionally, frequency domain learning techniques can improve the representational capabilities of deep neural networks (DNNs) by decomposing frequency components that involve rich texture features, which remain underexplored for automated early PD diagnosis in OCT. To bridge this gap, we propose an Adaptive Wavelet Filter (AWF) that serves as the Practical Texture Amplifier, which fully leverages the merits of texture features from the retinal pathology view. Specifically, AWF first enhances feature map diversities and refines feature representations via channel mixer, then emphasizes informative texture feature representations with the well-designed adaptive wavelet filtering token mixer with the aid of frequency domain learning. By embedding AWFs into the DNN stem, AWFNet is constructed for automated early PD screening from OCT images. Additionally, we introduce a novel Balanced Confidence (BC) loss to boost early PD screening performance and trustworthiness of AWFNet, by mining the potential of sample-wise predicted probabilities across all classes and class frequency prior. The extensive experiments manifest the superiority of AWFNet with BC over state-of-the-art methods in terms of early PD screening performance and trustworthiness. Xiaoqing Zhang 0001, Hanfeng Shi, Haili Ye, Tao Xu 0031, Jiang Liu 0001 |
Expert Syst. Appl. | 1 |
| 2026 | Token pyramid pooling-driven style adapter learning with dual-view balanced loss for imbalanced diabetic retinopathy grading
Jilu Zhao, Xiaoqing Zhang 0001, Hanxi Sun, Qiushi Nie, Zunjie Xiao, Linxia Xiao, Fengyun Zhang, Jiang Liu 0001 |
Pattern Recognit. | 2 |
| 2026 | Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields in Efficient CNNs for Fair Medical Image ClassificationabstractEfficient convolutional neural network (CNN) architecture design has attracted growing research interests. However, they typically apply single receptive field (RF), small asymmetric RFs, or pyramid RFs to learn different feature representations, still encountering two significant challenges in medical image classification tasks: i) They have limitations in capturing diverse lesion characteristics efficiently, e.g., tiny, coordination, small and salient, which have unique roles on the classification results, especially imbalanced medical image classification. ii) The predictions generated by those CNNs are often unfair/biased, bringing a high risk when employing them to real-world medical diagnosis conditions. To tackle these issues, we develop a new concept, Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields (ERoHPRF), to simultaneously boost medical image classification performance and fairness. This concept aims to mimic the multi-expert consultation mode by applying the well-designed heterogeneous pyramid RF bag to capture lesion characteristics with varying significances effectively via convolution operations with multiple heterogeneous kernel sizes. Additionally, ERoHPRF introduces an expert-like structural reparameterization technique to merge its parameters with the two-stage strategy, ensuring competitive computation cost and inference speed through comparisons to a single RF. To manifest the effectiveness and generalization ability of ERoHPRF, we incorporate it into mainstream efficient CNN architectures. The extensive experiments show that our proposed ERoHPRF maintains a better trade-off than state-of-the-art methods in terms of medical image classification, fairness, and computation overhead. The code of this paper is available at https://github.com/XiaoLing12138/Expert-Like-Reparameterization-of-Heterogeneous-Pyramid-Receptive-Fields. Xiaoqing Zhang 0001, Zunjie Xiao, Lingxi Hu, Risa Higashita, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Structural uncertainty estimation for medical image segmentation
Xiaoqing Zhang 0001, Huihong Zhang, Sanqian Li, Risa Higashita, Jiang Liu 0001 |
Medical Image Anal. | 2 |
| 2025 | VSR-Net: Vessel-Like Structure Rehabilitation Network With Graph ClusteringabstractThe morphologies of vessel-like structures, such as blood vessels and nerve fibres, play significant roles in disease diagnosis, e.g., Parkinson's disease. Although deep network-based refinement segmentation and topology-preserving segmentation methods recently have achieved promising results in segmenting vessel-like structures, they still face two challenges: 1) existing methods often have limitations in rehabilitating subsection ruptures in segmented vessel-like structures; 2) they are typically overconfident in predicted segmentation results. To tackle these two challenges, this paper attempts to leverage the potential of spatial interconnection relationships among subsection ruptures from the structure rehabilitation perspective. Based on this perspective, we propose a novel Vessel-like Structure Rehabilitation Network (VSR-Net) to both rehabilitate subsection ruptures and improve the model calibration based on coarse vessel-like structure segmentation results. VSR-Net first constructs subsection rupture clusters via a Curvilinear Clustering Module (CCM). Then, the well-designed Curvilinear Merging Module (CMM) is applied to rehabilitate the subsection ruptures to obtain the refined vessel-like structures. Extensive experiments on six 2D/3D medical image datasets show that VSR-Net significantly outperforms state-of-the-art (SOTA) refinement segmentation methods with lower calibration errors. Additionally, we provide quantitative analysis to explain the morphological difference between the VSR-Net's rehabilitation results and ground truth (GT), which are smaller compared to those between SOTA methods and GT, demonstrating that our method more effectively rehabilitates vessel-like structures. Haili Ye, Xiaoqing Zhang 0001, Huazhu Fu, Jiang Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Adaptive Dual-Axis Style-Based Recalibration Network With Class-Wise Statistics Loss for Imbalanced Medical Image ClassificationabstractSalient and small lesions (e.g., microaneurysms on fundus) both play significant roles in real-world disease diagnosis under medical image examinations. Although deep neural networks (DNNs) have achieved promising medical image classification performance, they often have limitations in capturing both salient and small lesion information, restricting performance improvement in imbalanced medical image classification. Recently, with the advent of DNN-based style transfer in medical image generation, the roles of clinical styles have attracted great interest, as they are crucial indicators of lesions. Motivated by this observation, we propose a novel Adaptive Dual-Axis Style-based Recalibration (ADSR) module, leveraging the potential of clinical styles to guide DNNs in effectively learning salient and small lesion information from a dual-axis perspective. ADSR first emphasizes salient lesion information via global style-based adaptation, then captures small lesion information with pixel-wise style-based fusion. We construct an ADSR-Net for imbalanced medical image classification by stacking multiple ADSR modules. Additionally, DNNs typically adopt cross-entropy loss for parameter optimization, which ignores the impacts of class-wise predicted probability distributions. To address this, we introduce a new Class-wise Statistics Loss (CWS) combined with CE to further boost imbalanced medical image classification results. Extensive experiments on five imbalanced medical image datasets demonstrate not only the superiority of ADSR-Net and CWS over state-of-the-art (SOTA) methods but also their improved confidence calibration results. For example, ADSR-Net with the proposed loss significantly outperforms CABNet50 by 21.39% and 27.82% in F1 and B-ACC while reducing 3.31% and 4.57% in ECE and BS on ISIC2018. Xiaoqing Zhang 0001, Zunjie Xiao, Jingzhe Ma, Jilu Zhao, Shuai Zhang 0029, Runzhi Li, Yi Pan 0001, Jiang Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | GlanceSeg: Real-Time Microaneurysm Lesion Segmentation With Gaze-Map-Guided Foundation Model for Early Detection of Diabetic RetinopathyabstractEarly-stage diabetic retinopathy (DR) presents challenges in clinical diagnosis due to inconspicuous and minute microaneurysms (MAs), resulting in limited research in this area. Additionally, the potential of emerging foundation models, such as the segment anything model (SAM), in medical scenarios remains rarely explored. In this work, we propose a human-in-the-loop, label-free early DR diagnosis framework called GlanceSeg, based on SAM. GlanceSeg enables real-time segmentation of MA lesions as ophthalmologists review fundus images. Our human-in-the-loop framework integrates the ophthalmologist's gaze maps, allowing for rough localization of minute lesions in fundus images. Subsequently, a saliency map is generated based on the located region of interest, which provides prompt points to assist the foundation model in efficiently segmenting MAs. Finally, a domain knowledge filtering (DKF) module refines the segmentation of minute lesions. We conducted experiments on two newly-built public datasets, i.e., IDRiD and Retinal-Lesions, and validated the feasibility and superiority of GlanceSeg through visualized illustrations and quantitative measures. Additionally, we demonstrated that GlanceSeg improves annotation efficiency for clinicians and further enhances segmentation performance through fine-tuning using annotations. The clinician-friendly GlanceSeg is able to segment small lesions in real-time, showing potential for clinical applications. Hongyang Jiang 0001, Mengdi Gao, Zirong Liu, Xiaoqing Zhang 0001, Wu Yuan 0001, Jiang Liu 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Pyramid Pixel Context Adaption Network for Medical Image Classification With Supervised Contrastive LearningabstractSpatial attention (SA) mechanism has been widely incorporated into deep neural networks (DNNs), significantly lifting the performance in computer vision tasks via long-range dependency modeling. However, it may perform poorly in medical image analysis. Unfortunately, the existing efforts are often unaware that long-range dependency modeling has limitations in highlighting subtle lesion regions. To overcome this limitation, we propose a practical yet lightweight architectural unit, pyramid pixel context adaption (PPCA) module, which exploits multiscale pixel context information to recalibrate pixel position in a pixel-independent manner dynamically. PPCA first applies a well-designed cross-channel pyramid pooling (CCPP) to aggregate multiscale pixel context information, then eliminates the inconsistency among them by the well-designed pixel normalization (PN), and finally estimates per pixel attention weight via a pixel context integration. By embedding PPCA into a DNN with negligible overhead, the PPCA network (PPCANet) is developed for medical image classification. In addition, we introduce supervised contrastive learning to enhance feature representation by exploiting the potential of label information via supervised contrastive loss (CL). The extensive experiments on six medical image datasets show that the PPCANet outperforms state-of-the-art (SOTA) attention-based networks and recent DNNs. We also provide visual analysis and ablation study to explain the behavior of PPCANet in the decision-making process. Xiaoqing Zhang 0001, Zunjie Xiao, Yanlin Chen 0004, Jilu Zhao, Jiang Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | WaveFormer: A Wavelet Transformer for Parkinson Disease's Retinal Layer Segmentation in OCTabstractPathology symptoms of Parkinson disease (PD) are different from those of retinal diseases in the retinal layers, which are subtle. However, segmenting pathology information of PD from retinal layers automatically based on optical coherence tomography (OCT) images has not been studied before. Although existing Transformer-based segmentation methods have achieved good segmentation results, they have limitations in capturing local context information. Convolutional neural networks (CNNs) can construct local context dependencies among pixels, which is complementary to Transformers. Particularly, edge information extraction is significant for accurate retinal layer segmentation, which is ignored by both Transformers and CNNs but can be captured by frequency domain learning methods. To fully leverage the advantages of Transformers, CNNs, and frequency domain learning methods, we propose a Wavelet Transformer (WaveFormer) for retinal layer segmentation based on OCT images. In the WaveFormer, we design a Wavelet Spatial Attention block to exploit the potential of frequency information. Based on these advantages, WaveFormer can be data-efficient in limited OCT images of PD. The experimental results on the OCT-PD segmentation dataset show that our WaveFormer outperforms existing Transformers and CNNs. For example, WaveFormer outperforms Swin-UNet by 3.41% of IoU. Yanlin Chen 0004, Xiaoqing Zhang 0001, Tianao Wang, Haili Ye, Jiang Liu 0001 |
IJCNN | 2 |
| 2024 | Efficient pyramid channel attention network for pathological myopia recognition with pretraining-and-finetuning
Xiaoqing Zhang 0001, Jilu Zhao, Xiangtian Zhou, Jiang Liu 0001 |
Artif. Intell. Medicine | 1 |
| 2024 | DCAMIL: Eye-tracking guided dual-cross-attention multi-instance learning for refining fundus disease detection
Hongyang Jiang 0001, Mengdi Gao, Jingqi Huang, Xiaoqing Zhang 0001, Jiang Liu 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Regional context-based recalibration network for cataract recognition in AS-OCT
Xiaoqing Zhang 0001, Zunjie Xiao, Risa Higashita, Jiang Liu 0001 |
Pattern Recognit. | 1 |
| 2024 | An Identity-Preserved Framework for Human Motion TransferabstractHuman motion transfer (HMT) aims to generate a video clip for the target subject by imitating the source subject’s motion. Although previous methods have achieved good results in synthesizing good-quality videos, they lose sight of individualized motion information from the source and target motions, which is significant for the realism of the motion in the generated video. To address this problem, we propose a novel identity-preserved HMT network, termedIDPres. This network is a skeleton-based approach that uniquely incorporates the target’s individualized motion and skeleton information to augment identity representations. This integration significantly enhances the realism of movements in the generated videos. Our method focuses on the fine-grained disentanglement and synthesis of motion. To improve the representation learning capability in latent space and facilitate the training ofIDPres, we introduce three training schemes. These schemes enableIDPresto concurrently disentangle different representations and accurately control them, ensuring the synthesis of ideal motions. To evaluate the proportion of individualized motion information in the generated video, we are the first to introduce a new quantitative metric called Identity Score (ID-Score), motivated by the success of gait recognition methods in capturing identity information. Moreover, we collect an identity-motion paired dataset,Dancer101, consisting of solo-dance videos of 101 subjects from the public domain, providing a benchmark to prompt the development of HMT methods. Extensive experiments demonstrate that the proposedIDPresmethod surpasses existing state-of-the-art techniques in terms of reconstruction accuracy, realistic motion, and identity preservation. Jingzhe Ma, Xiaoqing Zhang 0001, Shiqi Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Prior-SSL: A Thickness Distribution Prior and Uncertainty Guided Semi-supervised Learning Method for Choroidal Segmentation in OCT Images
Huihong Zhang, Xiaoqing Zhang 0001, Yinlin Zhang, Risa Higashita, Jiang Liu 0001 |
ICANN (2) | 2 |
| 2023 | Oct Image Blind Despeckling Based on Gradient Guided Filter with Speckle Statistical PriorabstractOptical coherence tomography (OCT) imaging technique has been widely used for ocular disease diagnosis. However, speckles occur in OCT images due to the property of coherent imaging, inevitably affecting the visual quality and clinical analysis. To alleviate this problem, we propose a novel gradient-guided speckle image filtering method (GGSF) with structure enhancement for directly removing speckles in OCT images. Specifically, the multiplicative characteristic of speckle noise is incorporated into the guided filtering processing for modeling raw OCT images. To avoid getting trapped in image distortions, we further employ gradient regularization to integrate the structure prior information into the guided speckle image filtering procedure. Additionally, we introduce the statistical property of speckle noise obeying a gamma distribution into the least square method solver for the resulting non-convex GGSF model. Experimental results on the AS-OCT dataset demonstrate the effectiveness of GGSF for OCT image despeckling compared with competitive methods. Furthermore, we validate the benefits of GGSF for subsequent clinical analysis with the CM-OCT dataset. Sanqian Li, Muxing Xiong, Xiaoqing Zhang 0001, Risa Higashita, Jiang Liu 0001 |
ICASSP | 4 |
| 2023 | DMINet: A lightweight dual-mixed channel-independent network for cataract recognitionabstractCataracts are the leading cause of visual impairment and blindness globally attracting abroad attention from society. Over the years researchers have developed many state-of-the-art convolutional neural networks (CNNs) to recognize cataract severity levels based on different ophthalmic images. However most current works focus on improving cataract recognition performance by designing complex CNNs often ignoring resource-constrained medical device limitations. To this problem this paper proposes a novel dual-mixed channel-independent convolution (DMIConv) method which takes advantage of the multiscale convolution kernels by combining a depthwise convolution with a depthwise dilated convolution sequentially. Moreover we build a lightweight dual-mixed channel-independent network (DMINet) to recognize cataracts. To verify the effectiveness and efficiency of DMINet we conduct extensive experiments on a clinical anterior segment optical coherence tomography (AS-OCT) dataset of nuclear cataract (NC) and a publicly available OCT dataset. The results show that our proposed DMINet keeps a better tradeoff between the model complexity and the classification performance than efficient CNNs e.g DMINet outperforms MixNet by 3.34% of accuracy by using 4.58 % fewer parameters Qiuyang Yan, Jilu Zhao, Xiaoqing Zhang 0001, Risa Higashita, Jiang Liu 0001 |
IJCNN | 6 |
| 2023 | HA-Net: Hierarchical Attention Network Based on Multi-Task Learning for Ciliary Muscle Segmentation in AS-OCTabstractCiliary muscle segmentation in Anterior Segment Optical Coherence Tomography (AS-OCT) images is critical significance, yet challenging due to ambiguous boundaries. In this paper, we propose a hierarchical attention multi-task network, HA-Net, based on U-Net for ciliary muscle segmentation using AS-OCT images. The network comprises a primary task for ciliary muscle segmentation and two auxiliary tasks for signed distance map regression and key point localization. The signed distance map is employed to incorporate shape priors into the model and delineate the ciliary muscle boundary, while key point localization guides the model to focus on ambiguous regions. Notably, in contrast to the widely-used multi-task model that generates results in parallel, we introduce a hierarchical attention module to exploit the affiliation prior of three tasks for generating outputs serially. Experimental results on CM544 dataset demonstrate that HA-Net outperforms state-of-the-art methods in ciliary muscle segmentation, with 0.9178 Dice score and 7.11 pixels HD95. Additionally, as a by-product of the multi-task model, key point localization facilitates the measurement of ciliary muscle thickness in clinical analysis. Xiaoqing Zhang 0001, Sanqian Li, Risa Higashita, Jiang Liu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Channel-Wise and Spatial Feature Recalibration Network for Nuclear Cataract ClassificationabstractNuclear cataract (NC) is a prior age-related disease for blindness and vision impairment globally. Anterior segment optical coherence tomography (AS-OCT) image is a new ophthalmology image, which can capture the lens nucleus region clearly compared with other ophthalmic images, e.g., slit lamp images. Clinical research has suggested that features e.g., mean from AS-OCT images have varying correlations with NC severity levels. However, existing convolutional neural network (CNN) based NC classification works have not incorporated the clinical features into the network design to improve the performance. To this end, we propose a novel channel-wise and spatial feature recalibration network (CSFR-Net) to predict NC severity levels automatically, which is built on a stack of channel-wise and spatial feature recalibration (CSFR) modules. In each CSFR module, we construct a channel-wise feature recalibration block and a spatial feature recalibration block to recalibrate intermediate feature maps dynamically. This feature recalibration strategy enables CSFR-Net to highlight feature representations and suppress unnecessary ones in a global-and-local manner. We conduct extensive experiments on a clinical AS-OCT image dataset and CIFAR benchmarks. The results show that our CSFR-Net achieves better performance than state-of-the-art methods with less model complexity. Xiaoqing Zhang 0001, Gelei Xu, Junyong Shen, Zunjie Xiao, Qiuyang Yan, Risa Higashita, Jiang Liu 0001 |
ICME | 1 |
| 2022 | Interaction-Oriented Feature Decomposition for Medical Image Lesion Detection
Junyong Shen, Xiaoqing Zhang 0001, Zhongxi Qiu, Tingming Deng, Yanwu Xu 0001, Jiang Liu 0001 |
MICCAI (3) | 3 |
| 2022 | A Novel Local-Global Spatial Attention Network for Cortical Cataract Classification in AS-OCT
Zunjie Xiao, Xiaoqing Zhang 0001, Qingyang Sun, Zhuofei Wei, Gelei Xu, Risa Higashita, Jiang Liu 0001 |
PRCV (2) | 2 |
| 2022 | Adaptive feature squeeze network for nuclear cataract classification in AS-OCT image
Xiaoqing Zhang 0001, Zunjie Xiao, Risa Higashita, Jiang Liu 0001 |
J. Biomed. Informatics | 1 |
| 2022 | RANet: Network intrusion detection with group-gating convolutional neural network
Xiaoqing Zhang 0001, Zhao Tian 0005, Wei Liu 0043, Yifa Li, Wei She |
J. Netw. Comput. Appl. | 1 |
| 2022 | CCA-Net: Clinical-awareness attention network for nuclear cataract classification in AS-OCT
Xiaoqing Zhang 0001, Zunjie Xiao, Lingxi Hu, Gelei Xu, Risa Higashita, Jiang Liu 0001 |
Knowl. Based Syst. | 1 |
| 2022 | Attention to region: Region-based integration-and-recalibration networks for nuclear cataract classification using AS-OCT imagesabstractNuclear cataract (NC) is a leading eye disease for blindness and vision impairment globally. Accurate and objective NC grading/classification is essential for clinically early intervention and cataract surgery planning. Anterior segment optical coherence tomography (AS-OCT) images are capable of capturing the nucleus region clearly and measuring the opacity of NC quantitatively. Recently, clinical research has suggested that the opacity correlation and repeatability between NC severity levels and the average nucleus density on AS-OCT images is high with the interclass and intraclass analysis. Moreover, clinical research has suggested that opacity distribution is uneven on the nucleus region, indicating that the opacities from different nucleus regions may play different roles in NC diagnosis. Motivated by the clinical priors, this paper proposes a simple yet effective region-based integration-and-recalibration attention (RIR), which integrates multiple feature map region representations and recalibrates the weights of each region via softmax attention adaptively. This region recalibration strategy enables the network to focus on high contribution region representations and suppress less useful ones. We combine the RIR block with the residual block to form a Residual-RIR module, and then a sequence of Residual-RIR modules are stacked to a deep network named region-based integration-and-recalibration network (RIR-Net), to predict NC severity levels automatically. The experiments on a clinical AS-OCT image dataset and two OCT datasets demonstrate that our method outperforms strong baselines and previous state-of-the-art methods. Furthermore, attention weight visualization analysis and ablation studies verify the capability of our RIR-Net for adjusting the relative importance of different regions in feature maps dynamically, agreeing with the clinical research. Xiaoqing Zhang 0001, Zunjie Xiao, Huazhu Fu, Yanwu Xu 0001, Risa Higashita, Jiang Liu 0001 |
Medical Image Anal. | 1 |
| 2021 | Multimedia Meets Archaeology: A Novel Interdisciplinary Teaching ApproachabstractMultimedia information processing course includes image processing, text processing, video processing, audio processing, graphics, and animation. Classical multimedia information processing course is to lecture these contents as independent course units, making the course teaching inconsistently and student's learning interest and attention lost easily. Archaeology is a cross-disciplinary field, where an archaeological research project needs to use a variety of multimedia information processing technologies. This paper introduces a novel cross-disciplinary course teaching approach to combine traditional multimedia information processing techniques with archaeological research content in the new engineering era, which increases students' attention and interest in the multimedia information processing course. Teaming up with the archaeology professor, we present two multimedia-archaeology projects: intelligent pottery fragments splicing and intelligent Oracle inscription recognition. We have addressed a few challenges in our courses: 1) how to guide students efficiently to implement different project contents, e.g., using the scanner to acquire three-dimensional porcelain fragment data skillfully. 2) How to inspire students to learn and use various multimedia processing technologies and archaeology knowledge. 3) How the teacher adapts the teaching content according to the dynamic interests of the students. To address these challenges, students are grouped into two project teams based on their strengths and interests. We introduce collaborative learning and active learning strategies to help students learn and use different knowledge and address project problems and learning problems. We also invite the archaeological professor to teach basic archaeology knowledge in the class. Furthermore, to better understand the students' learning situation, we present a weekly project progress report approach, which can also help the teacher adjust the teaching content. This teaching approach can enhance the continuity of multimedia information processing teaching and stimulate students' enthusiasm and creativity in learning. Moreover, it can deepen the cultural atmosphere of the teaching in an engineering course. Xiaoqing Zhang 0001, Shengjie Ye, Zunjie Xiao, Jigen Tang, Jiang Liu 0001 |
FIE | 1 |
| 2021 | Gated Channel Attention Network for Cataract Classification on AS-OCT Image
Zunjie Xiao, Xiaoqing Zhang 0001, Risa Higashita, Jiang Liu 0001 |
ICONIP (3) | 2 |
| 2020 | Attention-based Saliency Hashing for Ophthalmic Image RetrievalabstractDeep hashing methods have been proved to be effective for the large-scale medical image search assisting reference-based diagnosis for clinicians. However, when the salient region plays a maximal discriminative role in ophthalmic image, existing deep hashing methods do not fully exploit the learning ability of the deep network to capture the features of salient regions pointedly. The different grades or classes of ophthalmic images may be share similar overall performance but have subtle differences that can be differentiated by mining salient regions. To address this issue, we propose a novel end-to-end network, named Attention-based Saliency Hashing (ASH), for learning compact hash-code to represent ophthalmic images. ASH embeds a spatial-attention module to focus more on the representation of salient regions and highlights their essential role in differentiating ophthalmic images. Benefiting from the spatial-attention module, the information of salient regions can be mapped into the hash-code for similarity calculation. Extensive experiments on two different modalities of ophthalmic image datasets demonstrate that the proposed ASH can further improve the retrieval performance compared to the state-of-the-art deep hashing methods due to the huge contributions of the spatial-attention module. Jiansheng Fang, Yanwu Xu 0001, Xiaoqing Zhang 0001, Jiang Liu 0001 |
BIBM | 3 |
| 2020 | Probabilistic Latent Factor Model for Collaborative Filtering with Bayesian InferenceabstractLatent Factor Model (LFM) is one of the most successful methods for Collaborative filtering (CF) in the recommendation system, in which both users and items are projected into a joint latent factor space. Base on matrix factorization applied usually in pattern recognition, LFM models user-item interactions as inner products of factor vectors of user and item in that space and can be efficiently solved by least square methods with optimal estimation. However, such optimal estimation methods are prone to overfitting due to the extreme sparsity of user-item interactions. In this paper, we propose a Bayesian treatment for LFM, named Bayesian Latent Factor Model (BLFM). Based on observed user-item interactions, we build a probabilistic factor model in which the regularization is introduced via placing prior constraint on latent factors, and the likelihood function is established over observations and parameters. Then we draw samples of latent factors from the posterior distribution with Variational Inference (VI) to predict expected value. We further make an extension to BLFM, called BLFMBias, incorporating user-dependent and item-dependent biases into the model for enhancing performance. Extensive experiments on the movie rating dataset show the effectiveness of our proposed models by compared with several strong baselines. Jiansheng Fang, Xiaoqing Zhang 0001, Yanwu Xu 0001, Ming Yang 0039, Jiang Liu 0001 |
ICPR | 2 |
| 2020 | A Novel Deep Learning Method for Nuclear Cataract Classification Based on Anterior Segment Optical Coherence Tomography ImagesabstractNuclear cataract is one of the most common types of cataract. In the recent, ophthalmologists are increasingly using anterior segment optical coherence tomography (AS-OCT) images to diagnose many ocular diseases including cataract. The relationship between cataract and the lens opacity based on AS-OCT images has been being studied in clinical pioneer research. However, using AS-OCT images to classify cataract automatically based on computer-aided diagnosis (CAD) technique has not been seriously studied. This paper proposes a novel Convolutional Neural Network (CNN) model named GraNet for nuclear cataract classification based on AS-OCT images. In the GraNet, we introduce a grading block to learn high-level feature representations based on the pointwise convolution method. To further improve the classification performance, we propose a simple and efficient cross-training method is comprised of focal loss and cross-entropy loss. Extensive experiments are conducted on the AS-OCT image dataset, the results demonstrate that the proposed methods achieve better nuclear cataract classification results than baselines. Xiaoqing Zhang 0001, Zunjie Xiao, Risa Higashita, Jiansheng Fang, Jiang Liu 0001 |
SMC | 1 |