VLDB 2026 Research / reviewers in the wild / expert
Yuanjie Zheng
dblp:44/3875
· DBLP profile ↗
93ranked-venue papers
21as first author
38since 2021 · last 2026
0000-0002-5786-2491ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 13 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 13 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 8 first-author · 11 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CaliDiff: Multi-rater annotation calibrating diffusion probabilistic model towards medical image segmentation
Junxia Wang, Jing Wang 0138, Baijing Chen, Yuanjie Zheng |
Medical Image Anal. | 6 |
| 2026 | LSDiff: Diffusion-Guided Level Set for Low-Contrast Lesion Boundary Segmentation
Wenhui Huang 0002, Jing Wang 0138, James C. Gee, Yuanjie Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Text-Image Co-Alignment for Weakly Supervised Polyp SegmentationabstractFully supervised polyp segmentation relies on costly pixel-level annotations. Although semi- and weakly supervised methods reduce annotation requirements, they still depend on partial mask supervision. Text-supervised segmentation is a promising alternative; however, for polyps, the key challenge is to ground instance-specific phrases to the correct lesion region under cluttered backgrounds and large appearance variations. Existing approaches often rely on coarse text-image alignment, limiting precise region-level semantic correspondence. In this paper, we propose Text-Image Co-Alignment (TICoA), a text-supervised framework for polyp segmentation. TICoA leverages large language models (LLMs)-generated structured clinical descriptions as weak supervision and formulates segmentation as a fine-grained phrase-region co-alignment problem. Through contrastive learning, TICoA explicitly associates query phrases with corresponding image regions to achieve robust semantic grounding under weak supervision. Architecturally, we adopt a State-Space Model (Mamba) to efficiently model long-range dependencies with linear computational complexity. To support effective cross-modal interaction, we further design a dedicated Mamba Fusion module with a Bi-Dimension Fusion (BiDF) strategy, which progressively propagates information along spatial and channel dimensions. Experiments on polyp datasets, with additional validation on skin lesion segmentation, demonstrate that TICoA is competitive with state-of-the-art weakly supervised methods. Our code and data are available at https://github.com/silentyuchen/TICoA. Wenhui Huang 0002, Zhen Pan, Yedi Zhang, Jingzhen He, James C. Gee, Yuanjie Zheng |
IEEE Trans. Medical Imaging | 7 |
| 2025 | MAMBA-Based Weakly Supervised Medical Image Segmentation with Cross-Modal Textual Information
Zhen Pan, Wenhui Huang 0002, Yuanjie Zheng |
MICCAI (8) | 3 |
| 2025 | MambaMER: Adaptive EEG-Guided Multimodal Emotion Recognition with Mamba
Xiangle Ping, Wenhui Huang 0002, Yuanjie Zheng |
MICCAI (1) | 3 |
| 2025 | Inpaint-Outpaint Synergy: Mask Refinement for Trimap-Free MattingabstractImage matting is a fundamental task in computer vision that focuses on the precise separation of foreground objects from their backgrounds in images. This process is essential for numerous applications, such as image editing, film production, and augmented reality. Traditional methods often rely on a trimap, a predefined region that helps to distinguish the foreground from the background. However, generating an accurate trimap requires the provision of raw alpha matte, which is labor intensive and prone to drawing errors, limiting the applicability of the relevant methods in practical applications. In this paper, a novel inpaint and outpaint synergy matting approach (IOSM) is proposed to generate masks for image matting tasks without supervision by iterating through the inpaint and outpaint processes, avoiding the dependence on trimap. Specifically, the inpaint process is able to eliminate the false-positive regions present in the initial mask, while the outpaint process reduces the false-negative regions by expanding the pixels to the outer regions. The above process reduces inaccurate regions in the initial mask by means of adversarial updating, providing accurate target information for the subsequent matting stage. By iteratively combining these two processes, a more accurate mask is generated, which is then fed into the mask-to-matte (MTM) module along with the original image to obtain the final alpha matte. This approach allows for the seamless integration of the mask with the original image, improving the matting task and resulting in higher-quality matte outcomes. Experimental results demonstrate that IOSM outperforms other mainstream methods on the AIM-500, Distinctions-646 and PPM-100 datasets. Our project page is available at:https://github.com/xuecheng990531/IOSM. Xuecheng Li, Yuanjie Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | RD-FGM: A novel model for high-quality and diverse food image generation and ingredient classification
Jing Wang 0138, Yuanjie Zheng, Junxia Wang, Sujuan Hou |
Expert Syst. Appl. | 2 |
| 2024 | MP2PMatch: A Mask-guided Part-to-Part Matching network based on transformer for occluded person re-identification
Guilin Lv, Yanhui Ding, Yuanjie Zheng |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | A generic plug & play diffusion-based denosing module for medical image segmentation
Guangju Li, Dehu Jin, Yuanjie Zheng, Jia Cui, Wei Gai |
Neural Networks | 3 |
| 2024 | Cross-View Representation Learning: A Superior ContextIB Method for Logo ClassificationabstractLogo classification systems have become increasingly important in various industries for tasks, such as infringement detection and industrial production. However, challenges still exist in logo classification due to real-world image background interference, the high similarity between classes, labeling difficulties, and the insufficient representation of occlusion in single-view logos. Many existing algorithms fail to consider the data characteristics and the intrinsic information of multiple views, which limits their performance. To overcome these limitations, we developed a novel Cross-View Information Awareness Network (CVIA-Net) for logo classification. To differentiate between similar logo categories, the CVIA-Net novel learns context-shared features of the same category via a self-supervised way without labeled, which solves the problem of insufficient features due to occlusion. For single-view images, CVIA-Net establishes a “bottleneck” representation to address background interference. Extensive experiments on three datasets demonstrate that it outperforms state-of-the-art methods. The method is expected to advance the development of cross-view representation learning. Jing Wang 0138, Yuanjie Zheng, Zeyu Han, Mei Lv, Sujuan Hou |
IEEE Signal Process. Lett. | 2 |
| 2024 | CorrDiff: Corrective Diffusion Model for Accurate MRI Brain Tumor SegmentationabstractAccurate segmentation of brain tumors in MRI images is imperative for precise clinical diagnosis and treatment. However, existing medical image segmentation methods exhibit errors, which can be categorized into two types: random errors and systematic errors. Random errors, arising from various unpredictable effects, pose challenges in terms of detection and correction. Conversely, systematic errors, attributable to systematic effects, can be effectively addressed through machine learning techniques. In this paper, we propose a corrective diffusion model for accurate MRI brain tumor segmentation by correcting systematic errors. This marks the first application of the diffusion model for correcting systematic segmentation errors. Additionally, we introduce the Vector Quantized Variational Autoencoder (VQ-VAE) to compress the original data into a discrete coding codebook. This not only reduces the dimensionality of the training data but also enhances the stability of the correction diffusion model. Furthermore, we propose the Multi-Fusion Attention Mechanism, which can effectively enhances the segmentation performance of brain tumor images, and enhance the flexibility and reliability of the corrective diffusion model. Our model is evaluated on the BRATS2019, BRATS2020, and Jun Cheng datasets. Experimental results demonstrate the effectiveness of our model over state-of-the-art methods in brain tumor segmentation. Wenhui Huang 0002, Yuanjie Zheng |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | TransMatch: A Transformer-Based Multilevel Dual-Stream Feature Matching Network for Unsupervised Deformable Image RegistrationabstractFeature matching, which refers to establishing the correspondence of regions between two images (usually voxel features), is a crucial prerequisite of feature-based registration. For deformable image registration tasks, traditional feature-based registration methods typically use an iterative matching strategy for interest region matching, where feature selection and matching are explicit, but specific feature selection schemes are often useful in solving application-specific problems and require several minutes for each registration. In the past few years, the feasibility of learning-based methods, such as VoxelMorph and TransMorph, has been proven, and their performance has been shown to be competitive compared to traditional methods. However, these methods are usually single-stream, where the two images to be registered are concatenated into a 2-channel whole, and then the deformation field is output directly. The transformation of image features into interimage matching relationships is implicit. In this paper, we propose a novel end-to-end dual-stream unsupervised framework, named TransMatch, where each image is fed into a separate stream branch, and each branch performs feature extraction independently. Then, we implement explicit multilevel feature matching between image pairs via the query-key matching idea of the self-attention mechanism in the Transformer model. Comprehensive experiments are conducted on three 3D brain MR datasets, LPBA40, IXI, and OASIS, and the results show that the proposed method achieves state-of-the-art performance in several evaluation metrics compared to the commonly utilized registration methods, including SyN, NiftyReg, VoxelMorph, CycleMorph, ViT-V-Net, and TransMorph, demonstrating the effectiveness of our model in deformable medical image registration. Yuanjie Zheng, James C. Gee |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Deep Learning for Logo Detection: A SurveyabstractLogo detection has gradually become a research hotspot in the field of computer vision and multimedia for its various applications, such as social media monitoring, intelligent transportation, and video advertising recommendation. Recent advances in this area are dominated by deep learning-based solutions, where many datasets, learning strategies, network architectures, and loss functions have been employed. This article reviews the advance in applying deep learning techniques to logo detection. First, we discuss a comprehensive account of public datasets designed to facilitate performance evaluation of logo detection algorithms, which tend to be more diverse, more challenging, and more reflective of real life. Next, we perform an in-depth analysis of the existing logo detection strategies and their strengths and weaknesses of each learning strategy. Subsequently, we summarize the applications of logo detection in various fields, from intelligent transportation and brand monitoring to copyright and trademark compliance. Finally, we analyze the potential challenges and present the future directions for the development of logo detection. This study aims better to inform readers about the current state of logo detection and encourage more researchers to get involved in logo detection. Sujuan Hou, Weiqing Min, Yanna Zhao, Yuanjie Zheng, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Link Traffic-Delay Mapping Model Learning Based on Multi-Class Samples in Software-Defined NetworksabstractDelays are crucial factors in the service management of networks, especially software-defined networks. Unfortunately, it is very difficult to accurately model a traffic-delay mapping without any assumptions on an uncertain network. In this article, we present a machine learning-based solution to generate a mapping between link traffic and link delay in software-defined networks. The proposed solution only requires a small number of link delay samples from the production network. The small number of link delay samples is not sufficient for learning link traffic-delay mapping. To solve the above problem, we extend the link delay-related data via a sample transfer method and a distributed path delay data collection method without the assistance of the controller. We design a link traffic-delay mapping learning solution using the above three classes of data. This solution uses a traffic segment-based statistical mechanism to deduce the mean link delay effectively from the collected path delay information and implements effective sample transfer via a distance-based approximation. On the basis of specially designed deep learning structures and training procedures, the proposed learning solution effectively builds traffic-delay mapping models using the samples transferred from an experimental network and the samples of the production network. Xinchang Zhang 0001, Maoli Wang, Yuanjie Zheng, Dongjie Liu |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | A Cross-direction Task Decoupling Network for Small Logo DetectionabstractLogo detection plays an integral role in many applications. However, handling small logos is still difficult since they occupy too few pixels in the image, which burdens the extraction of discriminative features. The aggregation of small logos also brings a great challenge to the classification and localization of logos. To solve these problems, we creatively propose Cross-direction Task Decoupling Network (CTDNet) for small logo detection. We first introduce Cross-direction Feature Pyramid (CFP) to realize cross-direction feature fusion by adopting horizontal transmission and vertical transmission. In addition, Multi-frequency Task Decoupling Head (MTDH) decouples the classification and localization tasks into two branches. A multi-frequency attention convolution branch is designed to achieve more accurate regression by combining discrete cosine transform and convolution creatively. Comprehensive experiments on four logo datasets demonstrate the effectiveness and efficiency of the proposed method. Sujuan Hou, Xingzhuo Li, Weiqing Min, Jing Wang 0138, Yuanjie Zheng, Shuqiang Jiang |
ICME | 6 |
| 2023 | Instance-Aware Diffusion Model for Gland Segmentation in Colon Histology Images
Mengxue Sun, Wenhui Huang 0002, Yuanjie Zheng |
MICCAI (6) | 3 |
| 2023 | Few-shot logo detectionabstractAbstract The proliferation of deep learning has driven research into deep learning‐based logo detection, which usually needs a large number of annotated data to train the model. However, due to the occasional appearance of new brands or the high cost of annotation, the number of training data is limited. Against this backdrop, the authors adapt the few‐shot object detection into logo detection, and thus present a cutting‐edge method called Double Classification Head (DCH) for Few‐Shot Logo Detection (DCH‐FSLogo), which aims at detecting the unseen logo classes using few annotated data. Unlike the traditional few‐shot detection, some logo objects are similar to their backgrounds and have diverse shapes as well. For this reason, the authors adopt balanced feature pyramid and deformable Region of Interest pooling in DCH‐FSLogo, this enhances the feature extraction capability and adapts to the different logo shapes. In addition, we introduce the DCH for few‐shot logo detection to detect logo objects using few annotated data. Specifically, we use an extra classification head for the base classes to ease the influence from the novel classes. The experimental results on four datasets, namely: FlickrLogos‐32, FoodLogoDet‐1500‐100, LogoDet‐3K‐100 and QMUL‐OpenLogo‐100, demonstrate that our method achieves better performance. Sujuan Hou, Wenjie Liu 0001, Karim Awudu, Zhixiang Jia, Weikuan Jia, Yuanjie Zheng |
IET Comput. Vis. | 6 |
| 2023 | Driver Drowsiness EEG Detection Based on Tree Federated Learning and Interpretable NetworkabstractAccurate identification of driver's drowsiness state through Electroencephalogram (EEG) signals can effectively reduce traffic accidents, but EEG signals are usually stored in various clients in the form of small samples. This study attempts to construct an efficient and accurate privacy-preserving drowsiness monitoring system, and proposes a fusion model based on tree Federated Learning (FL) and Convolutional Neural Network (CNN), which can not only identify and explain the driver's drowsiness state, but also integrate the information of different clients under the premise of privacy protection. Each client uses CNN with the Global Average Pooling (GAP) layer and shares model parameters. The tree FL transforms communication relationships into a graph structure, and model parameters are transmitted in parallel along connected branches of the graph. Moreover, the Class Activation Mapping (CAM) is used to find distinctive EEG features for representing specific classes. On EEG data of 11 subjects, it is found that this method has higher average accuracy, F1-score and AUC than the traditional classification method, reaching 73.56%, 73.26% and 78.23%, respectively. Compared with the traditional FL algorithm, this method better protects the driver's privacy and improves communication efficiency. Huiyu Zhou 0001, Weikuan Jia, Yuanjie Zheng |
Int. J. Neural Syst. | 6 |
| 2023 | TISS-net: Brain tumor image synthesis and segmentation using cascaded dual-task networks and error-prediction consistencyabstractAccurate segmentation of brain tumors from medical images is important for diagnosis and treatment planning, and it often requires multi-modal or contrast-enhanced images. However, in practice some modalities of a patient may be absent. Synthesizing the missing modality has a potential for filling this gap and achieving high segmentation performance. Existing methods often treat the synthesis and segmentation tasks separately or consider them jointly but without effective regularization of the complex joint model, leading to limited performance. We propose a novel brain Tumor Image Synthesis and Segmentation network (TISS-Net) that obtains the synthesized target modality and segmentation of brain tumors end-to-end with high performance. First, we propose a dual-task-regularized generator that simultaneously obtains a synthesized target modality and a coarse segmentation, which leverages a tumor-aware synthesis loss with perceptibility regularization to minimize the high-level semantic domain gap between synthesized and real target modalities. Based on the synthesized image and the coarse segmentation, we further propose a dual-task segmentor that predicts a refined segmentation and error in the coarse segmentation simultaneously, where a consistency between these two predictions is introduced for regularization. Our TISS-Net was validated with two applications: synthesizing FLAIR images for whole glioma segmentation, and synthesizing contrast-enhanced T1 images for Vestibular Schwannoma segmentation. Experimental results showed that our TISS-Net largely improved the segmentation accuracy compared with direct segmentation from the available modalities, and it outperformed state-of-the-art image synthesis-based segmentation methods. Jianghao Wu 0001, Lu Wang 0002, Shuojue Yang, Yuanjie Zheng, Jonathan Shapey, Tom Vercauteren, Sotirios Bisdas, Robert Bradford, Shakeel R. Saeed, Neil Kitchen, Sébastien Ourselin, Shaoting Zhang 0001, Guotai Wang |
Neurocomputing | 5 |
| 2023 | AMNet: Adaptive multi-level network for deformable registration of 3D brain MR images
Tongtong Che, Xiuying Wang 0001, Kun Zhao 0014, Debin Zeng, Qiongling Li, Yuanjie Zheng, Jian Wang 0120 |
Medical Image Anal. | 7 |
| 2023 | Information bottleneck-based interpretable multitask network for breast cancer classification and segmentation
Junxia Wang, Yuanjie Zheng, Xinmeng Li, Chongjing Wang, James Gee, Wenhui Huang 0002 |
Medical Image Anal. | 2 |
| 2023 | A Holistically-Guided Decoder for Deep Representation Learning With Applications to Semantic Segmentation and Object DetectionabstractBoth high-level and high-resolution feature representations are of great importance in various visual understanding tasks. To acquire high-resolution feature maps with high-level semantic information, one common strategy is to adopt dilated convolutions in the backbone networks to extract high-resolution feature maps, such as the dilatedFCN-based methods for semantic segmentation. However, due to many convolution operations are conducted on the high-resolution feature maps, such methods have large computational complexity and memory consumption. To balance the performance and efficiency, there also exist encoder-decoder structures that gradually recover the spatial information by combining multi-level feature maps from a feature encoder, such as the FPN architecture for object detection and the U-Net for semantic segmentation. Although being more efficient, the performances of existing encoder-decoder methods for semantic segmentation are far from comparable with the dilatedFCN-based methods. In this paper, we propose one novel holistically-guided decoder which is introduced to obtain the high-resolution semantic-rich feature maps via the multi-scale features from the encoder. The decoding is achieved via novel holistic codeword generation and codeword assembly operations, which take advantages of both the high-level and low-level features from the encoder features. With the proposed holistically-guided decoder, we implement the EfficientFCN architecture for semantic segmentation and HGD-FPN for object detection and instance segmentation. The EfficientFCN achieves comparable or even better performance than state-of-the-art methods with only 1/3 of their computational costs for semantic segmentation on PASCAL Context, PASCAL VOC, ADE20K datasets. Meanwhile, the proposed HGD-FPN achieves higher mean Average Precision (mAP) when integrated into several object detection frameworks with ResNet-50 encoding backbones. Junjun He, Yuanjie Zheng, Shuai Yi, Xiaogang Wang 0001, Hongsheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Image Matting With Deep Gaussian ProcessabstractWe observe a common characteristic between the classical propagation-based image matting and the Gaussian process (GP)-based regression. The former produces closer alpha matte values for pixels associated with a higher affinity, while the outputs regressed by the latter are more correlated for more similar inputs. Based on this observation, we reformulate image matting as GP and find that this novel matting-GP formulation results in a set of attractive properties. First, it offers an alternative view on and approach to propagation-based image matting. Second, an application of kernel learning in GP brings in a novel deep matting-GP technique, which is pretty powerful for encapsulating the expressive power of deep architecture on the image relative to its matting. Third, an existing scalable GP technique can be incorporated to further reduce the computational complexity to$\mathcal {O}(n)$from$\mathcal {O}(n^{3})$of many conventional matting propagation techniques. Our deep matting-GP provides an attractive strategy toward addressing the limit of widespread adoption of deep learning techniques to image matting for which a sufficiently large labeled dataset is lacking. A set of experiments on both synthetically composited images and real-world images show the superiority of the deep matting-GP to not only the classical propagation-based matting techniques but also modern deep learning-based approaches. Yuanjie Zheng, Yunshuai Yang, Tongtong Che, Sujuan Hou, Wenhui Huang 0002, Yue Gao 0002, Ping Tan 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Decoding Color Visual Working Memory from EEG Signals Using Graph Convolutional Neural NetworksabstractColor has an important role in object recognition and visual working memory (VWM). Decoding color VWM in the human brain is helpful to understand the mechanism of visual cognitive process and evaluate memory ability. Recently, several studies showed that color could be decoded from scalp electroencephalogram (EEG) signals during the encoding stage of VWM, which process visible information with strong neural coding. Whether color could be decoded from other VWM processing stages, especially the maintaining stage which processes invisible information, is still unknown. Here, we constructed an EEG color graph convolutional network model (ECo-GCN) to decode colors during different VWM stages. Based on graph convolutional networks, ECo-GCN considers the graph structure of EEG signals and may be more efficient in color decoding. We found that (1) decoding accuracies for colors during the encoding, early, and late maintaining stages were 81.58%, 79.36%, and 77.06%, respectively, exceeding those during the pre-stimuli stage (67.34%), and (2) the decoding accuracy during maintaining stage could predict participants' memory performance. The results suggest that EEG signals during the maintaining stage may be more sensitive than behavioral measurement to predict the VWM performance of human, and ECo-GCN provides an effective approach to explore human cognitive function. Xiaowei Che, Yuanjie Zheng, Sutao Song, Shouxin Li |
Int. J. Neural Syst. | 2 |
| 2022 | Automatic Seizure Identification from EEG Signals Based on Brain Connectivity LearningabstractEpilepsy is a neurological disorder caused by brain dysfunction, which could cause uncontrolled behavior, loss of consciousness and other hazards. Electroencephalography (EEG) is an indispensable auxiliary tool for clinical diagnosis. Great progress has been made by current seizure identification methods. However, the performance of the methods on different patients varies a lot. In order to deal with this problem, we propose an automatic seizure identification method based on brain connectivity learning. The connectivity of different brain regions is modeled by a graph. Different from the manually defined graph structure, our method can extract the optimal graph structure and EEG features in an end-to-end manner. Combined with the popular graph attention neural network (GAT), this method achieves high performance and stability on different patients from the CHB-MIT dataset. The average values of accuracy, sensitivity, specificity, F1-score and AUC of the proposed model are 98.90%, 98.33%, 98.48%, 97.72% and 98.54%, respectively. The standard deviations of the above five indicators are 0.0049, 0.0125, 0.0116 and 0.0094, respectively. Compared with the existing seizure identification methods, the stability of the proposed model is improved by 78-95%. Yanna Zhao, Mingrui Xue, Changxu Dong, Jiatong He, Dengyu Chu, Gaobo Zhang, Fangzhou Xu, Xinting Ge, Yuanjie Zheng |
Int. J. Neural Syst. | 9 |
| 2022 | Self-adapting spiking neural P systems with refractory period and propagation delay
Yuzhen Zhao, Xiyu Liu 0001, Minghe Sun, Feng Qi 0002, Yuanjie Zheng |
Inf. Sci. | 6 |
| 2022 | SymReg-GAN: Symmetric Image Registration With Generative Adversarial NetworksabstractSymmetric image registration estimates bi-directional spatial transformations between images while enforcing an inverse-consistency. Its capability of eliminating bias introduced inevitably by generic single-directional image registration allows more precise analysis in different interdisciplinary applications of image registration, e.g., computational anatomy and shape analysis. However, most existing symmetric registration techniques especially for multimodal images are limited by low speed from the commonly-used iterative optimization, hardship in exploring inter-modality relations or high labor cost for labeling data. We propose SymReg-GAN to shatter these limits, which is a novel generative adversarial networks (GAN) based approach to symmetric image registration. We formulate symmetric registration of unimodal/multimodal images as a conditional GAN and train it with a semi-supervised strategy. The registration symmetry is realized by introducing a loss for encouraging that the cycle composed of the geometric transformation from one image to another and its reverse should bring an image back. The semi-supervised learning enables both the precious labeled data and large amounts of unlabeled data to be fully exploited. Experimental results from six public brain magnetic resonance imaging (MRI) datasets and 1 our own computed tomography (CT) and MRI dataset demonstrate the superiority of SymReg-GAN to several existing state-of-the-art methods. Yuanjie Zheng, Xiaodan Sui, Yanyun Jiang, Tongtong Che, Shaoting Zhang 0001, Jie Yang 0002, Hongsheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Multi-feature deep information bottleneck network for breast cancer classification in contrast enhanced spectral mammography
Jingqi Song, Yuanjie Zheng, Jing Wang 0138, Muhammad Zakir Ullah, Xuecheng Li, Zhenxing Zou, Guocheng Ding |
Pattern Recognit. | 2 |
| 2022 | An Attention Based Bidirectional LSTM Method to Predict the Binding of TCR and EpitopeabstractThe T-cell epitope prediction has always been a long-term challenge in immunoinformatics and bioinformatics. Studying the specific recognition between T-cell receptor (TCR) and peptide-major histocompatibility complex (p-MHC) complexes can help us better understand the immune mechanism, it's also make a signification contribution in developing vaccines and targeted drugs. Meanwhile, more advanced methods are needed for distinguishing TCRs binding from different epitopes. In this paper, we introduce a hybrid model composed of bidirectional long short-term memory networks (BiLSTM), attention and convolutional neural networks (CNN) that can identified the binding of TCRs to epitopes. The BiLSTM can more completely extract amino acid forward and backward information in the sequence, and attention mechanism can focus on amino acids at certain positions from complex sequences to capture the most important feature, then CNN was used to further extract salient features to predict the binding of TCR-epitope. In McPAS dataset, the AUC value (the area under ROC curve) of naive TCR-epitope binding is 0.974 and specific TCR-epitope binding is 0.887. The model has achieved better prediction results than other existing models (TCRGP, ERGO, NetTCR), and some experiments are used to analyze the advantages of our model. The algorithm is available at https://github.com/bijingshu/BiAttCNN.git. Jingshu Bi, Yuanjie Zheng, Chongjing Wang, Yanhui Ding |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | LogoDet-3K: A Large-scale Image Dataset for Logo DetectionabstractLogo detection has been gaining considerable attention because of its wide range of applications in the multimedia field, such as copyright infringement detection, brand visibility monitoring, and product brand management on social media. In this article, we introduce LogoDet-3K, the largest logo detection dataset with full annotation, which has 3,000 logo categories, about 200,000 manually annotated logo objects, and 158,652 images. LogoDet-3K creates a more challenging benchmark for logo detection, for its higher comprehensive coverage and wider variety in both logo categories and annotated objects compared with existing datasets. We describe the collection and annotation process of our dataset and analyze its scale and diversity in comparison to other datasets for logo detection. We further propose a strong baseline method Logo-Yolo, which incorporates Focal loss and CIoU loss into the basic YOLOv3 framework for large-scale logo detection. It obtains about 4% improvement on the average performance compared with YOLOv3, and greater improvements compared with reported several deep detection models on LogoDet-3K. We perform extensive evaluation on three other existing datasets to further verify on both logo detection and retrieval tasks, and we demonstrate better generalization ability of LogoDet-3K on logo detection and retrieval tasks. The LogoDet-3K dataset is used to promote large-scale logo-related research. The code and LogoDet-3K can be found at https://github.com/Wangjing1551/LogoDet-3K-Dataset. Jing Wang 0138, Weiqing Min, Sujuan Hou, Shengnan Ma, Yuanjie Zheng, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Exploiting Probabilistic Siamese Visual Tracking with a Conditional Variational AutoencoderabstractVisual tracking is a fundamental capability for robots tasked with humans and environment interaction. However, state-of-the-art visual tracking methods are still prone to failures and are imprecise when applied to challenging stereos, and their results are generally confidence agonistic. These methods depend on an embedded deep learning model to provide deterministic features or regression maps. A deterministic output with low confidence can result in disastrous consequences and lacks evidence needed for subsequent operations. Moreover, training data ambiguities or noise in the observations (so-called data uncertainty) can also lead to inherent uncertainty. In this paper, we focus on exploiting probabilistic Siamese visual tracking with a conditional variational autoencoder (CVAE). First, we build a bridge between the Siamese architecture and the CVAE and propose a novel Bayesian visual tracking method. Second, the proposed method generates a complete probability distribution that enables the production of multiple plausible tracking outputs. Third, CVAE conditioned by ground truth data encodes a low-dimensional latent space and conducts noise-injection training to prevent overfitting. Our proposed tracking method outperformed the state-of-the-art trackers on the VOT2016, VOT2018 and TColor-128 datasets. Wenhui Huang 0002, Jason Gu, Peiyong Duan, Sujuan Hou, Yuanjie Zheng |
ICRA | 5 |
| 2021 | Synthesis of Contrast-Enhanced Spectral Mammograms from Low-Energy Mammograms Using cGAN-Based Synthesis Network
Yanyun Jiang, Yuanjie Zheng, Weikuan Jia, Sutao Song, Yanhui Ding |
MICCAI (7) | 2 |
| 2021 | FoodLogoDet-1500: A Dataset for Large-Scale Food Logo Detection via Multi-Scale Feature Decoupling NetworkabstractFood logo detection plays an important role in the multimedia for its wide real-world applications, such as food recommendation of the self-service shop and infringement detection on e-commerce platforms. A large-scale food logo dataset is urgently needed for developing advanced food logo detection algorithms. However, there are no available food logo datasets with food brand information. To support efforts towards food logo detection, we introduce the dataset FoodLogoDet-1500, a new large-scale publicly available food logo dataset, which has 1,500 categories, about 100,000 images and about 150,000 manually annotated food logo objects. We describe the collection and annotation process of FoodLogoDet-1500, analyze its scale and diversity, and compare it with other logo datasets. To the best of our knowledge, FoodLogoDet-1500 is the first largest publicly available high-quality dataset for food logo detection. The challenge of food logo detection lies in the large-scale categories and similarities between food logo categories. For that, we propose a novel food logo detection method Multi-scale Feature Decoupling Network (MFDNet), which decouples classification and regression into two branches and focuses on the classification branch to solve the problem of distinguishing multiple food logo categories. Specifically, we introduce the feature offset module, which utilizes the deformation-learning for optimal classification offset and can effectively obtain the most representative features of classification in detection. In addition, we adopt a balanced feature pyramid in MFDNet, which pays attention to global information, balances the multi-scale feature maps, and enhances feature extraction capability. Comprehensive experiments on FoodLogoDet-1500 and other two popular benchmark logo datasets demonstrate the effectiveness of the proposed method. The code and FoodLogoDet-1500 can be found at https://github.com/hq03/FoodLogoDet-1500-Dataset. Weiqing Min, Jing Wang 0138, Sujuan Hou, Yuanjie Zheng, Shuqiang Jiang |
ACM Multimedia | 5 |
| 2021 | Cross-View Representation Learning for Multi-View Logo Classification with Information BottleneckabstractMulti-view logo classification is a challenging task due to the cross-view misalignment of logo image varies under different viewpoints, large intra-classes and small inter-classes variation of logo appearance. Cross-view data can represent objects from different views and thus provide complementary information for data analysis. However, most existing multi-view algorithms usually maximize the correlation between different views for consistency. Those methods ignore the interaction among different views and may cause semantic bias during the process of common feature learning. In this paper, we investigate the information bottleneck (IB) to the multi-view learning for extracting the different view common features of one category, named Dual-View Information Bottleneck representation (Dual-view IB). To the best of our knowledge, this is the first cross-view learning method for logo classification. Specifically, we maximize the mutual information between the representations of the two views to achieve the preservation of key features in the classification task, while eliminating the redundant information that is not shared between the two views. In addition, due to the unbalance of samples and limited computing resources, we further introduce a novel Pair Batch Data Augmentation (PB) algorithm for Dual-view IB model, which applies augmentations from a learned policy based on replicates instances of two samples within the same batch. Comprehensive experiments on three existing benchmark datasets, which demonstrate the effectiveness of the proposed method that outperforms the methods in the state of the art. The proposed method is expected to further the development of cross-view representation learning. Jing Wang 0138, Yuanjie Zheng, Jingqi Song, Sujuan Hou |
ACM Multimedia | 2 |
| 2021 | Graph Attention Network with Focal Loss for Seizure Detection on Electroencephalography SignalsabstractAutomatic seizure detection from electroencephalogram (EEG) plays a vital role in accelerating epilepsy diagnosis. Previous researches on seizure detection mainly focused on extracting time-domain and frequency-domain features from single electrodes, while paying little attention to the positional correlations between different EEG channels of the same subject. Moreover, data imbalance is common in seizure detection scenarios where the duration of nonseizure periods is much longer than the duration of seizures. To cope with the two challenges, a novel seizure detection method based on graph attention network (GAT) is presented. The approach acts on graph-structured data and takes the raw EEG data as input. The positional relationship between different EEG signals is exploited by GAT. The loss function of the GAT model is redefined using the focal loss to tackle data imbalance problem. Experiments are conducted on the CHB-MIT dataset. The accuracy, sensitivity and specificity of the proposed method are 98.89[Formula: see text], 97.10[Formula: see text] and 99.63[Formula: see text], respectively. Yanna Zhao, Gaobo Zhang, Changxu Dong, Fangzhou Xu, Yuanjie Zheng |
Int. J. Neural Syst. | 6 |
| 2021 | Cascaded MultiTask 3-D Fully Convolutional Networks for Pancreas SegmentationabstractAutomatic pancreas segmentation is crucial to the diagnostic assessment of diabetes or pancreatic cancer. However, the relatively small size of the pancreas in the upper body, as well as large variations of its location and shape in retroperitoneum, make the segmentation task challenging. To alleviate these challenges, in this article, we propose a cascaded multitask 3-D fully convolution network (FCN) to automatically segment the pancreas. Our cascaded network is composed of two parts. The first part focuses on fast locating the region of the pancreas, and the second part uses a multitask FCN with dense connections to refine the segmentation map for fine voxel-wise segmentation. In particular, our multitask FCN with dense connections is implemented to simultaneously complete tasks of the voxel-wise segmentation and skeleton extraction from the pancreas. These two tasks are complementary, that is, the extracted skeleton provides rich information about the shape and size of the pancreas in retroperitoneum, which can boost the segmentation of pancreas. The multitask FCN is also designed to share the low- and mid-level features across the tasks. A feature consistency module is further introduced to enhance the connection and fusion of different levels of feature maps. Evaluations on two pancreas datasets demonstrate the robustness of our proposed method in correctly segmenting the pancreas in various settings. Our experimental results outperform both baseline and state-of-the-art methods. Moreover, the ablation study shows that our proposed parts/modules are critical for effective multitask learning. Jie Xue 0001, Kelei He, Dong Nie, Ehsan Adeli-Mosabbeb, Zhenshan Shi, Seong-Whan Lee, Yuanjie Zheng, Xiyu Liu 0001, Dengwang Li, Dinggang Shen |
IEEE Trans. Cybern. | 7 |
| 2021 | SDOF-GAN: Symmetric Dense Optical Flow Estimation With Generative Adversarial NetworksabstractThere is a growing consensus in computer vision that symmetric optical flow estimation constitutes a better model than a generic asymmetric one for its independence of the selection of source/target image. Yet, convolutional neural networks (CNNs), that are considered the de facto standard vision model, deal with the asymmetric case only in most cutting-edge CNNs-based optical flow techniques. We bridge this gap by introducing a novel model named SDOF-GAN: symmetric dense optical flow with generative adversarial networks (GANs). SDOF-GAN realizes a consistency between the forward mapping (source-to-target) and the backward one (target-to-source) by ensuring that they are inverse of each other with an inverse network. In addition, SDOF-GAN leverages a GAN model for which the generator estimates symmetric optical flow fields while the discriminator differentiates the "real" ground-truth flow field from a "fake" estimation by assessing the flow warping error. Finally, SDOF-GAN is trained in a semi-supervised fashion to enable both the precious labeled data and large amounts of unlabeled data to be fully-exploited. We demonstrate significant performance benefits of SDOF-GAN on five publicly-available datasets in contrast to several representative state-of-the-art models for optical flow estimation. Tongtong Che, Yuanjie Zheng, Yunshuai Yang, Sujuan Hou, Weikuan Jia, Jie Yang 0002, Chen Gong 0002 |
IEEE Trans. Image Process. | 2 |
| 2021 | Solving Jigsaw Puzzles via Nonconvex Quadratic Programming With the Projected Power MethodabstractJigsaw puzzles consist of reconstructing a picture that has been divided into many interlocking pieces. This paper describes an automatic global method for solving the square-piece jigsaw puzzle problem in which neither the orientations nor the locations of the jigsaw pieces are known. This hard combinatorial sorting task is formulated as a nonconvex quadratic programming problem that is solved via the projected power method. Specifically, this work aims to specify the locations and orientations of puzzle pieces by maximizing a constrained quadratic function that resolves an optimized permutation matrix composed of the noisy pairwise affinities between jigsaw pieces. The experimental results obtained in the MIT, McGill and Pomeranz datasets indicate that our method outperforms state-of-the-art techniques. Fang Yan 0003, Yuanjie Zheng, Jinyu Cong, Liu Liu 0014, Dacheng Tao, Sujuan Hou |
IEEE Trans. Multim. | 2 |
| 2020 | Logo-2K+: A Large-Scale Logo Dataset for Scalable Logo ClassificationabstractLogo classification has gained increasing attention for its various applications, such as copyright infringement detection, product recommendation and contextual advertising. Compared with other types of object images, the real-world logo images have larger variety in logo appearance and more complexity in their background. Therefore, recognizing the logo from images is challenging. To support efforts towards scalable logo classification task, we have curated a dataset, Logo-2K+, a new large-scale publicly available real-world logo dataset with 2,341 categories and 167,140 images. Compared with existing popular logo datasets, such as FlickrLogos-32 and LOGO-Net, Logo-2K+ has more comprehensive coverage of logo categories and larger quantity of logo images. Moreover, we propose a Discriminative Region Navigation and Augmentation Network (DRNA-Net), which is capable of discovering more informative logo regions and augmenting these image regions for logo classification. DRNA-Net consists of four sub-networks: the navigator sub-network first selected informative logo-relevant regions guided by the teacher sub-network, which can evaluate its confidence belonging to the ground-truth logo class. The data augmentation sub-network then augments the selected regions via both region cropping and region dropping. Finally, the scrutinizer sub-network fuses features from augmented regions and the whole image for logo classification. Comprehensive experiments on Logo-2K+ and other three existing benchmark datasets demonstrate the effectiveness of proposed method. Logo-2K+ and the proposed strong baseline DRNA-Net are expected to further the development of scalable logo image recognition, and the Logo-2K+ dataset can be found at https://github.com/msn199959/Logo-2k-plus-Dataset. Jing Wang 0138, Weiqing Min, Sujuan Hou, Shengnan Ma, Yuanjie Zheng, Haishuai Wang, Shuqiang Jiang |
AAAI | 5 |
| 2020 | Multi-organ Segmentation via Co-training Weight-Averaged Models from Few-Organ Datasets
Rui Huang 0001, Yuanjie Zheng, Shaoting Zhang 0001, Hongsheng Li 0001 |
MICCAI (4) | 2 |
| 2020 | A CTR prediction model based on user interest via attention mechanism
Huichuan Duan, Yuanjie Zheng, Yu Wang 0228 |
Appl. Intell. | 3 |
| 2020 | Revealing False Positive Features in Epileptic EEG IdentificationabstractFeature selection plays a vital role in the detection and discrimination of epileptic seizures in electroencephalogram (EEG) signals. The state-of-the-art EEG classification techniques commonly entail the extraction of the multiple features that would be fed into classifiers. For some techniques, the feature selection strategies have been used to reduce the dimensionality of the entire feature space. However, most of these approaches focus on the performance of classifiers while neglecting the association between the feature and the EEG activity itself. To enhance the inner relationship between the feature subset and the epileptic EEG task with a promising classification accuracy, we propose a machine learning-based pipeline using a novel feature selection algorithm built upon a knockoff filter. First, a number of temporal, spectral, and spatial features are extracted from the raw EEG signals. Second, the proposed feature selection algorithm is exploited to obtain the optimal subgroup of features. Afterwards, three classifiers including [Formula: see text]-nearest neighbor (KNN), random forest (RF) and support vector machine (SVM) are used. The experimental results on the Bonn dataset demonstrate that the proposed approach outperforms the state-of-the-art techniques, with accuracy as high as 99.93% for normal and interictal EEG discrimination and 98.95% for interictal and ictal EEG classification. Meanwhile, it has achieved satisfactory sensitivity (95.67% in average), specificity (98.83% in average), and accuracy (98.89% in average) over the Freiburg dataset. Jian Lian, Yunfeng Shi, Yan Zhang 0060, Weikuan Jia, Xiaojun Fan, Yuanjie Zheng |
Int. J. Neural Syst. | 6 |
| 2020 | MMCL-Net: Spinal disease diagnosis in global mode using progressive multi-task joint learning
Yanfei Hong, Benzheng Wei, Zhongyi Han, Xiang Li 0114, Yuanjie Zheng, Shuo Li 0001 |
Neurocomputing | 5 |
| 2020 | A novel forecasting model for the long-term fluctuation of time series based on polar fuzzy information granules
Chao Luo 0001, Yuanjie Zheng |
Inf. Sci. | 3 |
| 2020 | Algebraic dynamics of k-valued fuzzy cognitive maps and its stabilization
Chao Luo 0001, Yuanjie Zheng |
Knowl. Based Syst. | 3 |
| 2020 | Multimodality registration for ocular multispectral images via co-embedding
Yan Zhang 0060, Jian Lian, Weikuan Jia, Chengjiang Li, Yuanjie Zheng |
Neural Comput. Appl. | 5 |
| 2020 | Even faster retinal vessel segmentation via accelerated singular value decomposition
Yan Zhang 0060, Jian Lian, Weikuan Jia, Chengjiang Li, Yuanjie Zheng |
Neural Comput. Appl. | 6 |
| 2020 | Time-free cell-like P systems with multiple promoters/inhibitors
Yuzhen Zhao, Xiyu Liu 0001, Minghe Sun, Jianhua Qu, Yuanjie Zheng |
Theor. Comput. Sci. | 5 |
| 2020 | Controllability of k-Valued Fuzzy Cognitive MapsabstractFuzzy cognitive maps (FCMs) as a kind of knowledge-based tools are widely applied to model complex dynamical systems using causal relations. Besides the representation and reasoning of systems behaviors, how to control the given systems into a desirable target by causal objects established by FCMs is also an open problem. Although, so far, there are some existing works about the applications of FCMs on the control-related problems, it is still a lack of the theoretical analysis in this domain. In this paper, the controllability of k-valued FCMs is studied. To improve the universality of models, a temporal extension of generalized FCMs is implemented. By means of semitensor product, the algebraic representation of k-valued FCMs with controls is established and a generalized formula of control-depending network transition matrices is achieved. A necessary and sufficient condition is proved to determine the control-depending fixed points of k-valued FCMs with temporalization. By utilizing three kinds of controls, the controllability of the discrete FCMs is discussed, respectively. The reachability condition of a specific target state from a given initial state at time s is studied, and the reachable set along with the corresponding reachable probability are also provided by analytic formula. Results provide a way to make FCMs evolving into the designed states by controls, which can further conduct the behaviors of the modeled systems in reality. Examples are shown to demonstrate the effectiveness and feasibility of the proposed scheme. Chao Luo 0001, Haiyue Wang, Yuanjie Zheng |
IEEE Trans. Fuzzy Syst. | 3 |
| 2020 | Deep-Learning-Based Small Surface Defect Detection via an Exaggerated Local Variation-Based Generative Adversarial NetworkabstractSurface detection of small defects plays a vital role in manufacturing and has attracted broad interest. It remains challenging primarily due to the small size of the defect relative to the large surface and the rare occurrence of defects. To address this problem, in this article we propose a novel machine vision approach for automatically identifying the tiny flaws that may appear in a single image. First, the presented defect exaggeration approach produces both the flawless image and the corresponding exaggerated version of the defect by taking the variations in the image as regularization terms. Second, a generative adversarial network (GAN) in conjunction with a convolutional neural network (CNN) is proposed to guarantee the accuracy of tiny surface defect detection by producing exaggerated defect image samples. Furthermore, the limited dataset of the training samples for defect detection is enlarged by exploiting the GAN technique with the variation exaggerated images. To evaluate the performance of our proposed method, we conduct comparison experiments between the state-of-the-art techniques with and without the proposed algorithm as well as comparison experiments between the state-of-the-art techniques and our method. The experimental results on different types of surface image samples demonstrate that the proposed method can significantly improve the performance of the state-of-the-art approaches while achieving a defect detection accuracy of 99.2%. Jian Lian, Weikuan Jia, Masoumeh Zareapoor, Yuanjie Zheng, Deepak Kumar Jain 0001, Neeraj Kumar 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2020 | Segmenting Diabetic Retinopathy Lesions in Multispectral Images Using Low-Dimensional Spatial-Spectral Matrix RepresentationabstractMultispectral imaging (MSI) provides a sequence of en-face fundus spectral slices and allows for the examination of structures and signatures throughout the thickness of retina to characterize diabetic retinopathy (DR) lesions comprehensively. Manual interpretation of MSI images is commonly conducted by qualitatively analyzing both the spatial and spectral properties of multiple spectral slices. Meanwhile, there exist few computer-based algorithms that can effectively exploit the spatial and spectral information of MSI images for the diagnosis of DR. We propose a new approach that can quantify the spatial-spectral features of MSI retinal images for automatic DR lesion segmentation. It combines a generalized low-rank approximation of matrices with a supervised regularization term to generate low-dimensional spatial-spectral representations using the feature vectors in all spectral slices. Experimental results showed that the proposed approach is very effective for the segmentation of DR lesions in MSI images, which suggests it as an interesting tool for assisting ophthalmologists in diagnosing, analyzing, and managing DR lesions in MSI. Wanzhen Jiao, Yunfeng Shi, Jian Lian, Bojun Zhao, Yue Min Zhu, Yuanjie Zheng |
IEEE J. Biomed. Health Informatics | 8 |
| 2019 | A novel reconstructed training-set SVM with roulette cooperative coevolution for financial time series classification
Chao Luo 0001, Yuanjie Zheng |
Expert Syst. Appl. | 3 |
| 2019 | Efficient clustering approach for adaptive unsupervised colour image segmentationabstractThis study proposes a clustering‐based colour image segmentation approach consisting of a novel initialisation technique. Colour image segmentation transforms image pixels into regions and a prerequisite for image analysis and computer vision applications. Therefore, colour image segmentation is considered one of the most important processes in image understanding and pattern recognition. This study presents an efficient and adaptive unsupervised approach based on bottom‐up red–green–blue (RGB) colour histogram search approach to achieve colour image segmentation. Firstly, the RGB histogram is processed through a double‐scan procedure to determine significant modes in each histogram. In the next step, each mode is processed through a bottom‐up histogram search approach, completing RGB triplet. The RGB triplets are utilised as the cluster centroids, clustering the pixels into regions and producing the final segmented image. The authors proposed method was compared with several other unsupervised image segmentation algorithms with an extensive experiment performed on various image segmentation evaluation benchmarks. Experimental results show that the proposed algorithm outperforms state‐of‐the‐art algorithms both in terms of features integrity and execution speed. Zubair Khan, Jie Yang 0002, Yuanjie Zheng |
IET Image Process. | 3 |
| 2019 | Long-term prediction of time series based on stepwise linear division algorithm and time-variant zonary fuzzy information granules
Chao Luo 0001, Chenhao Tan, Yuanjie Zheng |
Int. J. Approx. Reason. | 3 |
| 2019 | Whale optimized mixed kernel function of support vector machine for colorectal cancer diagnosis
Hong Liu 0013, Yuanjie Zheng, Dianjie Lu, Chen Lyu 0001 |
J. Biomed. Informatics | 3 |
| 2019 | A novel optimized GA-Elman neural network algorithm
Weikuan Jia, Dean Zhao, Yuanjie Zheng, Sujuan Hou |
Neural Comput. Appl. | 3 |
| 2019 | Ocular multi-spectral imaging deblurring via regularization of mutual information
Guoqiang Ren, Jian Lian, Zheng Xu 0001, Mingqu Fan, Yuanjie Zheng |
Pattern Recognit. Lett. | 5 |
| 2018 | Deep Propagation Based Image MattingabstractIn this paper, we propose a deep propagation based image matting framework by introducing deep learning into learning an alpha matte propagation principal. Our deep learning architecture is a concatenation of a deep feature extraction module, an affinity learning module and a matte propagation module. These three modules are all differentiable and can be optimized jointly via an end-to-end training process. Our framework results in a semantic-level pairwise similarity of pixels for propagation by learning deep image representations adapted to matte propagation. It combines the power of deep learning and matte propagation and can therefore surpass prior state-of-the-art matting techniques in terms of both accuracy and training complexity, as validated by our experimental results from 243K images created based on two benchmark matting databases. Yu Wang 0228, Peiyong Duan, Jianwei Lin, Yuanjie Zheng |
IJCAI | 5 |
| 2018 | Deep Random Walk for Drusen Segmentation from Fundus Images
Fang Yan 0003, Jia Cui, Yu Wang 0228, Hong Liu 0013, Hui Liu 0007, Benzheng Wei, Yilong Yin, Yuanjie Zheng |
MICCAI (2) | 8 |
| 2018 | Deblurring retinal optical coherence tomography via a convolutional neural network with anisotropic and double convolution layerabstractVarious image pre‐processing tasks in optical coherence tomography (OCT) systems involve reversing degradation effects (e.g. deblurring). Current deblurring research mainly focuses on how to build suitable degradation models using deconvolution operators. However, model‐based solutions may not work well in many scenarios. To solve this problem, the authors propose a non‐model architecture, called a deep convolutional neural network, to address parameter‐free situations. The proposed solution employs a deep learning strategy to bridge the gap between traditional model‐based methods and neural network architectures. Experiments on retinal OCT images demonstrate that the proposed approach achieves superior performance compared with the state‐of‐the‐art model‐based OCT deblurring methods. Jian Lian, Sujuan Hou, Xiaodan Sui, Fangzhou Xu, Yuanjie Zheng |
IET Comput. Vis. | 5 |
| 2018 | Graph cut based automatic aorta segmentation with an adaptive smoothness constraint in 3D abdominal CT images
Xiang Deng 0002, Yuanjie Zheng, Xiaoming Xi, Yilong Yin |
Neurocomputing | 2 |
| 2018 | Classifying advertising video by topicalizing high-level semantic concepts
Sujuan Hou, Shangbo Zhou, Wenjie Liu 0001, Yuanjie Zheng |
Multim. Tools Appl. | 4 |
| 2018 | Multiscale Rotation-Invariant Convolutional Neural Networks for Lung Texture ClassificationabstractWe propose a new multiscale rotation-invariant convolutional neural network (MRCNN) model for classifying various lung tissue types on high-resolution computed tomography. MRCNN employs Gabor-local binary pattern that introduces a good property in image analysis-invariance to image scales and rotations. In addition, we offer an approach to deal with the problems caused by imbalanced number of samples between different classes in most of the existing works, accomplished by changing the overlapping size between the adjacent patches. Experimental results on a public interstitial lung disease database show a superior performance of the proposed method to state of the art. Qiangchang Wang, Yuanjie Zheng, Gongping Yang 0001, Weidong Jin, Xinjian Chen 0001, Yilong Yin |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Learning Deep Match Kernels for Image-Set ClassificationabstractImage-set classification has recently generated great popularity due to its widespread applications in computer vision. The great challenges arise from effectively and efficiently measuring the similarity between image sets with high inter-class ambiguity and huge intra-class variability. In this paper, we propose deep match kernels (DMK) to directly measure the similarity between image sets in the match kernel framework. Specifically, we build deep local match kernels between images upon arc-cosine kernels, which can faithfully characterize the similarity between images by mimicking deep neural networks, we introduce anchors to aggregate those deep local match kernels into a global match kernel between image sets, which is learned in a supervised way by kernel alignment and therefore more discriminative. The DMK provides the first match kernel framework for image-set classification, which removes specific assumptions usually required in previous approaches and is computationally more efficient. We conduct extensive experiments on four datasets for three diverse image-set classification tasks. The DMK achieves high performance and consistently surpasses state-of-the-art methods, showing its great effectiveness for image-set classification. Haoliang Sun, Xiantong Zhen, Yuanjie Zheng, Gongping Yang 0001, Yilong Yin, Shuo Li 0001 |
CVPR | 3 |
| 2017 | Choroid segmentation from Optical Coherence Tomography with graph-edge weights learned from deep convolutional neural networks
Xiaodan Sui, Yuanjie Zheng, Benzheng Wei, Hongsheng Bi, Xuemei Pan, Yilong Yin, Shaoting Zhang 0001 |
Neurocomputing | 2 |
| 2017 | Corrigendum to "Hierarchical retinal blood vessel segmentation based on feature and ensemble learning" [Neurocomputing 149 (2015) 708-717]
Shuangling Wang, Yilong Yin, Guibao Cao, Benzheng Wei, Yuanjie Zheng, Gongping Yang 0001 |
Neurocomputing | 5 |
| 2017 | Guest Editorial: Special issue on advances in computing techniques for big medical image data
Yuanjie Zheng, Shaoting Zhang 0001, Junzhou Huang, Tom Weidong Cai |
Neurocomputing | 1 |
| 2017 | Multi-layer multi-view topic model for classifying advertising video
Sujuan Hou, Ling Chen 0006, Dacheng Tao, Shangbo Zhou, Wenjie Liu 0001, Yuanjie Zheng |
Pattern Recognit. | 6 |
| 2017 | Scalable Mammogram Retrieval Using Composite Anchor Graph Hashing With Iterative QuantizationabstractContent-based image retrieval (CBIR) shows great significance in clinical decision-making, which explores the visual content of medical images rather than keywords, tags, or descriptions. It provides doctors an image-guided approach to explore relevant cases that could offer doctors instructive reference. Mammogram screening has been known to be widely used in the early stage diagnosis of breast cancer and could reduce its morbidity and mortality. In this paper, we aim to develop a scalable CBIR method for a large repository of mammogram. To this end, we extend the original Anchor Graph Hashing (AGH) and propose a new unsupervised hashing algorithm, named as composite AGH with iterative quantization (C-AGH-ITQ), which compresses mammographic regions of interest (ROIs) into compact binary codes and enables real-time searching in Hamming space. Multimodal features and different distance metrics are integrated, performing upon a composite Anchor Graph. To improve the effectiveness of the hash code, quantization error is further iteratively minimized by introducing an orthogonal rotation matrix. We evaluate the presented C-AGH-ITQ algorithm on a data set of 11 533 mammographic ROIs obtained from the Digital Database for Screening Mammography. Our method obtains more than 84% retrieval precision and 93% classification accuracy (using$k$NN prediction), which demonstrates that hash codes produced by C-AGH-ITQ well capture the visual similarities between mammographic images. In addition, since C-AGH-ITQ ensures linear complexity of the training procedure and constant time for query, our system is readily applicable to large-scale mammogram databases and has the potential to provide abundant clinical cases as reference. Jingjing Liu 0001, Shaoting Zhang 0001, Wei Liu 0005, Cheng Deng 0002, Yuanjie Zheng, Dimitris N. Metaxas |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Mammographic Mass Segmentation with Online Learned Shape and Appearance Priors
Menglin Jiang, Shaoting Zhang 0001, Yuanjie Zheng, Dimitris N. Metaxas |
MICCAI (2) | 3 |
| 2015 | Hierarchical retinal blood vessel segmentation based on feature and ensemble learning
Shuangling Wang, Yilong Yin, Guibao Cao, Benzheng Wei, Yuanjie Zheng, Gongping Yang 0001 |
Neurocomputing | 5 |
| 2014 | Landmark matching based retinal image alignment by enforcing sparsity in correspondence matrix
Yuanjie Zheng, Ebenezer Daniel, Allan A. Hunter III, Rui Xiao 0001, Jianbin Gao, Hongsheng Li 0001, Maureen G. Maguire, David H. Brainard, James C. Gee |
Medical Image Anal. | 1 |
| 2014 | Solving a Special Type of Jigsaw Puzzles: Banknote Reconstruction From a Large Number of FragmentsabstractIn this paper, we propose a method to solve a special type of jigsaw puzzles, reconstructing banknotes from a large number of fragments based on fragments' images. Existing jigsaw puzzle assembly algorithms have difficulty solving this problem effectively. A main limitation of these methods is that they do not leverage the following important observations: 1) an intact banknote's image is known and thus can be used as prior information; 2) if two aligned fragments overlap each other, they must not be from a same banknote. Based on these two important observations, a three-step method is proposed to reconstruct banknotes from their fragments. Each fragment is first aligned to its original position on the banknote by a RANSAC method. After evaluating every two aligned fragments' relationships, all fragments are embedded into a lower dimensional space and then clustered into small groups using a modified agglomerative clustering method. Fragments in a same cluster are likely to be from a same banknote. Experiments on both synthetic and real data demonstrate the effectiveness of our proposed method. Hongsheng Li 0001, Yuanjie Zheng, Shaoting Zhang 0001, Jian Cheng 0003 |
IEEE Trans. Multim. | 2 |
| 2013 | Optic Disc and Cup Segmentation from Color Fundus Photograph Using Graph Cut with Priors
Yuanjie Zheng, Dwight Stambolian, Joan O'Brien, James C. Gee |
MICCAI (2) | 1 |
| 2013 | A Generative Model for OCT Retinal Layer Segmentation by Integrating Graph-Based Multi-surface Searching and Image Registration
Yuanjie Zheng, Rui Xiao 0001, James C. Gee |
MICCAI (1) | 1 |
| 2013 | Single-Image Vignetting Correction from Gradient Distribution SymmetriesabstractWe present novel techniques for single-image vignetting correction based on symmetries of two forms of image gradients: semicircular tangential gradients (SCTG) and radial gradients (RG). For a given image pixel, an SCTG is an image gradient along the tangential direction of a circle centered at the presumed optical center and passing through the pixel. An RG is an image gradient along the radial direction with respect to the optical center. We observe that the symmetry properties of SCTG and RG distributions are closely related to the vignetting in the image. Based on these symmetry properties, we develop an automatic optical center estimation algorithm by minimizing the asymmetry of SCTG distributions, and also present two methods for vignetting estimation based on minimizing the asymmetry of RG distributions. In comparison to prior approaches to single-image vignetting correction, our methods do not rely on image segmentation and they produce more accurate results. Experiments show our techniques to work well for a wide range of images while achieving a speed-up of 3-5 times compared to a state-of-the-art method. Yuanjie Zheng, Stephen Lin 0001, Sing Bing Kang, Rui Xiao 0001, James C. Gee, Chandra Kambhamettu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Adaptive Multi-cluster Fuzzy C-Means Segmentation of Breast Parenchymal Tissue in Digital Mammography
Brad M. Keller, Diane Nathan, Yuanjie Zheng, James C. Gee, Emily F. Conant, Despina Kontos |
MICCAI (3) | 4 |
| 2010 | Estimation of image bias field with sparsity constraintsabstractWe propose a new scheme to estimate image bias field through introducing two sparsity constraints. One is that the bias-free image has concise representation with image gradients or coefficients of other image transformations. The other constraint is that model fit on the bias field should be as concise as possible. The new scheme enables adaptive specifications of the estimated bias field's smoothness, and results in extremely accurate solutions with more efficient optimization techniques, e.g. linear programming. These distinguish our approaches from many previous methods. Our techniques can be applied to intensity inhomogeneity correction of medical images, illumination and vignetting estimation of images captured by digital cameras. Yuanjie Zheng, James C. Gee |
CVPR | 1 |
| 2010 | Sparse Unbiased Analysis of Anatomical Variance in Longitudinal Imaging
Brian B. Avants, Philip A. Cook, Corey McMillan, Murray Grossman, Nicholas J. Tustison, Yuanjie Zheng, James C. Gee |
MICCAI (1) | 6 |
| 2010 | N4ITK: Improved N3 Bias CorrectionabstractA variant of the popular nonparametric nonuniform intensity normalization (N3) algorithm is proposed for bias field correction. Given the superb performance of N3 and its public availability, it has been the subject of several evaluation studies. These studies have demonstrated the importance of certain parameters associated with the B-spline least-squares fitting. We propose the substitution of a recently developed fast and robust B-spline approximation routine and a modified hierarchical optimization scheme for improved bias field correction over the original N3 algorithm. Similar to the N3 algorithm, we also make the source code, testing, and technical documentation of our contribution, which we denote as "N4ITK," available to the public through the Insight Toolkit of the National Institutes of Health. Performance assessment is demonstrated using simulated data from the publicly available Brainweb database, hyperpolarized (3)He lung image data, and 9.4T postmortem hippocampus data. Nicholas J. Tustison, Brian B. Avants, Philip A. Cook, Yuanjie Zheng, Alexander Egan, Paul A. Yushkevich, James C. Gee |
IEEE Trans. Medical Imaging | 4 |
| 2009 | Single-image optical center estimation from vignetting and tangential gradient symmetryabstractIn this paper, we propose a method for estimating the optical center of a camera given only a single image with vignetting. This is accomplished by identifying the center of the vignetting effect in the image through an analysis of semicircular tangential gradients (SCTGs). For a given image pixel, the SCTG is the image gradient along the tangential direction of the circle centered at the currently estimated optical center and passing through the pixel. We show that for natural images with vignetting, the distribution of SCTGs is generally symmetric if the optical center is estimated accurately, but is skewed otherwise. By minimizing the asymmetry of the SCTG distribution with nonlinear optimization, our method is able to obtain reliable estimates of the optical center. Experiments on simulated and real vignetting images demonstrate the effectiveness of this technique. Yuanjie Zheng, Chandra Kambhamettu, Stephen Lin 0001 |
CVPR | 1 |
| 2009 | Learning based digital mattingabstractWe cast some new insights into solving the digital matting problem by treating it as a semi-supervised learning task in machine learning. A local learning based approach and a global learning based approach are then produced, to fit better the scribble based matting and the trimap based matting, respectively. Our approaches are easy to implement because only some simple matrix operations are needed. They are also extremely accurate because they can efficiently handle the nonlinear local color distributions by incorporating the kernel trick, that are beyond the ability of many previous works. Our approaches can outperform many recent matting methods, as shown by the theoretical analysis and comprehensive experiments. The new insights may also inspire several more works. Yuanjie Zheng, Chandra Kambhamettu |
ICCV | 1 |
| 2009 | Automatic Correction of Intensity Nonuniformity from Sparseness of Gradient Distribution in Medical Images
Yuanjie Zheng, Murray Grossman, Suyash P. Awate, James C. Gee |
MICCAI (1) | 1 |
| 2009 | Single-Image Vignetting CorrectionabstractIn this paper, we propose a method for robustly determining the vignetting function given only a single image. Our method is designed to handle both textured and untextured regions in order to maximize the use of available information. To extract vignetting information from an image, we present adaptations of segmentation techniques that locate image regions with reliable data for vignetting estimation. Within each image region, our method capitalizes on the frequency characteristics and physical properties of vignetting to distinguish it from other sources of intensity variation. Rejection of outlier pixels is applied to improve the robustness of vignetting estimation. Comprehensive experiments demonstrate the effectiveness of this technique on a broad range of images with both simulated and natural vignetting effects. Causes of failures using the proposed algorithm are also analyzed. Yuanjie Zheng, Stephen Lin 0001, Chandra Kambhamettu, Jingyi Yu 0001, Sing Bing Kang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | FuzzyMatte: A computationally efficient scheme for interactive mattingabstractIn this paper, we propose an online interactive matting algorithm, which we call FuzzyMatte. Our framework is based on computing the fuzzy connectedness (FC) [20] from each unknown pixel to the known foreground and background. FC effectively captures the adjacency and similarity between image elements and can be efficiently computed using the strongest connected path searching algorithm. The final alpha value at each pixel can then be calculated from its FC. While many previous methods need to completely recompute the matte when new inputs are provided, FuzzyMatte effectively integrates these new inputs with the previously estimated matte by efficiently recomputing the FC value for a small subset of pixels. Thus, the computational overhead between each iteration of the refinement is significantly reduced. We demonstrate FuzzyMatte on a wide range of images. We show that FuzzyMatte updates the matte in an online interactive setting and generates high quality matte for complex images. Yuanjie Zheng, Chandra Kambhamettu, Jingyi Yu 0001, Thomas L. Bauer, Karl V. Steiner |
CVPR | 1 |
| 2008 | Single-image vignetting correction using radial gradient symmetryabstractIn this paper, we present a novel single-image vignetting method based on the symmetric distribution of the radial gradient (RG). The radial gradient is the image gradient along the radial direction with respect to the image center. We show that the RG distribution for natural images without vignetting is generally symmetric. However, this distribution is skewed by vignetting. We develop two variants of this technique, both of which remove vignetting by minimizing asymmetry of the RG distribution. Compared with prior approaches to single-image vignetting correction, our method does not require segmentation and the results are generally better. Experiments show our technique works for a wide range of images and it achieves a speed-up of 4’5 times compared with a state-of-the-art method. Yuanjie Zheng, Jingyi Yu 0001, Sing Bing Kang, Stephen Lin 0001, Chandra Kambhamettu |
CVPR | 1 |
| 2008 | Estimation of Ground-Glass Opacity Measurement in CT Lung Images
Yuanjie Zheng, Chandra Kambhamettu, Thomas L. Bauer, Karl V. Steiner |
MICCAI (2) | 1 |
| 2007 | Lung Nodule Growth Analysis from 3D CT Data with a Coupled Segmentation and Registration FrameworkabstractIn this paper we propose a new framework to simultaneously segment and register lung and tumor in serial CT data. Our method assumes nonrigid transformation on lung deformation and rigid structure on the tumor. We use the B- Spline-based nonrigid transformation to model the lung deformation while imposing rigid transformation on the tumor to preserve the volume and the shape of the tumor. In particular, we set the control points within the tumor to form a control mesh and thus assume the tumor region follows the same rigid transformation as the control mesh. For segmentation, we apply a 2D graph-cut algorithm on the 3D lung and tumor datasets. By iteratively performing segmentation and registration, our method achieves highly accurate segmentation and registration on serial CT data. Finally, since our method eliminates the possible volume variations of the tumor during registration, we can further estimate accurately the tumor growth, an important evidence in lung cancer diagnosis. Initial experiments on five sets of patients ' serial CT data show that our method is robust and reliable. Yuanjie Zheng, Karl V. Steiner, Thomas L. Bauer, Jingyi Yu 0001, Dinggang Shen, Chandra Kambhamettu |
ICCV | 1 |
| 2007 | Segmentation and Classification of Breast Tumor Using Dynamic Contrast-Enhanced MR Images
Yuanjie Zheng, Sajjad Baloch, Sarah Englander, Mitchell D. Schnall, Dinggang Shen |
MICCAI (2) | 1 |
| 2007 | De-enhancing the Dynamic Contrast-Enhanced Breast MRI for Robust Registration
Yuanjie Zheng, Jingyi Yu 0001, Chandra Kambhamettu, Sarah Englander, Mitchell D. Schnall, Dinggang Shen |
MICCAI (1) | 1 |
| 2006 | Single-Image Vignetting CorrectionabstractIn this paper, we propose a method for determining the vignetting function given only a single image. Our method is designed to handle both textured and untextured regions in order to maximize the use of available information. To extract vignetting information from an image, we present adaptations of segmentation techniques that locate image regions with reliable data for vignetting estimation. Within each image region, our method capitalizes on frequency characteristics and physical properties of vignetting to distinguish it from other sources of intensity variation. The vignetting data acquired from regions are weighted according to a presented reliability measure to promote robustness in estimation. Comprehensive experiments demonstrate the effectiveness of this technique on a broad range of images. Yuanjie Zheng, Stephen Lin 0001, Sing Bing Kang |
CVPR (1) | 1 |
| 2004 | Unsupervised Segmentation on Image with JSEG Using Soft Class Map
Yuanjie Zheng, Jie Yang 0002, Yue Zhou 0005 |
IDEAL | 1 |
| 2004 | Unsupervised Image Segmentation with Fuzzy Connectedness
Yuanjie Zheng, Jie Yang 0002, Yue Zhou 0005 |
PRICAI | 1 |