Jin-Gang Yu

dblp:142/2835 · DBLP profile ↗
← Back
37ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-2148-2726ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 DC-RRG: Diagnosis-centered cascaded radiology report generation
Zinan Hong, Zijian Zhou 0002, Miaojing Shi, Jin-Gang Yu, Jingping Yun, Shuangping Huang
Expert Syst. Appl.5
2026 REST: Holistic Learning for End-to-End Semantic Segmentation of Whole-Scene Remote Sensing Imagery
abstract
Semantic segmentation of remote sensing imagery (RSI) is a fundamental task that aims at assigning a category label to each pixel. To pursue precise segmentation with one or more fine-grained categories, semantic segmentation often requires holistic segmentation of whole-scene RSI (WRI), which is normally characterized by a large size. However, conventional deep learning methods struggle to handle holistic segmentation of WRI due to the memory limitations of the graphics processing unit (GPU), thus requiring to adopt suboptimal strategies such as cropping or fusion, which result in performance degradation. Here, we introduce the Robust End-to-end semantic Segmentation architecture for whole-scene remoTe sensing imagery (REST). REST is the first intrinsically endtoend framework for truly holistic segmentation of WRI, supporting a wide range of encoders and decoders in a plugandplay fashion. It enables seamless integration with mainstream semantic segmentation methods, and even more advanced foundation models. Specifically, we propose a novel spatial parallel interaction mechanism (SPIM) within REST to overcome GPU memory constraints and achieve global context awareness. Unlike traditional parallel methods, SPIM enables REST to process a WRI effectively and efficiently by combining parallel computation with a divideandconquer strategy. Both theoretical analysis and experiments demonstrate that REST attains nearlinear throughput scalability as additional GPUs are employed. Extensive experiments demonstrate that REST consistently outperforms existing cropping-based and fusion-based methods across a variety of scenarios, ranging from single-class to multi-class segmentation, from multispectral to hyperspectral imagery, and from satellite to drone platforms. The robustness and versatility of REST are expected to offer a promising solution for the holistic segmentation of WRI, with the potential for further extension to large-size medical imagery segmentation.
Wei Chen 0089, Lorenzo Bruzzone, Bo Dang 0002, Yuan Gao 0015, Youming Deng, Jin-Gang Yu, Liangqi Yuan, Yansheng Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Bridging Knowledge Discrepancy in Retinal Image Analysis Through Federated Multi-task Learning
Jing Yang 0046, Jin-Gang Yu, Feng Gao 0023, Shuting Yang, Du Cai, Jiacheng Wang 0002, Liansheng Wang 0002
MICCAI (14)3
2025 Few-Shot Learning for Annotation-Efficient Nucleus Instance Segmentation
abstract
Nucleus instance segmentation from histopathology images suffers from the extremely laborious and expert-dependent annotation of nucleus instances. As a promising solution to this task, annotation-efficient deep learning paradigms have recently attracted much research interest, such as weakly-/semi-supervised learning, generative adversarial learning, etc. In this paper, we propose to formulate annotation-efficient nucleus instance segmentation from the perspective of few-shot learning (FSL). Our work was motivated by that, with the prosperity of computational pathology, an increasing number of fully-annotated datasets are publicly accessible, and we hope to leverage these external datasets to assist nucleus instance segmentation on the target dataset which only has very limited annotation. To achieve this goal, we adopt the meta-learning based FSL paradigm, which however has to be tailored in two substantial aspects before adapting to our task. First, since the novel classes may be inconsistent with those of the external dataset, we extend the basic definition of few-shot instance segmentation (FSIS) to generalized few-shot instance segmentation (GFSIS). Second, to cope with the intrinsic challenges of nucleus segmentation, including touching between adjacent cells, cellular heterogeneity, etc., we further introduce a structural guidance mechanism into the GFSIS network, finally leading to a unified Structurally-Guided Generalized Few-Shot Instance Segmentation (SGFSIS) framework. Extensive experiments on a couple of publicly accessible datasets demonstrate that, SGFSIS can outperform other annotation-efficient learning baselines, including semi-supervised learning, simple transfer learning, etc., with comparable performance to fully supervised learning with around 10% annotations.
Zihao Wu 0004, Jie Yang 0002, Danyi Li, Yuan Gao 0015, Changxin Gao, Gui-Song Xia, Yuanqing Li 0001, Jin-Gang Yu
IEEE Trans. Medical Imaging10
2024 Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference Cost
abstract
We aim at exploiting additional auxiliary labels from an independent (auxiliary) task to boost the primary task performance which we focus on, while preserving a single task inference cost of the primary task. While most existing auxiliary learning methods are optimization-based relying on loss weights/gradients manipulation, our method is architecture-based with a flexible asymmetric structure for the primary and auxiliary tasks, which produces different networks for training and inference. Specifically, starting from two single task networks/branches (each representing a task), we propose a novel method with evolving networks where only primary-to-auxiliary links exist as the cross-task connections after convergence. These connections can be removed during the primary task inference, resulting in a single-task inference cost. We achieve this by formulating a Neural Architecture Search (NAS) problem, where we initialize bi-directional connections in the search space and guide the NAS optimization converging to an architecture with only the single-side primary-to-auxiliary connections. Moreover, our method can be incorporated with optimization-based auxiliary learning approaches. Extensive experiments with six tasks on NYU v2, CityScapes, and Taskonomy datasets using VGG, ResNet, and ViT backbones validate the promising performance. The codes are available at https://github.com/ethanygao/Aux-NAS.
Yuan Gao 0015, Wenhan Luo, Lin Ma 0002, Jin-Gang Yu, Gui-Song Xia, Jiayi Ma 0001
ICLR5
2024 DMTG: One-Shot Differentiable Multi-Task Grouping
abstract
We aim to address Multi-Task Learning (MTL) with a large number of tasks by Multi-Task Grouping (MTG). Given $N$ tasks, we propose to simultaneously identify the best task groups from $2^N$ candidates and train the model weights simultaneously in one-shot, with the high-order task-affinity fully exploited. This is distinct from the pioneering methods which sequentially identify the groups and train the model weights, where the group identification often relies on heuristics. As a result, our method not only improves the training efficiency, but also mitigates the objective bias introduced by the sequential procedures that potentially leads to a suboptimal solution. Specifically, we formulate MTG as a fully differentiable pruning problem on an adaptive network architecture determined by an unknown Categorical distribution. To categorize $N$ tasks into $K$ groups (represented by $K$ encoder branches), we initially set up $KN$ task heads, where each branch connects to all $N$ task heads to exploit the high-order task-affinity. Then, we gradually prune the $KN$ heads down to $N$ by learning a relaxed differentiable Categorical distribution, ensuring that each task is exclusively and uniquely categorized into only one branch. Extensive experiments on CelebA and Taskonomy datasets with detailed ablations show the promising performance and efficiency of our method. The codes are available at https://github.com/ethanygao/DMTG.
Yuan Gao 0015, Shuguo Jiang, Moran Li, Jin-Gang Yu, Gui-Song Xia
ICML4
2024 Learning to Holistically Detect Bridges From Large-Size VHR Remote Sensing Imagery
abstract
Bridge detection in remote sensing images (RSIs) plays a crucial role in various applications, but it poses unique challenges compared to the detection of other objects. In RSIs, bridges exhibit considerable variations in terms of their spatial scales and aspect ratios. Therefore, to ensure the visibility and integrity of bridges, it is essential to perform holistic bridge detection in large-size very-high-resolution (VHR) RSIs. However, the lack of datasets with large-size VHR RSIs limits the deep learning algorithms' performance on bridge detection. Due to the limitation of GPU memory in tackling large-size images, deep learning-based object detection methods commonly adopt the cropping strategy, which inevitably results in label fragmentation and discontinuous prediction. To ameliorate the scarcity of datasets, this paper proposes a large-scale dataset named GLH-Bridge comprising 6,000 VHR RSIs sampled from diverse geographic locations across the globe. These images encompass a wide range of sizes, varying from 2,048 × 2,048 to 16,384 × 16,384 pixels, and collectively feature 59,737 bridges. These bridges span diverse backgrounds, and each of them has been manually annotated, using both an oriented bounding box (OBB) and a horizontal bounding box (HBB). Furthermore, we present an efficient network for holistic bridge detection (HBD-Net) in large-size RSIs. The HBD-Net presents a separate detector-based feature fusion (SDFF) architecture and is optimized via a shape-sensitive sample re-weighting (SSRW) strategy. The SDFF architecture performs inter-layer feature fusion (IFF) to incorporate multi-scale context in the dynamic image pyramid (DIP) of the large-size image, and the SSRW strategy is employed to ensure an equitable balance in the regression weight of bridges with various aspect ratios. Based on the proposed GLH-Bridge dataset, we establish a bridge detection benchmark including the OBB and HBB tasks, and validate the effectiveness of the proposed HBD-Net. Additionally, cross-dataset generalization experiments on two publicly available datasets illustrate the strong generalization capability of the GLH-Bridge dataset.
Yansheng Li 0001, Yongjun Zhang 0002, Yihua Tan, Jin-Gang Yu, Song Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Complete Instances Mining for Weakly Supervised Instance Segmentation
abstract
Weakly supervised instance segmentation (WSIS) using only image-level labels is a challenging task due to the difficulty of aligning coarse annotations with the finer task. However, with the advancement of deep neural networks (DNNs), WSIS has garnered significant attention. Following a proposal-based paradigm, we encounter a redundant segmentation problem resulting from a single instance being represented by multiple proposals. For example, we feed a picture of a dog and proposals into the network and expect to output only one proposal containing a dog, but the network outputs multiple proposals. To address this problem, we propose a novel approach for WSIS that focuses on the online refinement of complete instances through the use of MaskIoU heads to predict the integrity scores of proposals and a Complete Instances Mining (CIM) strategy to explicitly model the redundant segmentation problem and generate refined pseudo labels. Our approach allows the network to become aware of multiple instances and complete instances, and we further improve its robustness through the incorporation of an Anti-noise strategy. Empirical evaluations on the PASCAL VOC 2012 and MS COCO datasets demonstrate that our method achieves state-of-the-art performance with a notable margin. Our implementation will be made available at https://github.com/ZechengLi19/CIM.
Zening Zeng, Jin-Gang Yu
IJCAI4
2023 Prototypical multiple instance learning for predicting lymph node metastasis of breast cancer from whole-slide pathological images
Jin-Gang Yu, Zihao Wu 0004, Shule Deng, Yuanqing Li 0001, Caifeng Ou, Chunjiang He, Baiye Wang, Pusheng Zhang
Medical Image Anal.1
2023 Pose-Guided Hierarchical Semantic Decomposition and Composition for Human Parsing
abstract
Human parsing is a fine-grained semantic segmentation task, which needs to understand human semantic parts. Most existing methods model human parsing as a general semantic segmentation, which ignores the inherent relationship among hierarchical human parts. In this work, we propose a pose-guided hierarchical semantic decomposition and composition framework for human parsing. Specifically, our method includes a semantic maintained decomposition and composition (SMDC) module and a pose distillation (PC) module. SMDC progressively disassembles the human body to focus on the more concise regions of interest in the decomposition stage and then gradually assembles human parts under the guidance of pose information in the composition stage. Notably, SMDC maintains the atomic semantic labels during both stages to avoid the error propagation issue of the hierarchical structure. To further take advantage of the relationship of human parts, we introduce pose information as explicit guidance for the composition. However, the discrete structure prediction in pose estimation is against the requirement of the continuous region in human parsing. To this end, we design a PC module to broadcast the maximum responses of pose estimation to form the continuous structure in the way of knowledge distillation. The experimental results on the look-into-person (LIP) and PASCAL-Person-Part datasets demonstrate the superiority of our method compared with the state-of-the-art methods, that is, 55.21% mean Intersection of Union (mIoU) on LIP and 69.88% mIoU on PASCAL-Person-Part.
Changqian Yu, Jin-Gang Yu, Changxin Gao, Nong Sang
IEEE Trans. Cybern.3
2023 Bayesian Collaborative Learning for Whole-Slide Image Classification
abstract
Whole-slide image (WSI) classification is fundamental to computational pathology, which is challenging in extra-high resolution, expensive manual annotation, data heterogeneity, etc. Multiple instance learning (MIL) provides a promising way towards WSI classification, which nevertheless suffers from the memory bottleneck issue inherently, due to the gigapixel high resolution. To avoid this issue, the overwhelming majority of existing approaches have to decouple the feature encoder and the MIL aggregator in MIL networks, which may largely degrade the performance. Towards this end, this paper presents a Bayesian Collaborative Learning (BCL) framework to address the memory bottleneck issue with WSI classification. Our basic idea is to introduce an auxiliary patch classifier to interact with the target MIL classifier to be learned, so that the feature encoder and the MIL aggregator in the MIL classifier can be learned collaboratively while preventing the memory bottleneck issue. Such a collaborative learning procedure is formulated under a unified Bayesian probabilistic framework and a principled Expectation-Maximization algorithm is developed to infer the optimal model parameters iteratively. As an implementation of the E-step, an effective quality-aware pseudo labeling strategy is also suggested. The proposed BCL is extensively evaluated on three publicly available WSI datasets, i.e., CAMELYON16, TCGA-NSCLC and TCGA-RCC, achieving an AUC of 95.6%, 96.0% and 97.5% respectively, which consistently outperforms all the methods compared. Comprehensive analysis and discussion will also be presented for in-depth understanding of the method. To promote future work, our source code is released at: https://github.com/Zero-We/BCL.
Jin-Gang Yu, Zihao Wu 0004, Shule Deng, Qihang Wu, Zhongtang Xiong, Tianyou Yu, Gui-Song Xia, Qingping Jiang, Yuanqing Li 0001
IEEE Trans. Medical Imaging1
2023 Learning Relative Feature Displacement for Few-Shot Open-Set Recognition
abstract
Few-shot learning (FSL) usually assumes that the query is drawn from the same label space as the support set, while queries from unknown classes may emerge unexpectedly in many open-world application scenarios. Such an open-set issue will limit the practical deployment of FSL systems, which remains largely unexplored. In this paper, we investigate the problem of few-shot open-set recognition (FSOR) and propose a novel solution, called Relative Feature Displacement Network (RFDNet), which empowers FSL systems to reject queries from unknown classes while accurately classifying those from known classes. First, we suggest a different relative feature displacement learning (RFDL) paradigm for FSOR, i.e., meta-learning a feature displacement relative to a pretrained reference feature embedding, based on our insightful observations on the randomness drift issue of previous meta-learning based for FSOR methods, as well as the generalization ability of the feature embedding pretrained for general classification. Second, we design the RFDNet framework to implement the RFDL paradigm, which is mainly featured by a task-aware RFD generator and a marginal open-set loss. Comprehensive experiments on three public datasets, i.e., miniImageNet, CIFAR-FS and tieredImageNet, demonstrate that RFDNet can consistently outperform the state-of-the-art methods, achieving improvement of 5.2%, 2.0% and 1.7% respectively, in terms of AUROC for unknown-class rejection under the 5-way 5-shot setting.
Shule Deng, Jin-Gang Yu, Zihao Wu 0004, Hongxia Gao, Yansheng Li 0001, Yang Yang 0066
IEEE Trans. Multim.2
2023 Efficient Spatio-Temporal Contrastive Learning for Skeleton-Based 3-D Action Recognition
abstract
In this paper, we propose a simple yet effective self-supervised method called spatio-temporal contrastive learning (ST-CL) for 3D skeleton-based action recognition. ST-CL acquires action-specific features by regarding the spatio-temporal continuity of motion tendency as the supervisory signal. To yield effective representations, ST-CL first designs some novel contrastive proxy tasks by providing different spatio-temporal observation scenes for the same 3D action and pulling them together in the embedding space. Second, three key components are devised in the action encoding to efficiently extract representations in contrastive tasks: (1) Information Representation introduces the awareness of joint type when analyzing motion dynamics. (2) Non-local GCN learns a data-driven graph topology structure and promotes a spatial message passing among long-range joints in each frame. (3) Multi-Scale TCN makes larger receptive fields for capturing richer longe-range temporal dynamics amomg adjacent frames. In ST-CL, these effective proxy tasks yield useful representations and efficient action encoding further enhances the representation capacity. As validated on four large-scale datasets, ST-CL is a strong baseline with high performance and efficiency for the contrastive learning study of the skeleton data. Compared to previous self-supervised methods, the proposed ST-CL achieves significant improvement consistently with a smaller model size and better training efficiency.
Xuehao Gao, Yang Yang 0066, Maosen Li, Jin-Gang Yu, Shaoyi Du
IEEE Trans. Multim.5
2021 Deep Graph Matching Under Quadratic Constraint
abstract
Recently, deep learning based methods have demonstrated promising results on the graph matching problem, by relying on the descriptive capability of deep features extracted on graph nodes. However, one main limitation with existing deep graph matching (DGM) methods lies in their ignorance of explicit constraint of graph structures, which may lead the model to be trapped into local minimum in training. In this paper, we propose to explicitly formulate pairwise graph structures as a quadratic constraint incorporated into the DGM framework. The quadratic constraint minimizes the pairwise structural discrepancy between graphs, which can reduce the ambiguities brought by only using the extracted CNN features. Moreover, we present a differentiable implementation to the quadratic constrained-optimization such that it is compatible with the unconstrained deep learning optimizer. To give more precise and proper supervision, a well-designed false matching loss against class imbalance is proposed, which can better penalize the false negatives and false positives with less overfitting. Exhaustive experiments demonstrate that our method achieves competitive performance on real-world datasets. The code is available at: https://github.com/Zerg-Overmind/QC-DGM.
Quankai Gao, Fudong Wang 0001, Nan Xue 0001, Jin-Gang Yu, Gui-Song Xia
CVPR4
2021 Context-Aware Selective Label Smoothing for Calibrating Sequence Recognition Model
abstract
Despite the success of deep neural network (DNN) on sequential data (i.e., scene text and speech) recognition, it suffers from the over-confidence problem mainly due to overfitting in training with the cross-entropy loss, which may make the decision-making less reliable. Confidence calibration has been recently proposed as one effective solution to this problem. Nevertheless, the majority of existing confidence calibration methods aims at non-sequential data, which is limited if directly applied to sequential data since the intrinsic contextual dependency in sequences or the class-specific statistical prior is seldom exploited. To the end, we propose a Context-Aware Selective Label Smoothing (CASLS) method for calibrating sequential data. The proposed CASLS fully leverages the contextual dependency in sequences to construct confusion matrices of contextual prediction statistics over different classes. Class-specific error rates are then used to adjust the weights of smoothing strength in order to achieve adaptive calibration. Experimental results on sequence recognition tasks, including scene text recognition and speech recognition, demonstrate that our method can achieve the state-of-the-art performance.
Shuangping Huang, Yu Luo 0007, Zhenzhou Zhuang, Jin-Gang Yu, Mengchao He, Yongpan Wang
ACM Multimedia4
2021 A simple graph-based semi-supervised learning approach for imbalanced classification
Jianjin Deng, Jin-Gang Yu
Pattern Recognit.2
2021 Learning Deep Cross-Modal Embedding Networks for Zero-Shot Remote Sensing Image Scene Classification
abstract
Due to its wide applications, remote sensing (RS) image scene classification has attracted increasing research interest. When each category has a sufficient number of labeled samples, RS image scene classification can be well addressed by deep learning. However, in the RS big data era, it is extremely difficult or even impossible to annotate RS scene samples for all the categories in one time as the RS scene classification often needs to be extended along with the emergence of new applications that inevitably involve a new class of RS images. Hence, the RS big data era fairly requires a zero-shot RS scene classification (ZSRSSC) paradigm in which the classification model learned from training RS scene categories obeys the inference ability to recognize the RS image scenes from unseen categories, in common with the humans’ evolutionary perception ability. Unfortunately, zero-shot classification is largely unexploited in the RS field. This article proposes a novel ZSRSSC method based on locality-preservation deep cross-modal embedding networks (LPDCMENs). The proposed LPDCMENs, which can fully assimilate the pairwise intramodal and intermodal supervision in an end-to-end manner, aim to alleviate the problem of class structure inconsistency between two hybrid spaces (i.e., the visual image space and the semantic space). To pursue a stable and generalization ability, which is highly desired for ZSRSSC, a set of explainable constraints is specially designed to optimize LPDCMENs. To fully verify the effectiveness of the proposed LPDCMENs, we collect a new large-scale RS scene data set, including the instance-level visual images and class-level semantic representations (RSSDIVCS), where the general and domain knowledge is exploited to construct the class-level semantic representations. Extensive experiments show that the proposed ZSRSSC method based on LPDCMENs can obviously outperform the state-of-the-art methods, and the domain knowledge further improves the performance of ZSRSSC compared with the general knowledge. The collected RSSDIVCS will be made publicly available along with this article.
Yansheng Li 0001, Zhihui Zhu, Jin-Gang Yu, Yongjun Zhang 0002
IEEE Trans. Geosci. Remote. Sens.3
2020 FGN: Fully Guided Network for Few-Shot Instance Segmentation
abstract
Few-shot instance segmentation (FSIS) conjoins the few-shot learning paradigm with general instance segmentation, which provides a possible way of tackling instance segmentation in the lack of abundant labeled data for training. This paper presents a Fully Guided Network (FGN) for few-shot instance segmentation. FGN perceives FSIS as a guided model where a so-called support set is encoded and utilized to guide the predictions of a base instance segmentation network (i.e., Mask R-CNN), critical to which is the guidance mechanism. In this view, FGN introduces different guidance mechanisms into the various key components in Mask R-CNN, including Attention-Guided RPN, Relation-Guided Detector, and Attention-Guided FCN, in order to make full use of the guidance effect from the support set and adapt better to the inter-class generalization. Experiments on public datasets demonstrate that our proposed FGN can outperform the state-of-the-art methods.
Zhibo Fan, Jin-Gang Yu, Jiarong Ou, Changxin Gao, Gui-Song Xia, Yuanqing Li 0001
CVPR2
2020 Zero-Assignment Constraint for Graph Matching With Outliers
abstract
Graph matching (GM), as a longstanding problem in computer vision and pattern recognition, still suffers from numerous cluttered outliers in practical applications. To address this issue, we present the zero-assignment constraint (ZAC) for approaching the graph matching problem in the presence of outliers. The underlying idea is to suppress the matchings of outliers by assigning zero-valued vectors to the potential outliers in the obtained optimal correspondence matrix. We provide elaborate theoretical analysis to the problem, i.e., GM with ZAC, and figure out that the GM problem with and without outliers are intrinsically different, which enables us to put forward a sufficient condition to construct valid and reasonable objective function. Consequently, we design an efficient outlier-robust algorithm to significantly reduce the incorrect or redundant matchings caused by numerous outliers. Extensive experiments demonstrate that our method can achieve the state-of-the-art performance in terms of accuracy and efficiency, especially in the presence of numerous outliers.
Fudong Wang 0001, Nan Xue 0001, Jin-Gang Yu, Gui-Song Xia
CVPR3
2020 Censoring-Aware Deep Ordinal Regression for Survival Prediction from Pathological Images
Lichao Xiao, Jin-Gang Yu, Jiarong Ou, Shule Deng, Zhenhua Yang, Yuanqing Li 0001
MICCAI (5)2
2020 Pose-guided spatiotemporal alignment for video-based person Re-identification
Changxin Gao, Jin-Gang Yu, Nong Sang
Inf. Sci.3
2020 Exemplar-Based Recursive Instance Segmentation With Application to Plant Image Analysis
abstract
Instance segmentation is a challenging computer vision problem which lies at the intersection of object detection and semantic segmentation. Motivated by plant image analysis in the context of plant phenotyping, a recently emerging application field of computer vision, this paper presents the Exemplar-Based Recursive Instance Segmentation (ERIS) framework. A three-layer probabilistic model is firstly introduced to jointly represent hypotheses, voting elements, instance labels and their connections. Afterwards, a recursive optimization algorithm is developed to infer the maximum a posteriori (MAP) solution, which handles one instance at a time by alternating among the three steps of detection, segmentation and update. The proposed ERIS framework departs from previous works mainly in two respects. First, it is exemplar-based and model-free, which can achieve instance-level segmentation of a specific object class given only a handful of (typically less than 10) annotated exemplars. Such a merit enables its use in case that no massive manually-labeled data is available for training strong classification models, as required by most existing methods. Second, instead of attempting to infer the solution in a single shot, which suffers from extremely high computational complexity, our recursive optimization strategy allows for reasonably efficient MAP-inference in full hypothesis space. The ERIS framework is substantialized for the specific application of plant leaf segmentation in this work. Experiments are conducted on public benchmarks to demonstrate the superiority of our method in both effectiveness and efficiency in comparison with the state-of-the-art.
Jin-Gang Yu, Yansheng Li 0001, Changxin Gao, Hongxia Gao, Gui-Song Xia, Zhu Liang Yu, Yuanqing Li 0001
IEEE Trans. Image Process.1
2019 Fine-Grained Classification of Endoscopic Tympanic Membrane Images
abstract
Medical image based diagnosis often requires classification of images at sub-class level, which is essentially a fine-grained visual classification (FGVC) problem. Surprisingly, few prior works have considered this problem from the perspective of FGVC. Motivated by this fact, we present in this paper an FGVC method to boost the classification performance in the context of otitis media diagnosis with endoscopic tympanic membrane images. Our proposed method works in a weakly-supervised fashion, which only takes as input image-level class labels, without the necessity of expensive part annotations. An image-level convolutional neural network (C-NN) is first trained, which can generate saliency maps. The saliency maps can be used to localize discriminative local patches, over which another patch-level CNN can be trained. Both image-level and patch-level CNNs are then integrated for performance boosting. Experiments on real clinical data demonstrate that the proposed method can achieve promising performance.
Lichao Xiao, Jin-Gang Yu, Jiarong Ou
ICIP2
2018 Robust and Efficient Ellipse Fitting Using Tangent Chord Distance
Jiarong Ou, Jin-Gang Yu, Changxin Gao, Lichao Xiao
ACCV (3)2
2018 Spatially Attentive Correlation Filters for Visual Tracking
abstract
Although correlation filter based trackers have recently demonstrated excellent performance, they still suffer from the boundary effects. The cosine window is introduced to alleviate the boundary affects, which however may result in poor performance in case of occlusion or fast motion. To address this problem, we propose a simple yet effective framework, which builds a spatially attentive model with multiple features to guide the detection of the correlation filter based trackers. The proposed method not only can breakthrough the spatial extent of cosine window, but also can provides prior information about the target object. Moreover, to model a robust object prior, we propose a generic strategy for adaptive fusion and update of multiple features. Extensive experiments over multiple tracking benchmarks demonstrate the superior accuracy and real-time performance of our methods compared to the state-of-the-art trackers.
Huai Qin, Zhixiong Pi, Changqian Yu, Changxin Gao, Jin-Gang Yu, Nong Sang
ICIP5
2017 Robust Visual Tracking Using Exemplar-Based Detectors
abstract
Tracking by detection has become an attractive tracking technique, which treats tracking as an object detection problem and trains a detector to separate the target object from the background in each frame. While this strategy is effective to some extent, we argue that the task in tracking should be searching for a specific object instance instead of an object category. Based on this viewpoint, a novel framework based on object exemplar detectors is proposed for visual tracking. To build a specific and discriminative model to separate the object instance from the background, the proposed method trains an exemplar-based linear discriminant analysis (ELDA) classifier for the object exemplar, using the current tracked instance as the positive sample and massive negative samples obtained both offline and online. To improve the trackers' adaptivity, we use an ensemble of the above ELDA detectors and update them during the tracking to cover the variation in object appearance. Extensive experimental results on a large benchmark data set show that the proposed method outperforms many state-of-the-art trackers, demonstrating the effectiveness and robustness of the ELDA tracker.
Changxin Gao, Jin-Gang Yu, Rui Huang 0001, Nong Sang
IEEE Trans. Circuits Syst. Video Technol.3
2016 Temporally aligned pooling representation for video-based person re-identification
abstract
This paper proposes an effective Temporally Aligned Pooling Representation (TAPR) for video-based person re-identification. To extract the motion information from a sequence, we propose to track the superpixels of the lowest portions of human. To perform temporal alignment of videos, we propose to select the “best” walking cycle from the noisy motion information according to the intrinsic periodicity property of walking persons, that is fitted sinusoid in our implementation. To describe the video data in the selected walking cycle, we first divide the cycle into several segments according to the sinusoid, and then describe each segment by temporally aligned pooling. Extensive experimental results on the public datasets demonstrate the effectiveness of the proposed method compared with the state-of-the-art approaches.
Changxin Gao, Jin Wang 0019, Leyuan Liu 0001, Jin-Gang Yu, Nong Sang
ICIP4
2016 A novel spatio-temporal saliency approach for robust dim moving target detection from airborne infrared image sequences
Yansheng Li 0001, Yongjun Zhang 0002, Jin-Gang Yu, Yihua Tan, Jinwen Tian, Jiayi Ma 0001
Inf. Sci.3
2016 Globally consistent correspondence of multiple feature sets using proximal Gauss-Seidel relaxation
Jin-Gang Yu, Gui-Song Xia, Ashok Samal, Jinwen Tian
Pattern Recognit.1
2016 Image retrieval based on image-to-class similarity
Jun Chen 0019, Yong Wang 0036, Linbo Luo 0002, Jin-Gang Yu, Jiayi Ma 0001
Pattern Recognit. Lett.4
2016 A Computational Model for Object-Based Visual Saliency: Spreading Attention Along Gestalt Cues
abstract
The past few years have witnessed impressive progress on the research of salient object detection. Nevertheless , existing approaches still cannot perform satisfactorily in the case of complex scenes, particularly when the salient objects have non- uniform appearance or complicated shapes, and the background is complexly structured. One important reason for such limitations may be that these approaches commonly ignore the factor of perceptual grouping in saliency modeling. To address this issue, this paper presents a novel computational model for object -based visual saliency, which explicitly takes into consideration the connections between attention and perceptual grouping, and incorporates Gestalt grouping cues into saliency computation. Inspired by the sensory enhancement theory, we suggest a paradigm for object-based saliency modeling, that is, object-based saliency stems from spreading attention along Gestalt grouping cues. Computationally , three typical Gestalt cues, including proximity, similarity, and closure, are respectively extracted from the given image, which are then integrated by constructing a unified Gestalt graph. A new algorithm named personalized power iteration clustering is developed to effectively fulfill the spreading of attention information across the Gestalt graph. Intensive experiments have been carried out to demonstrate the superior performance of the proposed model in comparison to the state-of-the-art.
Jin-Gang Yu, Gui-Song Xia, Changxin Gao, Ashok Samal
IEEE Trans. Multim.1
2015 Kernel regression in mixed feature spaces for spatio-temporal saliency detection
Yansheng Li 0001, Yihua Tan, Jin-Gang Yu, Shengxiang Qi, Jinwen Tian
Comput. Vis. Image Underst.3
2015 Salient object detection via contrast information and object vision organization cues
Shengxiang Qi, Jin-Gang Yu, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian
Neurocomputing2
2014 Exemplar-based linear discriminant analysis for robust object tracking
abstract
Tracking-by-detection has become an attractive tracking technique, which treats tracking as a category detection problem. However, the task in tracking is to search for a specific object, rather than an object category as in detection. In this paper, we propose a novel tracking framework based on exemplar detector rather than category detector. The proposed tracker is an ensemble of exemplar-based linear discriminant analysis (ELDA) detectors. Each detector is quite specific and discriminative, because it is trained by a single object instance and massive negatives. To improve its adaptivity, we update both object and background models. Experimental results on several challenging video sequences demonstrate the effectiveness and robustness of our tracking algorithm.
Changxin Gao, Jin-Gang Yu, Rui Huang 0001, Nong Sang
ICIP3
2014 Visual saliency detection using feature activity weighted decorrelation cues
abstract
In this paper, a novel model based on feature activity weighted decorrelation cues is proposed for visual saliency detection in natural images. It consists of two parts: the feature decorrelation and feature information-activity. For the first part, Laplacian sparse coding and low-rank decomposition are used to extract decorrelated features from the scenes. For the second part, Incremental Coding Length is applied to measure the information-activity contained in features, which is then employed to weight the decorrelated features. Finally, visual saliency is estimated through a max pooling strategy. Experimental results on a publicly available benchmark demonstrate the effectiveness of our proposed model with good performance against the state-of-the-art methods.
Shengxiang Qi, Jin-Gang Yu, Ji Zhao 0001, Jie Ma 0003, Jinwen Tian
ICIP2
2014 Maximal Entropy Random Walk for Region-Based Visual Saliency
abstract
Visual saliency is attracting more and more research attention since it is beneficial to many computer vision applications. In this paper, we propose a novel bottom-up saliency model for detecting salient objects in natural images. First, inspired by the recent advance in the realm of statistical thermodynamics, we adopt a novel mathematical model, namely, the maximal entropy random walk (MERW) to measure saliency. We analyze the rationality and superiority of MERW for modeling visual saliency. Then, based on the MERW model, we establish a generic framework for saliency detection. Different from the vast majority of existing saliency models, our method is built on a purely region-based strategy, which is able to yield high-resolution saliency maps with well preserved object shapes and uniformly highlighted salient regions. In the proposed framework, the input image is first over-segmented into superpixels, which are taken as the primary units for subsequent procedures, and regional features are extracted. Then, saliency is measured according to two principles, i.e., uniqueness and visual organization, both implemented in a unified approach, i.e., the MERW model based on graph representation. Intensive experimental results on publicly available datasets demonstrate that our method outperforms the state-of-the-art saliency models.
Jin-Gang Yu, Ji Zhao 0001, Jinwen Tian, Yihua Tan
IEEE Trans. Cybern.1
2012 Urban area detection using multiple Kernel Learning and graph cut
abstract
This paper presents a new method for urban detection from high-spatial-resolution satellite images. Unlike traditional approaches using only texture information for urban detection, we integrate several complementary image features through multiple Kernel Learning framework, and demonstrate that fusing multiple features can help improving urban detection accuracy rate. Furthermore, since that most of supervised urban classification approaches are mainly based on block-based image interpretation, the resulting urban boundary is very coarse. To handle this, we formulate the urban boundary refinement as a binary labeling problem, and propose a graph cut based approach to solve it. Experimental results show that the proposed approach outperforms the existing algorithm in terms of detection accuracy.
Chao Tao 0001, Yihua Tan, Jin-Gang Yu, Jin-Wen Tian
IGARSS3