Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kuang-Jui Hsu

dblp:127/5013 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
0since 2021 · last 2020
0000-0003-4055-3585ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-authorArtificial intelligence and machine learning · 10 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Segmentation and scene understanding · 68% Learning paradigms · 6% Vision and language · 6%
Computer graphics and multimedia
4 papers
Image and video processing · 56% Multimedia analysis and retrieval · 44%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding › saliency detection › salient object detection
co-saliency detection
1.132020
Deep Co-Saliency Detection via Stacked Autoencoder-Enabled Fusion and Self-Trained CNNs · IEEE Trans. Multim. 2020
DeepCO3: Deep Instance Co-Segmentation by Co-Peak Search and Co-Saliency Detection · CVPR 2019
Unsupervised CNN-Based Co-saliency Detection with Graphical Optimization · ECCV (5) 2018
Computer vision › Segmentation and scene understanding › image segmentation
co-segmentation
0.822020
Deep Co-Saliency Detection via Stacked Autoencoder-Enabled Fusion and Self-Trained CNNs · IEEE Trans. Multim. 2020
Co-attention CNNs for Unsupervised Object Co-segmentation · IJCAI 2018
Machine learning › Learning paradigms
multiple instance learning
0.422019
Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior · NeurIPS 2019
Multiple Structured-Instance Learning for Semantic Segmentation with Uncertain Training Data · CVPR 2014
Computer vision › Vision and language › video grounding
video re-localization
0.412020
Weakly-Supervised Video Re-Localization with Multiscale Attention Model · AAAI 2020
Computer vision › Segmentation and scene understanding › instance segmentation
instance co-segmentation
0.412019
DeepCO3: Deep Instance Co-Segmentation by Co-Peak Search and Co-Saliency Detection · CVPR 2019
Computer vision › Segmentation and scene understanding
instance segmentation
0.412019
Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior · NeurIPS 2019
Machine learning › Optimization for machine learning › optimization
joint optimization
0.412019
Image Co-Saliency Detection and Co-Segmentation via Progressive Joint Optimization · IEEE Trans. Image Process. 2019
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection
0.412019
Weakly Supervised Salient Object Detection by Learning A Classifier-Driven Map Generator · IEEE Trans. Image Process. 2019
Computer vision › Segmentation and scene understanding
semantic segmentation
0.422014
Augmented Multiple Instance Regression for Inferring Object Contours in Bounding Boxes · IEEE Trans. Image Process. 2014
Multiple Structured-Instance Learning for Semantic Segmentation with Uncertain Training Data · CVPR 2014
Computer vision › Segmentation and scene understanding › instance segmentation
weakly supervised instance segmentation
0.412019
Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior · NeurIPS 2019
Computer vision › Segmentation and scene understanding › saliency detection › salient object detection
weakly supervised salient object detection
0.412019
Weakly Supervised Salient Object Detection by Learning A Classifier-Driven Map Generator · IEEE Trans. Image Process. 2019
Computer vision › 3D vision › multi-view geometry
homography estimation
0.212015
Matching Images With Multiple Descriptors: An Unsupervised Approach for Locally Adaptive Descriptor Selection · IEEE Trans. Image Process. 2015
Image and video processing › image matching
feature matching
0.212015
Matching Images With Multiple Descriptors: An Unsupervised Approach for Locally Adaptive Descriptor Selection · IEEE Trans. Image Process. 2015
Image and video processing
image registration
0.212015
Robust image alignment with multiple feature descriptors and matching-guided neighborhoods · CVPR 2015
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field
0.212014
Multiple Structured-Instance Learning for Semantic Segmentation with Uncertain Training Data · CVPR 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.212014
Multiple Structured-Instance Learning for Semantic Segmentation with Uncertain Training Data · CVPR 2014
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation
0.212014
Multiple Structured-Instance Learning for Semantic Segmentation with Uncertain Training Data · CVPR 2014
Computer vision › Image recognition and object detection
image classification
0.112019
Weakly Supervised Salient Object Detection by Learning A Classifier-Driven Map Generator · IEEE Trans. Image Process. 2019
Machine learning and data management › weak supervision
multiple instance learning
0.112014
Augmented Multiple Instance Regression for Inferring Object Contours in Bounding Boxes · IEEE Trans. Image Process. 2014

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 1.1co-attention loss · 0.9c3d features · 0.9attention mechanism · 0.9unsupervised learning · 0.8conditional random field · 0.6multiple instance learning · 0.6stacked autoencoder · 0.4self-training · 0.4progressive joint optimization · 0.4energy minimization over graph · 0.4CNN · 0.4one-class SVM · 0.2geodesic distance · 0.2descriptor fusion · 0.2affine transformation learning · 0.2augmented multiple instance regression · 0.2
YearPublicationVenuePosition
2020 Weakly-Supervised Video Re-Localization with Multiscale Attention Model
abstract
Video re-localization aims to localize a sub-sequence, called target segment, in an untrimmed reference video that is similar to a given query video. In this work, we propose an attention-based model to accomplish this task in a weakly supervised setting. Namely, we derive our CNN-based model without using the annotated locations of the target segments in reference videos. Our model contains three modules. First, it employs a pre-trained C3D network for feature extraction. Second, we design an attention mechanism to extract multiscale temporal features, which are then used to estimate the similarity between the query video and a reference video. Third, a localization layer detects where the target segment is in the reference video by determining whether each frame in the reference video is consistent with the query video. The resultant CNN model is derived based on the proposed co-attention loss which discriminatively separates the target segment from the reference video. This loss maximizes the similarity between the query video and the target segment while minimizing the similarity between the target segment and the rest of the reference video. Our model can be modified to fully supervised re-localization. Our method is evaluated on a public dataset and achieves the state-of-the-art performance under both weakly supervised and fully supervised settings.
Yung-Han Huang, Kuang-Jui Hsu, Shyh-Kang Jeng, Yen-Yu Lin
AAAI2
2020 Deep Co-Saliency Detection via Stacked Autoencoder-Enabled Fusion and Self-Trained CNNs
abstract
Image co-saliency detection via fusion-based or learning-based methods faces cross-cutting issues. Fusion-based methods often combine saliency proposals using a majority voting rule. Their performance hence highly depends on the quality and coherence of individual proposals. Learning-based methods typically require ground-truth annotations for training, which are not available for co-saliency detection. In this work, we present a two-stage approach to address these issues jointly. At the first stage, an unsupervised deep learning model with stacked autoencoder (SAE) is proposed to evaluate the quality of saliency proposals. It employs latent representations for image foregrounds, and auto-encodes foreground consistency and foreground-background distinctiveness in a discriminative way. The resultant model, SAE-enabled fusion (SAEF), can combine multiple saliency proposals to yield a more reliable saliency map. At the second stage, motivated by the fact that fusion often leads to over-smoothed saliency maps, we develop self-trained convolutional neural networks (STCNN) to alleviate this negative effect.STCNNtakes the saliency maps produced bySAEFas inputs. It propagates information from regions of high confidence to those of low confidence. During propagation, feature representations are distilled, resulting in sharper and better co-saliency maps. Our approach is comprehensively evaluated on three benchmarks, including MSRC, iCoseg, and Cosal2015, and performs favorably against the state-of-the-arts. In addition, we demonstrate that our method can be applied to object co-segmentation and object co-localization, achieving the state-of-the-art performance in both applications.
Chung-Chi Tsai, Kuang-Jui Hsu, Yen-Yu Lin, Xiaoning Qian, Yung-Yu Chuang
IEEE Trans. Multim.2
2019 DeepCO3: Deep Instance Co-Segmentation by Co-Peak Search and Co-Saliency Detection
abstract
In this paper, we address a new task called instance co-segmentation. Given a set of images jointly covering object instances of a specific category, instance co-segmentation aims to identify all of these instances and segment each of them, i.e. generating one mask for each instance. This task is important since instance-level segmentation is preferable for humans and many vision applications. It is also challenging because no pixel-wise annotated training data are available and the number of instances in each image is unknown. We solve this task by dividing it into two sub-tasks, co-peak search and instance mask segmentation. In the former sub-task, we develop a CNN-based network to detect the co-peaks as well as co-saliency maps for a pair of images. A co-peak has two endpoints, one in each image, that are local maxima in the response maps and similar to each other. Thereby, the two endpoints are potentially covered by a pair of instances of the same category. In the latter subtask, we design a ranking function that takes the detected co-peaks and co-saliency maps as inputs and can select the object proposals to produce the final results. Our method for instance co-segmentation and its variant for object colocalization are evaluated on four datasets, and achieve favorable performance against the state-of-the-art methods. The source codes and the collected datasets are available at https://github.com/KuangJuiHsu/DeepCO3/
Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang
CVPR1
2019 Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior
abstract
This paper presents a weakly supervised instance segmentation method that consumes training data with tight bounding box annotations. The major difficulty lies in the uncertain figure-ground separation within each bounding box since there is no supervisory signal about it. We address the difficulty by formulating the problem as a multiple instance learning (MIL) task, and generate positive and negative bags based on the sweeping lines of each bounding box. The proposed deep model integrates MIL into a fully supervised instance segmentation network, and can be derived by the objective consisting of two terms, i.e., the unary term and the pairwise term. The former estimates the foreground and background areas of each bounding box while the latter maintains the unity of the estimated object masks. The experimental results show that our method performs favorably against existing weakly supervised methods and even surpasses some fully supervised methods for instance segmentation on the PASCAL VOC dataset.
Cheng-Chun Hsu, Kuang-Jui Hsu, Chung-Chi Tsai, Yen-Yu Lin, Yung-Yu Chuang
NeurIPS2
2019 Weakly Supervised Salient Object Detection by Learning A Classifier-Driven Map Generator
abstract
Top-down saliency detection aims to highlight the regions of a specific object category, and typically relies on pixel-wise annotated training data. In this paper, we address the high cost of collecting such training data by a weakly supervised approach to object saliency detection, where only image-level labels, indicating the presence or absence of a target object in an image, are available. The proposed framework is composed of two collaborative CNN modules, an image-level classifier and a pixel-level map generator. While the former distinguishes images with objects of interest from the rest, the latter is learned to generate saliency maps by which the images masked by the maps can be better predicted by the former. In addition to the top-down guidance from class labels, the map generator is derived by also exploring other cues, including the background prior, superpixel- and object proposal-based evidence. The background prior is introduced to reduce false positives. Evidence from superpixels helps preserve sharp object boundaries. The clue from object proposals improves the integrity of highlighted objects. These different types of cues greatly regularize the training process and reduces the risk of overfitting, which happens frequently when learning CNN models with few training data. Experiments show that our method achieves superior results, even outperforming fully supervised methods.
Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang
IEEE Trans. Image Process.1
2019 Image Co-Saliency Detection and Co-Segmentation via Progressive Joint Optimization
abstract
We present a novel computational model for simultaneous image co-saliency detection and co-segmentation that concurrently explores the concepts of saliency and objectness in multiple images. It has been shown that the co-saliency detection via aggregating multiple saliency proposals by diverse visual cues can better highlight the salient objects; however, the optimal proposals are typically region-dependent and the fusion process often leads to blurred results. Co-segmentation can help preserve object boundaries, but it may suffer from complex scenes. To address these issues, we develop a unified method that addresses co-saliency detection and co-segmentation jointly via solving an energy minimization problem over a graph. Our method iteratively carries out the region-wise adaptive saliency map fusion and object segmentation to transfer useful information between the two complementary tasks. Through the optimization iterations, sharp saliency maps are gradually obtained to recover entire salient objects by referring to object segmentation, while these segmentations are progressively improved owing to the better saliency prior. We evaluate our method on four public benchmark data sets while comparing it to the state-of-the-art methods. Extensive experiments demonstrate that our method can provide consistently higher-quality results on both co-saliency detection and co-segmentation.
Chung-Chi Tsai, Weizhi Li, Kuang-Jui Hsu, Xiaoning Qian, Yen-Yu Lin
IEEE Trans. Image Process.3
2018 Unsupervised CNN-Based Co-saliency Detection with Graphical Optimization
Kuang-Jui Hsu, Chung-Chi Tsai, Yen-Yu Lin, Xiaoning Qian, Yung-Yu Chuang
ECCV (5)1
2018 Co-attention CNNs for Unsupervised Object Co-segmentation
abstract
Object co-segmentation aims to segment the common objects in images. This paper presents a CNN-based method that is unsupervised and end-to-end trainable to better solve this task. Our method is unsupervised in the sense that it does not require any training data in the form of object masks but merely a set of images jointly covering objects of a specific class. Our method comprises two collaborative CNN modules, a feature extractor and a co-attention map generator. The former module extracts the features of the estimated objects and backgrounds, and is derived based on the proposed co-attention loss which minimizes inter-image object discrepancy while maximizing intra-image figure-ground separation. The latter module is learned to generated co-attention maps by which the estimated figure-ground segmentation can better fit the former module. Besides, the co-attention loss, the mask loss is developed to retain the whole objects and remove noises. Experiments show that our method achieves superior results, even outperforming the state-of-the-art, supervised methods.
Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang
IJCAI1
2017 Weakly Supervised Saliency Detection with A Category-Driven Map Generator
Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang
BMVC1
2017 Real-time embedded implementation of robust speed-limit sign recognition using a novel centroid-to-contour description method
abstract
Traffic sign recognition is a very important function in automatic driving assistance systems (ADAS). This study addresses the design and implementation of a vision‐based ADAS based on an image‐based speed‐limit sign (SLS) recognition algorithm, which can automatically detect and recognise SLS on the road in real‐time. To improve the recognition rate of SLS having different orientations and scales in the image, this study also presents a new sign content description algorithm, which describes the detected road sign using centroid‐to‐contour (CtC) distances of the extracted sign content. The proposed CtC descriptor is robust to translation, rotation and scale changes of the SLS in the image. This advantage improves the recognition accuracy of a support vector machine classifier trained using a large database of traffic signs. The proposed SLS recognition method had been implemented on two different embedded platforms, each of them equipped with an ARM‐based Quad‐Core CPU running Android 4.4 operating system. Experimental results validate that the proposed method not only provides a high recognition rate, but also achieves real‐time performance up to 30 frames per second for processing 1280 × 720 video streams running on a commercial ARM‐based smartphone.
Chi-Yi Tsai, Hsien-Chen Liao, Kuang-Jui Hsu
IET Comput. Vis.3
2015 Robust image alignment with multiple feature descriptors and matching-guided neighborhoods
abstract
This paper addresses two issues hindering the advances in accurate image alignment. First, he performance of descriptor-based approaches to image alignment relies on the chosen descriptor, but the optimal descriptor typically varies from image to image, or even pixel to pixel. Second, the neighborhood structure for smoothness enforcement is usually predefined before alignment. However, object boundaries are often better discovered during alignment. The proposed approach tackles the two issues by adaptive descriptor selection and dynamic neighborhood construction. Specifically we associate each pixel to be aligned with an affine transformation, and integrate the learning of the pixel-specific transformations into image alignment. The transformations serve as the common domain for descriptor fusion, since the local consensus of each descriptor can be estimated by accessing the corresponding affine transformation t allows us to pick the most plausible descriptor for aligning each pixel. On the other hand more object-aware neighborhoods can be produced by referencing the consistency between the learned affine transformations of neighboring pixels. The promising results on popular image alignment benchmarks manifests the effectiveness of our approach.
Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang
CVPR1
2015 Matching Images With Multiple Descriptors: An Unsupervised Approach for Locally Adaptive Descriptor Selection
abstract
With the aim to improve the performance of feature matching, we present an unsupervised approach for adaptive description selection in the space of homographies. Inspired by the observation that the homographies of correct feature correspondences vary smoothly along the spatial domain, our approach stands on the unsupervised nature of feature matching, and can choose a good descriptor locally for matching each feature point, instead of using one global descriptor. To this end, the homography space serves as the domain for selecting various heterogeneous descriptors. Correspondences obtained by any descriptors are considered as points in the space, and their geometric coherence and spatial continuity are measured via computing the geodesic distances. In this way, mutual verification across different descriptors is allowed, and correct correspondences will be highlighted with a high degree of consistency short geodesic distances here. It follows that one-class SVM can be applied to identifying these correct correspondences, and achieves adaptive descriptor selection. The proposed approach is comprehensively compared with the state-of-the-art approaches, and evaluated on five benchmarks of image matching. The promising results manifest its effectiveness.
Yuan-Ting Hu, Yen-Yu Lin, Hsin-Yi Chen, Kuang-Jui Hsu, Bing-Yu Chen 0004
IEEE Trans. Image Process.4
2014 Multiple Structured-Instance Learning for Semantic Segmentation with Uncertain Training Data
abstract
We present an approach MSIL-CRF that incorporates multiple instance learning (MIL) into conditional random fields (CRFs). It can generalize CRFs to work on training data with uncertain labels by the principle of MIL. In this work, it is applied to saving manual efforts on annotating training data for semantic segmentation. Specifically, we consider the setting in which the training dataset for semantic segmentation is a mixture of a few object segments and an abundant set of objects' bounding boxes. Our goal is to infer the unknown object segments enclosed by the bounding boxes so that they can serve as training data for semantic segmentation. To this end, we generate multiple segment hypotheses for each bounding box with the assumption that at least one hypothesis is close to the ground truth. By treating a bounding box as a bag with its segment hypotheses as structured instances, MSIL-CRF selects the most likely segment hypotheses by leveraging the knowledge derived from both the labeled and uncertain training data. The experimental results on the Pascal VOC segmentation task demonstrate that MSIL-CRF can provide effective alternatives to manually labeled segments for semantic segmentation.
Feng-Ju Chang, Yen-Yu Lin, Kuang-Jui Hsu
CVPR3
2014 Augmented Multiple Instance Regression for Inferring Object Contours in Bounding Boxes
abstract
In this paper, we address the problem of the high annotation cost of acquiring training data for semantic segmentation. Most modern approaches to semantic segmentation are based upon graphical models, such as the conditional random fields, and rely on sufficient training data in form of object contours. To reduce the manual effort on pixel-wise annotating contours, we consider the setting in which the training data set for semantic segmentation is a mixture of a few object contours and an abundant set of bounding boxes of objects. Our idea is to borrow the knowledge derived from the object contours to infer the unknown object contours enclosed by the bounding boxes. The inferred contours can then serve as training data for semantic segmentation. To this end, we generate multiple contour hypotheses for each bounding box with the assumption that at least one hypothesis is close to the ground truth. This paper proposes an approach, called augmented multiple instance regression (AMIR), that formulates the task of hypothesis selection as the problem of multiple instance regression (MIR), and augments information derived from the object contours to guide and regularize the training process of MIR. In this way, a bounding box is treated as a bag with its contour hypotheses as instances, and the positive instances refer to the hypotheses close to the ground truth. The proposed approach has been evaluated on the Pascal VOC segmentation task. The promising results demonstrate that AMIR can precisely infer the object contours in the bounding boxes, and hence provide effective alternatives to manually labeled contours for semantic segmentation.
Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang
IEEE Trans. Image Process.1
2012 Knowledge Leverage from Contours to Bounding Boxes: A Concise Approach to Annotation
Jie-Zhi Cheng, Feng-Ju Chang, Kuang-Jui Hsu, Yen-Yu Lin
ACCV (1)3