Youbao Tang

dblp:20/8578 · DBLP profile ↗
← Back
31ranked-venue papers
17as first author
11since 2021 · last 2024
0000-0001-8719-3375ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 15 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2024 Bidirectional Autoregressive Diffusion Model for Dance Generation
abstract
Dance serves as a powerful medium for expressing human emotions, but the lifelike generation of dance is still a considerable challenge. Recently, diffusion models have showcased remarkable generative abilities across various domains. They hold promise for human motion generation due to their adaptable many-to-many nature. Nonetheless, current diffusion-based motion generation models often create entire motion sequences directly and unidirectionally, lacking focus on the motion with local and bidirectional enhancement. When choreographing high-quality dance movements, people need to take into account not only the musical context but also the nearby music-aligned dance motions. To authentically capture human behavior, we propose a Bidirectional Autoregressive Diffusion Model (BADM) for music-to-dance generation, where a bidirectional encoder is built to enforce that the generated dance is harmonious in both the forward and backward directions. To make the generated dance motion smoother, a local information decoder is built for local motion enhancement. The proposed framework is able to generate new motions based on the input conditions and nearby motions, which foresees individual motion slices iteratively and con-solidates all predictions. To further refine the synchronicity between the generated dance and the beat, the beat information is incorporated as an input to generate better music-aligned dance movements. Experimental results demonstrate that the proposed model achieves state-of-the-art performance compared to existing unidirectional approaches on the prominent benchmark for music-to-dance generation. The code and models are available: https://github.com/czzhang179/BADM.
Canyu Zhang 0002, Youbao Tang, Ruei-Sung Lin, Jing Xiao 0006, Song Wang 0002
CVPR2
2023 Prior-Enhanced Temporal Action Localization Using Subject-Aware Spatial Attention
abstract
Temporal action localization (TAL) aims to detect the boundary and identify the class of each action instance in a long untrimmed video. Current approaches treat video frames homogeneously, and tend to give background and key objects excessive attention. This limits their sensitivity to localize action boundaries. To this end, we propose a prior-enhanced temporal action localization method (PETAL), which only takes in RGB input and incorporates action subjects as priors. This proposal leverages action subjects’ information with a plug-and-play subject-aware spatial attention module (SA-SAM) to generate an aggregated and subject-prioritized representation. Experimental results on THUMOS-14 and ActivityNet-1.3 datasets demonstrate that the proposed PETAL achieves competitive performance using only RGB features, e.g., boosting mAP by 2.41% or 0.25% over the state-of-the-art approach that uses RGB features or with additional optical flow features on the THUMOS-14 dataset.
Youbao Tang, Ruei-Sung Lin, Haoqian Wang
ICASSP2
2022 Accurate and Robust Lesion RECIST Diameter Prediction and Segmentation with Transformers
Youbao Tang, Yirui Wang 0002, Shenghua He, Jing Xiao 0006, Ruei-Sung Lin
MICCAI (4)1
2022 Global-Local attention network with multi-task uncertainty loss for abnormal lymph node detection in MR images
Shuai Wang 0003, Yingying Zhu 0003, Sungwon Lee 0003, Daniel C. Elton, Thomas C. Shen, Youbao Tang, Yifan Peng 0002, Zhiyong Lu, Ronald M. Summers
Medical Image Anal.6
2022 SAM: Self-Supervised Learning of Pixel-Wise Anatomical Embeddings in Radiological Images
abstract
Radiological images such as computed tomography (CT) and X-rays render anatomy with intrinsic structures. Being able to reliably locate the same anatomical structure across varying images is a fundamental task in medical image analysis. In principle it is possible to use landmark detection or semantic segmentation for this task, but to work well these require large numbers of labeled data for each anatomical structure and sub-structure of interest. A more universal approach would learn the intrinsic structure from unlabeled images. We introduce such an approach, called Self-supervised Anatomical eMbedding (SAM). SAM generates semantic embeddings for each image pixel that describes its anatomical location or body part. To produce such embeddings, we propose a pixel-level contrastive learning framework. A coarse-to-fine strategy ensures both global and local anatomical information are encoded. Negative sample selection strategies are designed to enhance the embedding's discriminability. Using SAM, one can label any point of interest on a template image and then locate the same body part in other images by simple nearest neighbor searching. We demonstrate the effectiveness of SAM in multiple tasks with 2D and 3D image modalities. On a chest CT dataset with 19 landmarks, SAM outperforms widely-used registration algorithms while only taking 0.23 seconds for inference. On two X-ray datasets, SAM, with only one labeled template image, surpasses supervised methods trained on 50 labeled images. We also apply SAM on whole-body follow-up lesion matching in CT and obtain an accuracy of 91%. SAM can also be applied for improving image registration and initializing CNN weights.
Ke Yan 0006, Jinzheng Cai, Dakai Jin, Shun Miao, Dazhou Guo, Adam P. Harrison, Youbao Tang, Jing Xiao 0006, Jingjing Lu, Le Lu 0001
IEEE Trans. Medical Imaging7
2021 Deep Lesion Tracker: Monitoring Lesions in 4D Longitudinal Imaging Studies
abstract
Monitoring treatment response in longitudinal studies plays an important role in clinical practice. Accurately identifying lesions across serial imaging follow-up is the core to the monitoring procedure. Typically this incorporates both image and anatomical considerations. However, matching lesions manually is labor-intensive and time-consuming. In this work, we present deep lesion tracker (DLT), a deep learning approach that uses both appearance- and anatomical-based signals. To incorporate anatomical constraints, we propose an anatomical signal encoder, which prevents lesions being matched with visually similar but spurious regions. In addition, we present a new formulation for Siamese networks that avoids the heavy computational loads of 3D cross-correlation. To present our network with greater varieties of images, we also propose a self-supervised learning (SSL) strategy to train trackers with unpaired images, overcoming barriers to data collection. To train and evaluate our tracker, we introduce and release the first lesion tracking benchmark, consisting of 3891 lesion pairs from the public DeepLesion database. The proposed method, DLT, locates lesion centers with a mean error distance of 7mm. This is 5% better than a leading registration algorithm while running 14 times faster on whole CT volumes. We demonstrate even greater improvements over detector or similarity-learning alternatives. DLT also generalizes well on an external clinical test set of 100 longitudinal studies, achieving 88% accuracy. Finally, we plug DLT into an automatic tumor monitoring workflow where it leads to an accuracy of 85% in assessing lesion treatment responses, which is only 0.46% lower than the accuracy of manual inputs.
Jinzheng Cai, Youbao Tang, Ke Yan 0006, Adam P. Harrison, Jing Xiao 0006, Gigin Lin, Le Lu 0001
CVPR2
2021 Sequential Learning on Liver Tumor Boundary Semantics and Prognostic Biomarker Mining
Jieneng Chen, Ke Yan 0006, Youbao Tang, Shuwen Sun, Qiuping Liu, Lingyun Huang, Jing Xiao 0006, Alan L. Yuille, Ya Zhang 0002, Le Lu 0001
MICCAI (7)4
2021 Weakly-Supervised Universal Lesion Segmentation with Regional Level Set Loss
Youbao Tang, Jinzheng Cai, Ke Yan 0006, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Jingjing Lu, Gigin Lin, Le Lu 0001
MICCAI (2)1
2021 Lesion Segmentation and RECIST Diameter Prediction via Click-Driven Attention and Dual-Path Connection
Youbao Tang, Ke Yan 0006, Jinzheng Cai, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Jingjing Lu, Gigin Lin, Le Lu 0001
MICCAI (2)1
2021 A disentangled generative model for disease decomposition in chest X-rays via normal image synthesis
Youbao Tang, Yuxing Tang, Yingying Zhu 0003, Jing Xiao 0006, Ronald M. Summers
Medical Image Anal.1
2021 Learning From Multiple Datasets With Heterogeneous and Partial Labels for Universal Lesion Detection in CT
abstract
Large-scale datasets with high-quality labels are desired for training accurate deep learning models. However, due to the annotation cost, datasets in medical imaging are often either partially-labeled or small. For example, DeepLesion is such a large-scale CT image dataset with lesions of various types, but it also has many unlabeled lesions (missing annotations). When training a lesion detector on a partially-labeled dataset, the missing annotations will generate incorrect negative signals and degrade the performance. Besides DeepLesion, there are several small single-type datasets, such as LUNA for lung nodules and LiTS for liver tumors. These datasets have heterogeneous label scopes, i.e., different lesion types are labeled in different datasets with other types ignored. In this work, we aim to develop a universal lesion detection algorithm to detect a variety of lesions. The problem of heterogeneous and partial labels is tackled. First, we build a simple yet effective lesion detection framework named Lesion ENSemble (LENS). LENS can efficiently learn from multiple heterogeneous lesion datasets in a multi-task fashion and leverage their synergy by proposal fusion. Next, we propose strategies to mine missing annotations from partially-labeled datasets by exploiting clinical prior knowledge and cross-dataset knowledge transfer. Finally, we train our framework on four public lesion datasets and evaluate it on 800 manually-labeled sub-volumes in DeepLesion. Our method brings a relative improvement of 49% compared to the current state-of-the-art approach in the metric of average sensitivity. We have publicly released our manual 3D annotations of DeepLesion online.11https://github.com/viggin/DeepLesion_manual_test_set
Ke Yan 0006, Jinzheng Cai, Youjing Zheng, Adam P. Harrison, Dakai Jin, Youbao Tang, Yuxing Tang, Lingyun Huang, Jing Xiao 0006, Le Lu 0001
IEEE Trans. Medical Imaging6
2020 E2Net: An Edge Enhanced Network for Accurate Liver and Tumor Segmentation on CT Scans
Youbao Tang, Yuxing Tang, Yingying Zhu 0003, Jing Xiao 0006, Ronald M. Summers
MICCAI (4)1
2020 One Click Lesion RECIST Measurement and Segmentation on CT Scans
Youbao Tang, Ke Yan 0006, Jing Xiao 0006, Ronald M. Summers
MICCAI (4)1
2020 Cross-domain Medical Image Translation by Shared Latent Gaussian Mixture Model
Yingying Zhu 0003, Youbao Tang, Yuxing Tang, Daniel C. Elton, Sungwon Lee 0003, Perry J. Pickhardt, Ronald M. Summers
MICCAI (2)2
2020 Discrete Probability Distribution Prediction of Image Emotions with Shared Sparse Learning
abstract
Computationally modelling the affective content of images has been extensively studied recently because of its wide applications in entertainment, advertisement, and education. Significant progress has been made on designing discriminative features to bridge the affective gap. Assuming that viewers can reach a consensus on the emotion of images, most existing works focused on assigning the dominant emotion category or the average dimension values to an image. However, the image emotions perceived by viewers are subjective by nature with the influence of personal and situational factors. In this paper, we propose a novel machine learning approach that characterizes the categorical image emotions as a discrete probability distribution (DPD). To associate emotion with the visual features extracted from images, we present shared sparse learning to learn the combination coefficients, with which the DPD of an unseen image is predicted by linearly combining the DPDs of the training images. Furthermore, we extend our method to the setup where multi-features are available and learn the optimal weights for each feature to reflect the importance of different features. Extensive experiments are carried out on Abstract, Emotion6 and IESN datasets and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art approaches.
Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Xin Zhao 0020, Youbao Tang, Jungong Han, Hongxun Yao, Qingming Huang
IEEE Trans. Affect. Comput.5
2019 TUNA-Net: Task-Oriented UNsupervised Adversarial Network for Disease Recognition in Cross-domain Chest X-rays
Yuxing Tang, Youbao Tang, Veit Sandfort, Jing Xiao 0006, Ronald M. Summers
MICCAI (6)2
2019 MULAN: Multitask Universal Lesion Analysis Network for Joint Lesion Detection, Tagging, and Segmentation
Ke Yan 0006, Youbao Tang, Yifan Peng 0002, Veit Sandfort, Mohammadhadi Bagheri, Zhiyong Lu, Ronald M. Summers
MICCAI (6)2
2019 Salient Object Detection Using Cascaded Convolutional Neural Networks and Adversarial Learning
abstract
Salient object detection has received much attention and achieved great success in last several years. It is still challenging to get clear boundaries and consistent saliencies, which can be considered as the structural information of salient objects. A popular solution is to conduct some post-processes (e.g., conditional random field (CRF)) to refine these structural information. In this paper, a novel cascaded convolutional neural networks (CNNs) based method is proposed to implicitly learn these structural information via adversarial learning for salient object detection (we termed the proposed method as CCAL). A cascaded CNNs model is first designed as a generator G, which consists of an encoder-decoder network for global saliency estimation and a deep residual network for local saliency refinement. It is hard to explicitly learn such structural information due to the limitation of frequently-used pixel-wise loss functions. Instead, a discriminator D is then designed to distinguish the real salient maps (i.e., ground truths) from the fake ones produced by G, based on which an adversarial loss is introduced to optimize G. G and D are trained in a fully end-to-end fashion by following the strategy of conditional generative adversarial networks to make G well learn the structural information. At last, G is able to produce high quality salient maps without requiring any post-process to fool D. Experimental results on eight benchmark datasets demonstrate the effectiveness and efficiency (about 17 fps on graphics processing unit (GPU)) of the proposed method for salient object detection.
Youbao Tang, Xiangqian Wu 0002
IEEE Trans. Multim.1
2018 Accurate Weakly-Supervised Deep Lesion Segmentation Using Large-Scale Clinical Annotations: Slice-Propagated 3D Mask Generation from 2D RECIST
Jinzheng Cai, Youbao Tang, Le Lu 0001, Adam P. Harrison, Ke Yan 0006, Jing Xiao 0006, Lin Yang 0002, Ronald M. Summers
MICCAI (4)2
2018 CT-Realistic Lung Nodule Simulation from 3D Conditional Generative Adversarial Networks for Robust Lung Segmentation
Dakai Jin, Ziyue Xu 0001, Youbao Tang, Adam P. Harrison, Daniel J. Mollura
MICCAI (2)3
2018 Semi-automatic RECIST Labeling on CT Scans with Cascaded Convolutional Neural Networks
Youbao Tang, Adam P. Harrison, Mohammadhadi Bagheri, Jing Xiao 0006, Ronald M. Summers
MICCAI (4)1
2018 Scene Text Detection Using Superpixel-Based Stroke Feature Transform and Deep Learning Based Region Classification
abstract
Scene text detection is a crucial step in end-to-end scene text recognition, a greatly challenging problem in computer vision. This paper proposes a novel scene text detection method that involves superpixel-based stroke feature transform (SSFT) and deep learning based region classification (DLRC). The SSFT is developed for candidate character region (CCR) extraction, which consists in partitioning an input image into several regions via superpixel-based clustering, removing most regions based on predefined criteria satisfied by the characters, and refining the remaining regions to obtain CCRs by computing a stroke width map. The character regions are identified from the CCRs using DLRC, in which several hand-crafted low-level features, i.e., color, texture, and geometric features, and some deep convolution neural network (CNN) based high-level features are first extracted from the regions, and then these features are fused by using two fully connected networks (FCNs) for region classification. In the DLRC step, the deep feature extraction CNN and the feature fusion FCNs are jointly trained. Next, the extracted character regions are merged to form candidate text regions, from which the final scene texts are detected. The proposed method is evaluated on three publicly available datasets: ICDAR2011, ICDAR2013, and street view text. It achieves F -measures of 0.876, 0.885, and 0.631, respectively, which demonstrate the effectiveness of the proposed scene text detection method.
Youbao Tang, Xiangqian Wu 0002
IEEE Trans. Multim.1
2017 Salient Object Detection with Chained Multi-Scale Fully Convolutional Network
abstract
In this paper, we proposed a novel method for effective salient object detection by designing a chained multi-scale fully convolutional network (CMSFCN). CMSFCN contained multiple single-scale fully convolutional networks (SSFCNs), which were integrated successively by using chained connections and generated saliency prediction results from coarse to fine. The chained connections not only combined the saliency prediction result from previous SSFCN with the input image of current SSFCN, but also combined the intermediate features from previous SSFCN and current SSFCN. With these chained connections, the sequential SSFCNs in CMSFCN automatically learned complemental and discriminative features to improve the saliency predictions progressively. Therefore, after jointly training CMSFCN with an end-to-end manner, precise saliency prediction results were produced under a coarse-to-fine behaviour. Compared with seven state-of-the-art CNN based salient object detection approaches over five benchmark datasets, experimental results demonstrated the efficiency and effectiveness of CMSFCN.
Youbao Tang, Xiangqian Wu 0002
ACM Multimedia1
2017 Scene Text Detection and Segmentation Based on Cascaded Convolution Neural Networks
abstract
Scene text detection and segmentation are two important and challenging research problems in the field of computer vision. This paper proposes a novel method for scene text detection and segmentation based on cascaded convolution neural networks (CNNs). In this method, a CNN based text-aware candidate text region (CTR) extraction model (named detection network, DNet) is designed and trained using both the edges and the whole regions of text, with which coarse CTRs are detected. A CNN based CTR refinement model (named segmentation network, SNet) is then constructed to precisely segment the coarse CTRs into text to get the refined CTRs. With DNet and SNet, much fewer CTRs are extracted than with traditional approaches while more true text regions are kept. The refined CTRs are finally classified using a CNN based CTR classification model (named classification network, CNet) to get the final text regions. All of these CNN based models are modified from VGGNet-16. Extensive experiments on three benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance and greatly outperforms other scene text detection and segmentation approaches.
Youbao Tang, Xiangqian Wu 0002
IEEE Trans. Image Process.1
2016 Saliency Detection via Combining Region-Level and Pixel-Level Predictions with CNNs
Youbao Tang, Xiangqian Wu 0002
ECCV (8)1
2016 Scene Text Detection via Edge Cue and Multi-features
abstract
Inspired by the fact that edge is an important cue to distinguish texts from background, we propose a novel scene text detection method via edge cue and multiple features, which has two main parts, i.e. candidate character region (CCR) extraction and region classification. For CCR extraction, the edges are first extracted from the input image, which are then broken and merged based on color features to form the final edge image. For each edge connected component, a number of image patches are extracted by translating and scaling its boundary rectangle to generate the CCRs. For region classification, the character regions are extracted from the CCRs by using a region classification technique, which extracts both the hand-designed low-level features and deep convolution neural network based high-level features of the regions for classification. And then the character regions are merged to form the candidate text regions, based on which the final text region are detected by using the region classification technique. The proposed method is evaluated on two latest ICDAR benchmark datasets and the experimental results demonstrate that the proposed method outperforms the state-of-the-art approaches of scene text detection.
Youbao Tang, Xiangqian Wu 0002
ICFHR1
2016 Text-Independent Writer Identification via CNN Features and Joint Bayesian
abstract
This paper proposes a novel method for offline text-independent writer identification by using convolutional neural network (CNN) and joint Bayesian, which consists of two stages, i.e. feature extraction and writer identification. In the stage of feature extraction, since a large number of data is essential to train an effective CNN model with high generalizability and the amount of handwriting is limited in writer identification, a data augmentation technique is first developed to generate thousands of handwriting images for each writer. Then a deep CNN network is designed to extract discriminative features to represent the properties of different writing styles, which is trained by using the generated handwriting images. In the stage of writer identification, the training dataset is used to train the CNN model for feature extraction and the joint Bayesian technique is employed to accomplish the task of writer identification based on the extracted CNN features. The proposed method is tested on two standard benchmark datasets, i.e. ICDAR2013 and CVL dataset. Experimental results demonstrate that the proposed method gets the best performance compared to the state-of-the-art approaches.
Youbao Tang, Xiangqian Wu 0002
ICFHR1
2016 Deeply-Supervised Recurrent Convolutional Neural Network for Saliency Detection
abstract
This paper proposes a novel saliency detection method by developing a deeply-supervised recurrent convolutional neural network (DSRCNN), which performs a full image-to-image saliency prediction. For saliency detection, the local, global, and contextual information of salient objects is important to obtain a high quality salient map. To achieve this goal, the DSRCNN is designed based on VGGNet-16. Firstly, the recurrent connections are incorporated into each convolutional layer, which can make the model more powerful for learning the contextual information. Secondly, side-output layers are added to conduct the deeply-supervised operation, which can make the model learn more discriminative and robust features by effecting the intermediate layers. Finally, all of the side-outputs are fused to integrate the local and global information to get the final saliency detection results. Therefore, the DSRCNN combines the advantages of recurrent convolutional neural networks and deeply-supervised nets. The DSRCNN model is tested on five benchmark datasets, and experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art saliency detection approaches on all test datasets.
Youbao Tang, Xiangqian Wu 0002, Wei Bu
ACM Multimedia1
2015 Saliency Detection Based on Graph-Structural Agglomerative Clustering
abstract
This paper proposes a novel saliency detection method based on graph-structural agglomerative clustering (GSAC). In this method, a number of intermediate images with consecutive number of regions are firstly created by using GSAC to the input image with the maximum incremental path integral criterion. Then an initial salient map is computed based on the boundary connectivity of the regions in the intermediate images, with enforcing the early formed objects in the clustering process. Finally, the initial salient map is refined to get the final salient map by using the reconstruction errors of sparse coding and the object-bias prior. The experimental results demonstrate that the proposed method greatly outperforms the state-of-the-art approaches on two standard benchmark datasets.
Youbao Tang, Xiangqian Wu 0002, Wei Bu
ACM Multimedia1
2014 Text Line Segmentation Based on Matched Filtering and Top-Down Grouping for Handwritten Documents
abstract
This paper presents a novel text line segmentation method based on matched filtering and top-down grouping for handwritten documents. The proposed method consists of three distinct steps. Firstly, the foreground pixel density (FPD) of handwritten document image (HDI) is estimated, then FPD is used to decide the size of the generated filter which is the convolution of a band-shape filter and an isotropic LoG filter. Secondly, the centers of the text lines (CTLs) are extracted by performing filtering, binarizing, thinning and top-down grouping operation on HDI. Finally, the overlapping connected-components (OCCs) which travel through multiple text lines are separated, and then all OCCs are assigned to a label of CTLs by the nearest neighbor principle. The proposed method is tested on two public databases, and the experimental results show that the proposed method outperforms the state-of-the-art text line segmentation approaches in both of these databases.
Youbao Tang, Xiangqian Wu 0002, Wei Bu
Document Analysis Systems1
2014 Offline Text-Independent Writer Identification Based on Scale Invariant Feature Transform
abstract
This paper proposes a novel offline text-independent writer identification method based on scale invariant feature transform (SIFT), composed of training, enrollment, and identification stages. In all stages, an isotropic LoG filter is first used to segment the handwriting image into word regions (WRs). Then, the SIFT descriptors (SDs) of WRs and the corresponding scales and orientations (SOs) are extracted. In the training stage, an SD codebook is constructed by clustering the SDs of training samples. In the enrollment stage, the SDs of the input handwriting are adopted to form an SD signature (SDS) by looking up the SD codebook and the SOs are utilized to generate a scale and orientation histogram (SOH). In the identification stage, the SDS and SOH of the input handwriting are extracted and matched with the enrolled ones for identification. Experimental results on six public data sets (including three English data sets, one Chinese data set, and two hybrid-language data sets) demonstrate that the proposed method outperforms the state-of-the-art algorithms.
Xiangqian Wu 0002, Youbao Tang, Wei Bu
IEEE Trans. Inf. Forensics Secur.2