EDBT 2026 Demo / reviewers in the wild / expert
Kai Zhao 0012
dblp:72/2621-12
· DBLP profile ↗
27ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0002-2496-0829ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Predictive Resampling: Learning Input-Agnostic Downsampling for Efficient Aligned Vision RecognitionabstractImages are typically sampled on a uniform grid,despite their non-uniform information distribution—some regions are rich in content while others are not. The mismatch leads to inefficient computation allocation in deep learning models. To address this, recent studies have proposed predictive downsampling methodsthat adaptively downsample images based on predicted per-pixel importance, allocating more pixels to informative areas. However,these methods require high-resolution processing to accurately estimate importance, which undermines their efficiency:the prediction itself must process the full-resolution image,consuming most of the computational budget. This high-resolution importance prediction is necessary because each input may differ significantly in structure and content. In this paper, we take a different approach and introduce a learn-to-downsample paradigmtailored for aligned vision recognition tasks, such as face recognition and palmprint recognition, where input alignment ensures consistent spatial structure across images. This alignment ensures structural consistency across images, allowing a shared, input-agnostic downsampling template applicable to all inputs. Furthermore, instead of relying on implicit importance maps, we introduce a flow-based representation that explicitly models the spatial warping from the original image to the downsampled version. The flow representation is not only more efficient but also more controllable: we regularize the flow using its Jacobian determinant to precisely control the sampling density and coverage,enabling interpretable and tunable sampling patterns. Extensive experiments on two aligned recognition tasks, face and palmprint recognition, demonstrate that our method substantially reduces computational cost with minimal accuracy degradation, achieving a significantly better performance-efficiency trade-off than existing predictive downsampling methods. Kai Zhao 0012, Liting Ruan, Xiaoqiang Zhu, Xianchao Zhang 0002, Dan Zeng 0001 |
AAAI | 1 |
| 2026 | Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language ModelsabstractOpen-vocabulary camouflaged object segmentation (OVCOS) seeks to segment and classify camouflaged objects in arbitrary categories, presenting unique challenges due to visual ambiguity and unseen categories. Recent approaches typically adopt a two-stage paradigm: they first segment objects, and then classify the segmented regions using vision language models (VLMs). However, such methods (i) suffer from a domain gap caused by the mismatch between VLMs' full-image training and cropped-region inferencing, and (ii) depend on generic segmentation models optimized for well-delineated objects which are less effective for camouflaged objects. Without explicit guidance, generic segmentation models often overlook subtle boundaries, leading to imprecise segmentation. In this paper, we introduce a novel VLM-guided cascaded framework to address these issues in OVCOS. For segmentation, we leverage the segment anything model (SAM), guided by the VLM. Our framework uses VLM-derived features as explicit prompts to SAM, effectively directing attention to camouflaged regions and significantly improving localization accuracy. For classification, we avoid the domain gap introduced by hard cropping. Instead, we treat the segmentation output as a soft spatial prior using the alpha channel. This retains the full image context while providing precise spatial guidance, leading to more accurate and context-aware classification of camouflaged objects. The same VLM is shared between segmentation and classification to ensure efficiency and semantic consistency. Extensive experiments on both OVCOS and conventional camouflaged object segmentation benchmarks demonstrate the clear superiority of our method, highlighting the effectiveness of leveraging rich VLM semantics for both segmentation and classification of camouflaged objects. Our code and models are open-sourced at https://github.com/intcomp/camouflaged-vlm. Kai Zhao 0012, Wubang Yuan, Zheng Wang 0059, Guanyi Li, Xiaoqiang Zhu, Deng-Ping Fan, Dan Zeng 0001 |
Comput. Vis. Media | 1 |
| 2026 | PCa-Mamba: Spatiotemporal state space models for prostate cancer detection in multi-parametric MRI
Kai Zhao 0012, Alex Ling Yu Hung, Kaifeng Pang, Parsa Hajipour, Holden H. Wu, Steven S. Raman, Kyung Hyun Sung |
Medical Image Anal. | 1 |
| 2026 | MIST: A Benchmark and Baseline for Multi-Frame Infrared Small Target Detection in Complex MotionabstractMotion cues play a vital role in multi-frame infrared small target detection (MISTD). However, most targets in existing datasets exhibit regular and slow motion, which cannot reflect the complex and diverse motion patterns in real-world scenarios. This biased data distribution makes recent data-driven methods highly rely on simplified motion assumptions that tend to fail in irregular or fast motion, resulting in noisy feature representations cluttered with target-irrelevant factors. Hence, we stress that methods for MISTD should also work when targets are in complex motion. To enable this research, we propose a large-scale dataset called MIST for airborne infrared detection scenarios. The dataset is built on a synthetic data engine that models variations in pose, size, and intensity of moving targets while seamlessly blending them into real backgrounds for physical, geometric, and visual realism. Targets in MIST exhibit low signal-to-clutter ratios and complex motion, making it a promising yet challenging benchmark for developing algorithms focused on motion analysis. To tackle the challenges of MIST, we develop MISTNet, a robust baseline based on the Information Bottleneck theory. To handle irregular and fast motion, we propose a shifted neighborhood compensation block to efficiently model multi-scale correspondences for implicit motion compensation. To distill compact representations free from irrelevant cues, we design a progressive distillation decoder to hierarchically filter out redundancy while preserving target-relevant information. We benchmark 31 state-of-the-art methods and find that their performance on MIST drops significantly compared with that on the widely used NUDT-MIRSDT dataset. Our MISTNet outperforms all other methods by a large margin, with an over 6% gain in the IoU metric, demonstrating its superiority. The dataset, code, and model weights are available at https://github.com/GR-ray/MIST. Meihong Zhang, Gongyang Li, Guanyi Li, Kai Zhao 0012, Xianchao Zhang 0002, Dan Zeng 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | MRI Super-Resolution With Partial Diffusion ModelsabstractDiffusion models have achieved impressive performance on various image generation tasks, including image super-resolution. Despite their impressive performance, diffusion models suffer from high computational costs due to the large number of denoising steps. In this paper, we proposed a novel accelerated diffusion model, termed Partial Diffusion Models (PDMs), for magnetic resonance imaging (MRI) super-resolution. We observed that the latents of diffusing a pair of low- and high-resolution images gradually converge and become indistinguishable after a certain noise level. This inspires us to use certain low-resolution latent to approximate corresponding high-resolution latent. With the approximation, we can skip part of the diffusion and denoising steps, reducing the computation in training and inference. To mitigate the approximation error, we further introduced 'latent alignment' that gradually interpolates and approaches the high-resolution latents from the low-resolution latents. Partial diffusion models, in conjunction with latent alignment, essentially establish a new trajectory where the latents, unlike those in original diffusion models, gradually transition from low-resolution to high-resolution images. Experiments on three MRI datasets demonstrate that partial diffusion models achieve competetive super-resolution quality with significantly fewer denoising steps than original diffusion models. In addition, they can be incorporated with recent accelerated diffusion models to further enhance the efficiency. Kai Zhao 0012, Kaifeng Pang, Alex Ling Yu Hung, Haoxin Zheng, Kyung Hyun Sung |
IEEE Trans. Medical Imaging | 1 |
| 2024 | Cross-Slice Attention and Evidential Critical Loss for Uncertainty-Aware Prostate Cancer Detection
Alex Ling Yu Hung, Haoxin Zheng, Kai Zhao 0012, Kaifeng Pang, Demetri Terzopoulos, Kyung Hyun Sung |
MICCAI (8) | 3 |
| 2024 | CSAM: A 2.5D Cross-Slice Attention Module for Anisotropic Volumetric Medical Image SegmentationabstractA large portion of volumetric medical data, especially magnetic resonance imaging (MRI) data, is anisotropic, as the through-plane resolution is typically much lower than the in-plane resolution. Both 3D and purely 2D deep learning-based segmentation methods are deficient in dealing with such volumetric data since the performance of 3D methods suffers when confronting anisotropic data, and 2D methods disregard crucial volumetric information. Insufficient work has been done on 2.5D methods, in which 2D convolution is mainly used in concert with volumetric information. These models focus on learning the relationship across slices, but typically have many parameters to train. We offer a Cross-Slice Attention Module (CSAM) with minimal trainable parameters, which captures information across all the slices in the volume by applying semantic, positional, and slice attention on deep feature maps at different scales. Our extensive experiments using different network architectures and tasks demonstrate the usefulness and generalizability of CSAM. Associated code is available at https://github.com/aL3x-O-o-Hung/CSAM. Alex Ling Yu Hung, Haoxin Zheng, Kai Zhao 0012, Xiaoxi Du, Kaifeng Pang, Qi Miao, Steven S. Raman, Demetri Terzopoulos, Kyung Hyun Sung |
WACV | 3 |
| 2024 | Adaptive feature alignment for adversarial trainingabstractRecent studies reveal that Convolutional Neural Networks (CNNs) are typically vulnerable to adversarial attacks . Many adversarial defense methods have been proposed to improve the robustness against adversarial samples. Moreover, these methods can only defend adversarial samples of a specific strength, reducing their flexibility against attacks of varying strengths. Moreover, these methods often enhance adversarial robustness at the expense of accuracy on clean samples. In this paper, we first observed that features of adversarial images change monotonically and smoothly w.r.t the rising of attacking strength. This intriguing observation suggests that features of adversarial images with various attacking strengths can be approximated by interpolating between the features of adversarial images with the strongest and weakest attacking strengths. Due to the monotonicity property, the interpolation weight can be easily learned by a neural network . Based on the observation, we proposed the adaptive feature alignment (AFA) that automatically align features to defense adversarial attacks of various attacking strengths. During training, our method learns the statistical information of adversarial samples with various attacking strengths using a dual batchnorm architecture. In this architecture, each batchnorm process handles samples of a specific attacking strength. During inference, our method automatically adjusts to varying attacking strengths by linearly interpolating the dual-BN features. Unlike previous methods that need to either retrain the model or manually tune hyper-parameters for a new attacking strength, our method can deal with arbitrary attacking strengths with a single model without introducing any hyper-parameter. Additionally, our method improves the model robustness against adversarial samples without incurring much loss of accuracy on clean images. Experiments on CIFAR-10, SVHN and tiny-ImageNet datasets demonstrate that our method outperforms the state-of-the-art under various attacking strengths and even improve accuracy on clean samples. Code will be made open available upon acceptance. Kai Zhao 0012, Wei Shen 0002 |
Pattern Recognit. Lett. | 1 |
| 2024 | Refining Uncertain Features With Self-Distillation for Face Recognition and Person Re-IdentificationabstractDeep recognition models aim to recognize targets with various quality levels in uncontrolled application circumstances, and typically low-quality images usually retard the recognition performance dramatically. As such, a straightforward solution is to restore low-quality input images as pre-processing during deployment. However, this scheme cannot guarantee that deep recognition features of the processed images are conducive to recognition accuracy. How deep recognition features of low-quality images can be refined during training to optimize recognition models has largely escaped research attention in the field of metric learning. In this paper, we propose a quality-aware feature refinement framework based on the dedicated quality priors obtained according to the recognition performance, and a novel quality self-distillation algorithm to learn recognition models. We further show that the proposed scheme can significantly boost the performance of the recognition model with two popular deep recognition tasks, including face recognition and person re-identification. Extensive experimental results provide sufficient evidence on the effectiveness and impressive generalization capability of the proposed framework. Moreover, our framework can be essentially integrated with existing state-of-the-art classification loss functions and network architectures, without extra computation costs during deployment. The source code is available athttps://github.com/oufuzhao/QSD Fu-Zhao Ou, Kai Zhao 0012, Shiqi Wang 0001, Yuan-Gen Wang, Sam Kwong |
IEEE Trans. Multim. | 3 |
| 2023 | RPG-Palm: Realistic Pseudo-data Generation for Palmprint RecognitionabstractPalmprint recently shows great potential in recognition applications as it is a privacy-friendly and stable biometric. However, the lack of large-scale public palmprint datasets limits further research and development of palmprint recognition. In this paper, we propose a novel realistic pseudo-palmprint generation (RPG) model to synthesize palmprints with massive identities. We first introduce a conditional modulation generator to improve the intra-class diversity. Then an identity-aware loss is proposed to ensure identity consistency against unpaired training. We further improve the Bézier palm creases generation strategy to guarantee identity independence. Extensive experimental results demonstrate that synthetic pretraining significantly boosts the recognition model performance. For example, our model improves the state-of-the-art BézierPalm by more than 5% and 14% in terms of TAR@FAR=1e-6 under the 1 : 1 and 1 : 3 Open-set protocol. When accessing only 10% of the real training data, our method still outperforms ArcFace with 100% real training data, indicating that we are closer to real-data-free palmprint recognition. Jianlong Jin, Huaen Li, Kai Zhao 0012, Shouhong Ding, Yang Zhao 0002, Wei Jia 0001 |
ICCV | 5 |
| 2022 | ContrastMask: Contrastive Learning to Segment Every ThingabstractPartially-supervised instance segmentation is a task which requests segmenting objects from novel categories via learning on limited base categories with annotated masks thus eliminating demands of heavy annotation burden. The key to addressing this task is to build an effective class-agnostic mask segmentation model. Unlike previous methods that learn such models only on base categories, in this paper, we propose a new method, named ContrastMask, which learns a mask segmentation model on both base and novel categories under a unified pixel-level contrastive learning framework. In this framework, annotated masks of base categories and pseudo masks of novel categories serve as a prior for contrastive learning, where features from the mask regions (foreground) are pulled together, and are contrasted against those from the background, and vice versa. Through this framework, feature discrimination between foreground and background is largely improved, facilitating learning of the class-agnostic mask segmentation model. Exhaustive experiments on the COCO dataset demonstrate the superiority of our method, which outperforms previous state-of-the-arts. Kai Zhao 0012, Shouhong Ding, Yan Wang 0033, Wei Shen 0002 |
CVPR | 2 |
| 2022 | BézierPalm: A Free Lunch for Palmprint Recognition
Kai Zhao 0012, Chuhan Zhou, Shouhong Ding, Wei Jia 0001, Wei Shen 0002 |
ECCV (13) | 1 |
| 2022 | Rethinking mask heads for partially supervised instance segmentation
Kai Zhao 0012, Wei Shen 0002 |
Neurocomputing | 1 |
| 2022 | Deep Hough Transform for Semantic Line DetectionabstractWe focus on a fundamental task of detecting meaningful line structures, a.k.a., semantic line, in natural scenes. Many previous methods regard this problem as a special case of object detection and adjust existing object detectors for semantic line detection. However, these methods neglect the inherent characteristics of lines, leading to sub-optimal performance. Lines enjoy much simpler geometric property than complex objects and thus can be compactly parameterized by a few arguments. To better exploit the property of lines, in this paper, we incorporate the classical Hough transform technique into deeply learned representations and propose a one-shot end-to-end learning framework for line detection. By parameterizing lines with slopes and biases, we perform Hough transform to translate deep representations into the parametric domain, in which we perform line detection. Specifically, we aggregate features along candidate lines on the feature map plane and then assign the aggregated features to corresponding locations in the parametric domain. Consequently, the problem of detecting semantic lines in the spatial domain is transformed into spotting individual points in the parametric domain, making the post-processing steps, i.e., non-maximal suppression, more efficient. Furthermore, our method makes it easy to extract contextual line features that are critical for accurate line detection. In addition to the proposed method, we design an evaluation metric to assess the quality of line detection and construct a large scale dataset for the line detection task. Experimental results on our proposed dataset and another public dataset demonstrate the advantages of our method over previous state-of-the-art alternatives. The dataset and source code is available at https://mmcheng.net/dhtline/. Kai Zhao 0012, Qi Han 0007, Chang-Bin Zhang, Jun Xu 0019, Ming-Ming Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Distribution alignment for cross-device palmprint recognition
Kai Zhao 0012, Wei Shen 0002 |
Pattern Recognit. | 3 |
| 2022 | Adaptive Material Matching for Hyperspectral Imagery DestripingabstractDue to instrument instability, slit contamination, and light interference, hyperspectral images often suffer from striping artifacts, which greatly impairs the data quality. Real hyperspectral data are usually characterized by a small amount of historical data, complex material distribution, insignificant periodicity of noise, and so on, which brings significant challenges for the destriping task. However, the assumptions made by traditional destriping methods are often inconsistent with these characteristics. To this end, we propose a novel destriping method based on adaptive material matching (MAM) without making explicit assumptions of hyperspectral data. Specifically, to identify pixels that belong to the same material, we propose a principal material analysis (PMA) to adaptively generate thresholds within each superpixel. The pixels are matched by thresholding their vertical gradients and leveraging both inner stripe gradient feature (ISGF) and neighbor-stripe geometry feature (NSGF). Correction pixels selected from the same material can then be used to calculate the offsets and gains of pixels to adjust adjacent columns. To further improve the stability of the destriping process, we generate a set of correction candidates for each column and select the optimal candidate by considering the prior distribution and destriping nonuniformity. The stripe noise within the whole image is finally removed by iteratively performing the correction between adjacent columns. We compare the proposed model against traditional and deep learning methods on both synthetic and real hyperspectral images. The promising results indicate that MAM can effectively remove the image stripes, retain original image information, and improve the nonuniformity. Jia Li 0032, Junjie Zhang 0002, Kai Zhao 0012, Dan Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Structured sparsification with joint optimization of group convolution and channel shuffleabstractRecent advances in convolutional neural networks (CNNs) usually come with the expense of excessive computational overhead and memory footprint. Network compression aims to alleviate this issue by training compact models with comparable performance. However, existing compression techniques either entail dedicated expert design or compromise with a moderate performance drop. In this paper, we propose a novel structured sparsification method for efficient network compression. The proposed method automatically induces structured sparsity on the convolutional weights, thereby facilitating the implementation of the compressed model with the highly-optimized group convolution. We further address the problem of inter-group communication with a learnable channel shuffle mechanism. The proposed approach can be easily applied to compress many network architectures with a negligible performance drop. Extensive experimental results and analysis demonstrate that our approach gives a competitive performance against the recent network compression counterparts with a sound accuracy-complexity trade-off. Xinyu Zhang 0023, Kai Zhao 0012, Taihong Xiao, Ming-Ming Cheng, Ming-Hsuan Yang 0001 |
UAI | 2 |
| 2021 | Res2Net: A New Multi-Scale Backbone ArchitectureabstractRepresenting features at multiple scales is of great importance for numerous vision tasks. Recent advances in backbone convolutional neural networks (CNNs) continually demonstrate stronger multi-scale representation ability, leading to consistent performance gains on a wide range of applications. However, most existing methods represent the multi-scale features in a layer-wise manner. In this paper, we propose a novel building block for CNNs, namely Res2Net, by constructing hierarchical residual-like connections within one single residual block. The Res2Net represents multi-scale features at a granular level and increases the range of receptive fields for each network layer. The proposed Res2Net block can be plugged into the state-of-the-art backbone CNN models, e.g., ResNet, ResNeXt, and DLA. We evaluate the Res2Net block on all these models and demonstrate consistent performance gains over baseline models on widely-used datasets, e.g., CIFAR-100 and ImageNet. Further ablation studies and experimental results on representative computer vision tasks, i.e., object detection, class activation mapping, and salient object detection, further verify the superiority of the Res2Net over the state-of-the-art baseline methods. The source code and trained models are available on https://mmcheng.net/res2net/. Shanghua Gao, Ming-Ming Cheng, Kai Zhao 0012, Xinyu Zhang 0023, Ming-Hsuan Yang 0001, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Deep Differentiable Random Forests for Age EstimationabstractAge estimation from facial images is typically cast as a label distribution learning or regression problem, since aging is a gradual progress. Its main challenge is the facial feature space w.r.t. ages is inhomogeneous, due to the large variation in facial appearance across different persons of the same age and the non-stationary property of aging. In this paper, we propose two Deep Differentiable Random Forests methods, Deep Label Distribution Learning Forest (DLDLF) and Deep Regression Forest (DRF), for age estimation. Both of them connect split nodes to the top layer of convolutional neural networks (CNNs) and deal with inhomogeneous data by jointly learning input-dependent data partitions at the split nodes and age distributions at the leaf nodes. This joint learning follows an alternating strategy: (1) Fixing the leaf nodes and optimizing the split nodes and the CNN parameters by Back-propagation; (2) Fixing the split nodes and optimizing the leaf nodes by Variational Bounding. Two Deterministic Annealing processes are introduced into the learning of the split and leaf nodes, respectively, to avoid poor local optima and obtain better estimates of tree parameters free of initial values. Experimental results show that DLDLF and DRF achieve state-of-the-art performance on three age estimation datasets. Wei Shen 0002, Yilu Guo, Yan Wang 0033, Kai Zhao 0012, Bo Wang 0044, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Deep Hough Transform for Semantic Line Detection
Qi Han 0007, Kai Zhao 0012, Jun Xu 0019, Ming-Ming Cheng |
ECCV (9) | 2 |
| 2019 | RegularFace: Deep Face Recognition via Exclusive RegularizationabstractWe consider the face recognition task where facial images of the same identity (person) is expected to be closer in the representation space, while different identities be far apart. Several recent studies encourage the intra-class compactness by developing loss functions that penalize the variance of representations of the same identity. In this paper, we propose the `exclusive regularization' that focuses on the other aspect of discriminability -- the inter-class separability, which is neglected in many recent approaches. The proposed method, named RegularFace, explicitly distances identities by penalizing the angle between an identity and its nearest neighbor, resulting in discriminative face representations. Our method has intuitive geometric interpretation and presents unique benefits that are absent in previous works. Quantitative comparisons against prior methods on several open benchmarks demonstrate the superiority of our method. In addition, our method is easy to implement and requires only a few lines of python code on modern deep learning frameworks. Kai Zhao 0012, Ming-Ming Cheng |
CVPR | 1 |
| 2019 | Optimizing the F-Measure for Threshold-Free Salient Object DetectionabstractCurrent CNN-based solutions to salient object detection (SOD) mainly rely on the optimization of cross-entropy loss (CELoss). Then the quality of detected saliency maps is often evaluated in terms of F-measure. In this paper, we investigate an interesting issue: can we consistently use the F-measure formulation in both training and evaluation for SOD? By reformulating the standard F-measure we propose the relaxed F-measure which is differentiable w.r.t the posterior and can be easily appended to the back of CNNs as the loss function. Compared to the conventional cross-entropy loss of which the gradients decrease dramatically in the saturated area, our loss function, named FLoss, holds considerable gradients even when the activation approaches the target. Consequently, the FLoss can continuously force the network to produce polarized activations. Comprehensive benchmarks on several popular datasets show that FLoss outperforms the state-of-the-art with a considerable margin. More specifically, due to the polarized predictions, our method is able to obtain high-quality saliency maps without carefully tuning the optimal threshold, showing significant advantages in real-world applications. Kai Zhao 0012, Shanghua Gao, Wenguan Wang, Ming-Ming Cheng |
ICCV | 1 |
| 2018 | Deep Regression Forests for Age EstimationabstractAge estimation from facial images is typically cast as a nonlinear regression problem. The main challenge of this problem is the facial feature space w.r.t. ages is inhomogeneous, due to the large variation in facial appearance across different persons of the same age and the non-stationary property of aging patterns. In this paper, we propose Deep Regression Forests (DRFs), an end-to-end model, for age estimation. DRFs connect the split nodes to a fully connected layer of a convolutional neural network (CNN) and deal with inhomogeneous data by jointly learning input-dependant data partitions at the split nodes and data abstractions at the leaf nodes. This joint learning follows an alternating strategy: First, by fixing the leaf nodes, the split nodes as well as the CNN parameters are optimized by Back-propagation; Then, by fixing the split nodes, the leaf nodes are optimized by iterating a step-size free update rule derived from Variational Bounding. We verify the proposed DRFs on three standard age estimation benchmarks and achieve state-of-the-art results on all of them. Wei Shen 0002, Yilu Guo, Yan Wang 0033, Kai Zhao 0012, Bo Wang 0044, Alan L. Yuille |
CVPR | 4 |
| 2018 | Hi-Fi: Hierarchical Feature Integration for Skeleton DetectionabstractIn natural images, the scales (thickness) of object skeletons may dramatically vary among objects and object parts. Thus, robust skeleton detection requires powerful multi-scale feature integration ability. To address this issue, we present a new convolutional neural network (CNN) architecture by introducing a novel hierarchical feature integration mechanism, named Hi-Fi, to address the object skeleton detection problem. The proposed CNN-based approach intrinsically captures high-level semantics from deeper layers, as well as low-level details from shallower layers. By hierarchically integrating different CNN feature levels with bidirectional guidance, our approach (1) enables mutual refinement across features of different levels, and (2) possesses the strong ability to capture both rich object context and high-resolution details. Experimental results show that our method significantly outperforms the state-of-the-art methods in terms of effectively fusing features from very different scales, as evidenced by a considerable performance improvement on several benchmarks. Kai Zhao 0012, Wei Shen 0002, Shanghua Gao, Ming-Ming Cheng |
IJCAI | 1 |
| 2017 | Label Distribution Learning ForestsabstractLabel distribution learning (LDL) is a general learning framework, which assigns to an instance a distribution over a set of labels rather than a single label or multiple labels. Current LDL methods have either restricted assumptions on the expression form of the label distribution or limitations in representation learning, e.g., to learn deep features in an end-to-end manner. This paper presents label distribution learning forests (LDLFs) - a novel label distribution learning algorithm based on differentiable decision trees, which have several advantages: 1) Decision trees have the potential to model any general form of label distributions by a mixture of leaf node predictions. 2) The learning of differentiable decision trees can be combined with representation learning. We define a distribution-based loss function for a forest, enabling all the trees to be learned jointly, and show that an update function for leaf node predictions, which guarantees a strict decrease of the loss function, can be derived by variational bounding. The effectiveness of the proposed LDLFs is verified on several LDL tasks and a computer vision application, showing significant improvements to the state-of-the-art LDL methods. Wei Shen 0002, Kai Zhao 0012, Yilu Guo, Alan L. Yuille |
NIPS | 2 |
| 2017 | DeepSkeleton: Learning Multi-Task Scale-Associated Deep Side Outputs for Object Skeleton Extraction in Natural ImagesabstractObject skeletons are useful for object representation and object detection. They are complementary to the object contour, and provide extra information, such as how object scale (thickness) varies among object parts. But object skeleton extraction from natural images is very challenging, because it requires the extractor to be able to capture both local and non-local image context in order to determine the scale of each skeleton pixel. In this paper, we present a novel fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the different layers in the network and the skeleton scales they can capture, we introduce two scale-associated side outputs to each stage of the network. The network is trained by multi-task learning, where one task is skeleton localization to classify whether a pixel is a skeleton pixel or not, and the other is skeleton scale prediction to regress the scale of each skeleton pixel. Supervision is imposed at different stages by guiding the scale-associated side outputs toward the ground-truth skeletons at the appropriate scales. The responses of the multiple scale-associated side outputs are then fused in a scale-specific way to detect skeleton pixels using multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors. In addition, the usefulness of the obtained skeletons and scales (thickness) are verified on two object detection applications: foreground object segmentation and object proposal detection. Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Xiang Bai, Alan L. Yuille |
IEEE Trans. Image Process. | 2 |
| 2016 | Object Skeleton Extraction in Natural Images by Fusing Scale-Associated Deep Side OutputsabstractObject skeleton is a useful cue for object detection, complementary to the object contour, as it provides a structural representation to describe the relationship among object parts. While object skeleton extraction in natural images is a very challenging problem, as it requires the extractor to be able to capture both local and global image context to determine the intrinsic scale of each skeleton pixel. Existing methods rely on per-pixel based multi-scale feature computation, which results in difficult modeling and high time consumption. In this paper, we present a fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the sequential stages in the network and the skeleton scales they can capture, we introduce a scale-associated side output to each stage. We impose supervision to different stages by guiding the scale-associated side outputs toward groundtruth skeletons of different scales. The responses of the multiple scaleassociated side outputs are then fused in a scale-specific way to localize skeleton pixels with multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors. Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Zhijiang Zhang, Xiang Bai |
CVPR | 2 |