EDBT 2026 Demo / reviewers in the wild / expert
Hao Lu 0003
dblp:72/5422-3
· DBLP profile ↗
67ranked-venue papers
12as first author
49since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 5 first-author · 31 since 2021Artificial intelligence and machine learning · 43 · 8 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task-aware dynamic routing network for cross-domain few-shot learning
Yanan Li 0006, Haoyang Ye, Huabing Zhou, Tao Lu 0001, Hao Lu 0003 |
Neurocomputing | 5 |
| 2026 | Scalable portrait matte creation with layer diffusion and connectivity priors
Hao Lu 0003, Hua Huang 0001 |
Pattern Recognit. | 2 |
| 2026 | Densely activated self-attention for semantic segmentation
Liwen Xiao, Wenze Liu, Zhicheng Wang 0002, Yiran Wang 0005, Hao Lu 0003, Zhiguo Cao 0001 |
Pattern Recognit. | 6 |
| 2025 | Training Matting Models Without Alpha LabelsabstractThe labeling difficulty has been a longstanding problem in deep image matting. To escape from fine labels, this work explores using rough annotations such as trimaps coarsely indicating the foreground/background as supervision. We present that the cooperation between learned semantics from indicated known regions and proper assumed matting rules can help infer alpha values at transition areas. Inspired by the nonlocal principle in traditional image matting, we build a directional distance consistency loss (DDC loss) at each pixel neighborhood to constrain the alpha values conditioned on the input image. DDC loss forces the distance of similar pairs on the alpha matte and on its corresponding image to be consistent. In this way, the alpha values can be propagated from learned known regions to unknown transition areas. With only images and trimaps, a matting model can be trained under the supervision of a known loss and the proposed DDC loss. Experiments on AM-2K and P3M-10K dataset show that our paradigm achieves comparable performance with the fine-label-supervised baseline, while sometimes offers even more satisfying results than human-labeled ground truth. Wenze Liu, Zixuan Ye, Hao Lu 0003, Zhiguo Cao 0001, Xiangyu Yue 0001 |
AAAI | 3 |
| 2025 | MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation ModelsabstractRecent advances in generalizable 3D Gaussian Splatting have demonstrated promising results in real-time high-fidelity rendering without per-scene optimization, yet existing approaches still struggle to handle unfamiliar visual content during inference on novel scenes due to limited generalizability. To address this challenge, we introduce MonoSplat, a novel framework that leverages rich visual priors from pre-trained monocular depth foundation models for robust Gaussian reconstruction. Our approach consists of two key components: a Mono-Multi Feature Adapter that transforms monocular features into multi-view representations, coupled with an Integrated Gaussian Prediction module that effectively fuses both feature types for precise Gaussian generation. Through the Adapter’s lightweight attention mechanism, features are seamlessly aligned and aggregated across views while preserving valuable monocular priors, enabling the Prediction module to generate Gaussian primitives with accurate geometry and appearance. Through extensive experiments on diverse real-world datasets, we convincingly demonstrate that MonoSplat achieves superior reconstruction quality and generalization capability compared to existing methods while maintaining computational efficiency with minimal trainable parameters. Codes are available at https://github.com/CUHK-AIM-Group/MonoSplat. Yifan Liu 0010, Keyu Fan, Weihao Yu 0005, Chenxin Li, Hao Lu 0003, Yixuan Yuan |
CVPR | 5 |
| 2025 | OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and ReasoningabstractScoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks ($4\times$ more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios ($31$ diverse scenarios), and thorough evaluation metrics, with $10,000$ human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with $1,500$ manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below $50$ ($100$ in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The benchmark and evaluation scripts are available at https://github.com/Yuliang-Liu/MultimodalOCR. Zhebin Kuang, Jiajun Song, Mingxin Huang, Linghao Zhu, Qidi Luo, Xinyu Wang 0010, Hao Lu 0003, Guozhi Tang, Bin Shan, Chunhui Lin, Binghong Wu, Hao Feng 0009, Hao Liu 0003, Can Huang 0002, Jingqun Tang, Wei Chen 0088, Xiang Bai |
NeurIPS | 10 |
| 2025 | FADE: A Task-Agnostic Upsampling Operator for Encoder-Decoder Architectures
Hao Lu 0003, Wenze Liu, Hongtao Fu, Zhiguo Cao 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-Domain GeneralizationabstractText spotting, a task involving the extraction of textual information from image or video sequences, faces challenges in cross-domain adaption, such as image-to-image and image-to-video generalization. In this paper, we introduce a new method, termed VimTS, which enhances the generalization ability of the model by achieving better synergy among different tasks. Typically, we propose a Prompt Queries Generation Module and a Tasks-aware Adapter to effectively convert the original single-task model into a multi-task model suitable for both image and video scenarios with minimal additional parameters. The Prompt Queries Generation Module facilitates explicit interaction between different tasks, while the Tasks-aware Adapter helps the model dynamically learn suitable features for each task. Additionally, to further enable the model to learn temporal information at a lower cost, we propose a synthetic video text dataset (VTD-368 k) by leveraging the Content Deformation Fields (CoDeF) algorithm. Notably, our method outperforms the state-of-the-art method by an average of 2.6% in six cross-domain benchmarks such as TT-to-IC15, CTW1500-to-TT, and TT-to-CTW1500. For video-level cross-domain adaption, our method even surpasses the previous end-to-end video spotting method in ICDAR2015 video and DSText v2 by an average of 5.5% on the MOTA metric, using only image-level data. We further demonstrate that existing Large Multimodal Models exhibit limitations in generating cross-domain scene text spotting, in contrast to our VimTS model which requires significantly fewer parameters and data. Mingxin Huang, Linger Deng, Weijia Wu 0001, Hao Lu 0003, Chunhua Shen, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Vision Transformer Off-the-Shelf: A Surprising Baseline for Few-Shot Class-Agnostic CountingabstractClass-agnostic counting (CAC) aims to count objects of interest from a query image given few exemplars. This task is typically addressed by extracting the features of query image and exemplars respectively and then matching their feature similarity, leading to an extract-then-match paradigm. In this work, we show that CAC can be simplified in an extract-and-match manner, particularly using a vision transformer (ViT) where feature extraction and similarity matching are executed simultaneously within the self-attention. We reveal the rationale of such simplification from a decoupled view of the self-attention.The resulting model, termed CACViT, simplifies the CAC pipeline into a single pretrained plain ViT. Further, to compensate the loss of the scale and the order-of-magnitude information due to resizing and normalization in plain ViT, we present two effective strategies for scale and magnitude embedding. Extensive experiments on the FSC147 and the CARPK datasets show that CACViT significantly outperforms state-of-the-art CAC approaches in both effectiveness (23.60% error reduction) and generalization, which suggests CACViT provides a concise and strong baseline for CAC. Code will be available. Zhicheng Wang 0002, Liwen Xiao, Zhiguo Cao 0001, Hao Lu 0003 |
AAAI | 4 |
| 2024 | In-Context Matting
He Guo 0005, Zixuan Ye, Zhiguo Cao 0001, Hao Lu 0003 |
CVPR | 4 |
| 2024 | Unifying Automatic and Interactive Matting with Pretrained ViTsabstractAutomatic and interactive matting largely improve image matting by respectively alleviating the need for auxil-iary input and enabling object selection. Due to different settings on whether prompts exist, they either suffer from weakness in instance completeness or region details. Also, when dealing with different scenarios, directly switching between the two matting models introduces inconvenience and higher workload. Therefore, we wonder whether we can al-leviate the limitations of both settings while achieving unification to facilitate more convenient use. Our key idea is to offer saliency guidance for automatic mode to enable its attention to detailed regions, and also refine the instance completeness in interactive mode by replacing the binary mask guidance with a more probabilistic form. With different guidance for each mode, we can achieve unification through adaptable guidance, defined as saliency information in automatic mode and user cue for interactive one. It is instantiated as candidate feature in our method, an automatic switch for class token in pretrained ViTs and average feature of user prompts, controlled by the existence of user prompts. Then we use the candidate feature to generate a probabilistic similarity map as the guidance to alleviate the over-reliance on binary mask. Extensive experiments show that our method can adapt well to both automatic and inter-active scenarios with more light-weight framework. Code available at github.com/coconut/SMat. Zixuan Ye, Wenze Liu, He Guo 0005, Yujia Liang, Chaoyi Hong, Hao Lu 0003, Zhiguo Cao 0001 |
CVPR | 6 |
| 2024 | SCAPE: A Simple and Strong Category-Agnostic Pose Estimator
Yujia Liang, Zixuan Ye, Wenze Liu, Hao Lu 0003 |
ECCV (23) | 4 |
| 2024 | Counting Crowd by Weighing Counts: A Sequential Decision-Making PerspectiveabstractWe show that crowd counting can be formulated as a sequential decision-making (SDM) problem. Inspired by human counting, we evade one-step estimation mostly executed in existing counting models and decompose counting into sequential sub-decision problems. During implementation, a key insight is to interpret sequential counting as a physical process in reality-scale weighing. This analogy allows us to implement a novel "counting scale" termed LibraNet. Our idea is that, by placing a crowd image on the scale, LibraNet (agent) learns to place appropriate weights to match the count: at each step, one weight (action) is chosen from the weight box (the predefined action pool) conditioned on the image features and the placed weights (state) until the pointer (the agent output) informs balance. We investigate two forms of state definition and explore four types of LibraNet implementations under different learning paradigms, including deep Q-network (DQN), actor-critic (AC), imitation learning (IL), and mixed AC+IL. Experiments show that LibraNet indeed mimics scale weighing, that it outperforms or performs comparably against state-of-the-art approaches on five crowd counting benchmarks, that it can be used as a plug-in to improve off-the-shelf counting models, and particularly that it demonstrates remarkable cross-dataset generalization. Code and models are available at https://git.io/libranet. Hao Lu 0003, Liang Liu 0001, Hu Wang 0005, Zhiguo Cao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Find Beauty in the Rare: Contrastive Composition Feature Clustering for Nontrivial Cropping Box RegressionabstractAutomatic image cropping algorithms aim to recompose images like human-being photographers by generating the cropping boxes with improved composition quality. Cropping box regression approaches learn the beauty of composition from annotated cropping boxes. However, the bias of annotations leads to quasi-trivial recomposing results, which has an obvious tendency to the average location of training samples. The crux of this predicament is that the task is naively treated as a box regression problem, where rare samples might be dominated by normal samples, and the composition patterns of rare samples are not well exploited. Observing that similar composition patterns tend to be shared by the cropping boundaries annotated nearly, we argue to find the beauty of composition from the rare samples by clustering the samples with similar cropping boundary annotations, i.e., similar composition patterns. We propose a novel Contrastive Composition Clustering (C2C) to regularize the composition features by contrasting dynamically established similar and dissimilar pairs. In this way, common composition patterns of multiple images can be better summarized, which especially benefits the rare samples and endows our model with better generalizability to render nontrivial results. Extensive experimental results show the superiority of our model compared with prior arts. We also illustrate the philosophy of our design with an interesting analytical visualization. Yinpeng Chen, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong |
AAAI | 4 |
| 2023 | Infusing Definiteness into Randomness: Rethinking Composition Styles for Deep Image MattingabstractWe study the composition style in deep image matting, a notion that characterizes a data generation flow on how to exploit limited foregrounds and random backgrounds to form a training dataset. Prior art executes this flow in a completely random manner by simply going through the foreground pool or by optionally combining two foregrounds before foreground-background composition. In this work, we first show that naive foreground combination can be problematic and therefore derive an alternative formulation to reasonably combine foregrounds. Our second contribution is an observation that matting performance can benefit from a certain occurrence frequency of combined foregrounds and their associated source foregrounds during training. Inspired by this, we introduce a novel composition style that binds the source and combined foregrounds in a definite triplet. In addition, we also find that different orders of foreground combination lead to different foreground patterns, which further inspires a quadruplet-based composition style. Results under controlled experiments on four matting baselines show that our composition styles outperform existing ones and invite consistent performance improvement on both composited and real-world datasets. Code is available at: https://github.com/coconuthust/composition_styles Zixuan Ye, Yutong Dai 0001, Chaoyi Hong, Zhiguo Cao 0001, Hao Lu 0003 |
AAAI | 5 |
| 2023 | Learning Second-Order Attentive Context for Efficient Correspondence PruningabstractCorrespondence pruning aims to search consistent correspondences (inliers) from a set of putative correspondences. It is challenging because of the disorganized spatial distribution of numerous outliers, especially when putative correspondences are largely dominated by outliers. It's more challenging to ensure effectiveness while maintaining efficiency. In this paper, we propose an effective and efficient method for correspondence pruning. Inspired by the success of attentive context in correspondence problems, we first extend the attentive context to the first-order attentive context and then introduce the idea of attention in attention (ANA) to model second-order attentive context for correspondence pruning. Compared with first-order attention that focuses on feature-consistent context, second-order attention dedicates to attention weights itself and provides an additional source to encode consistent context from the attention map. For efficiency, we derive two approximate formulations for the naive implementation of second-order attention to optimize the cubic complexity to linear complexity, such that second-order attention can be used with negligible computational overheads. We further implement our formulations in a second-order context layer and then incorporate the layer in an ANA block. Extensive experiments demonstrate that our method is effective and efficient in pruning outliers, especially in high-outlier-ratio cases. Compared with the state-of-the-art correspondence pruning approach LMCNet, our method runs 14 times faster while maintaining a competitive accuracy. Weiyue Zhao, Hao Lu 0003, Zhiguo Cao 0001 |
AAAI | 3 |
| 2023 | ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in TransformerabstractIn recent years, end-to-end scene text spotting approaches are evolving to the Transformer-based framework. While previous studies have shown the crucial importance of the intrinsic synergy between text detection and recognition, recent advances in Transformer-based methods usually adopt an implicit synergy strategy with shared query, which can not fully realize the potential of these two interactive tasks. In this paper, we argue that the explicit synergy considering distinct characteristics of text detection and recognition can significantly improve the performance text spotting. To this end, we introduce a new model named Explicit Synergy-based Text Spotting Transformer framework (ESTextSpotter), which achieves explicit synergy by modeling discriminative and interactive features for text detection and recognition within a single decoder. Specifically, we decompose the conventional shared query into task-aware queries for text polygon and content, respectively. Through the decoder with the proposed vision-language communication module, the queries interact with each other in an explicit manner while preserving discriminative patterns of text detection and recognition, thus improving performance significantly. Additionally, we propose a task-aware query initialization scheme to ensure stable training. Experimental results demonstrate that our model significantly outperforms previous state-of-the-art methods. Code is available at https://github.com/mxin262/ESTextSpotter. Mingxin Huang, Jiaxin Zhang 0003, Dezhi Peng, Hao Lu 0003, Can Huang 0002, Xiang Bai |
ICCV | 4 |
| 2023 | Learning to Upsample by Learning to SampleabstractWe present DySample, an ultra-lightweight and effective dynamic upsampler. While impressive performance gains have been witnessed from recent kernel-based dynamic upsamplers such as CARAFE, FADE, and SAPA, they introduce much workload, mostly due to the time-consuming dynamic convolution and the additional sub-network used to generate dynamic kernels. Further, the need for high-res feature guidance of FADE and SAPA somehow limits their application scenarios. To address these concerns, we bypass dynamic convolution and formulate upsampling from the perspective of point sampling, which is more resource-efficient and can be easily implemented with the standard built-in function in PyTorch. We first showcase a naive design, and then demonstrate how to strengthen its upsampling behavior step by step towards our new upsampler, DySample. Compared with former kernel-based dynamic upsamplers, DySample requires no customized CUDA package and has much fewer parameters, FLOPs, GPU memory, and latency. Besides the light-weight characteristics, DySample outperforms other upsamplers across five dense prediction tasks, including semantic segmentation, object detection, instance segmentation, panoptic segmentation, and monocular depth estimation. Code is available at https://github.com/tiny-smart/dysample. Wenze Liu, Hao Lu 0003, Hongtao Fu, Zhiguo Cao 0001 |
ICCV | 2 |
| 2023 | Point-Query Quadtree for Crowd Counting, Localization, and MoreabstractWe show that crowd counting can be viewed as a decomposable point querying process. This formulation enables arbitrary points as input and jointly reasons whether the points are crowd and where they locate. The querying processing, however, raises an underlying problem on the number of necessary querying points. Too few imply underestimation; too many increase computational overhead. To address this dilemma, we introduce a decomposable structure, i.e., the point-query quadtree, and propose a new counting model, termed Point quEry Transformer (PET). PET implements decomposable point querying via data-dependent quadtree splitting, where each querying point could split into four new points when necessary, thus enabling dynamic processing of sparse and dense regions. Such a querying process yields an intuitive, universal modeling of crowd as both the input and output are interpretable and steerable. We demonstrate the applications of PET on a number of crowd-related tasks, including fully-supervised crowd counting and localization, partial annotation learning, and point annotation refinement, and also report state-of-the-art performance. For the first time, we show that a single counting model can address multiple crowd-related tasks across different learning paradigms. Code is available at https://github.com/cxliu0/PET. Hao Lu 0003, Zhiguo Cao 0001, Tongliang Liu |
ICCV | 2 |
| 2023 | Fast Full-frame Video Stabilization with Iterative OptimizationabstractVideo stabilization refers to the problem of transforming a shaky video into a visually pleasing one. The question of how to strike a good trade-off between visual quality and computational speed has remained one of the open challenges in video stabilization. Inspired by the analogy between wobbly frames and jigsaw puzzles, we propose an iterative optimization-based learning approach using synthetic datasets for video stabilization, which consists of two interacting submodules: motion trajectory smoothing and full-frame outpainting. First, we develop a two-level (coarse-to-fine) stabilizing algorithm based on the probabilistic flow field. The confidence map associated with the estimated optical flow is exploited to guide the search for shared regions through backpropagation. Second, we take a divide-and-conquer approach and propose a novel multi-frame fusion strategy to render full-frame stabilized views. An important new insight brought about by our iterative optimization approach is that the target video can be interpreted as the fixed point of nonlinear mapping for video stabilization. We formulate video stabilization as a problem of minimizing the amount of jerkiness in motion trajectories, which guarantees convergence with the help of fixed-point theory. Extensive experimental results are reported to demonstrate the superiority of the proposed approach in terms of computational speed and visual quality. The code will be available on GitHub. Weiyue Zhao, Xin Li 0005, Xianrui Luo, Hao Lu 0003, Zhiguo Cao 0001 |
ICCV | 6 |
| 2023 | SIERRA: A robust bilateral feature upsampler for dense prediction
Hongtao Fu, Wenze Liu, Zhiguo Cao 0001, Hao Lu 0003 |
Comput. Vis. Image Underst. | 5 |
| 2023 | From Open Set to Closed Set: Supervised Spatial Divide-and-Conquer for Object Counting
Haipeng Xiong, Hao Lu 0003, Liang Liu 0001, Chunhua Shen, Zhiguo Cao 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | A2B: Anchor to Barycentric Coordinate for Robust Correspondence
Weiyue Zhao, Hao Lu 0003, Zhiguo Cao 0001, Xin Li 0005 |
Int. J. Comput. Vis. | 2 |
| 2023 | Accurate Robotic Grasp Detection with Angular Label Smoothing
Min Shi 0005, Hao Lu 0003, Zhao-Xin Li, Dengming Zhu, Zhao-Qi Wang |
J. Comput. Sci. Technol. | 2 |
| 2023 | Learning Probabilistic Coordinate Fields for Robust CorrespondencesabstractWe introduce Probabilistic Coordinate Fields (PCFs), a novel geometric-invariant coordinate representation for image correspondence problems. In contrast to standard Cartesian coordinates, PCFs encode coordinates in correspondence-specific barycentric coordinate systems (BCS) with affine invariance. To know when and where to trust the encoded coordinates, we implement PCFs in a probabilistic network termed PCF-Net, which parameterizes the distribution of coordinate fields as Gaussian mixture models. By jointly optimizing coordinate fields and their confidence conditioned on dense flows, PCF-Net can work with various feature descriptors when quantifying the reliability of PCFs by confidence maps. An interesting observation of this work is that the learned confidence map converges to geometrically coherent and semantically consistent regions, which facilitates robust coordinate representation. By delivering the confident coordinates to keypoint/feature descriptors, we show that PCF-Net can be used as a plug-in to existing correspondence-dependent approaches. Extensive experiments on both indoor and outdoor datasets suggest that accurate geometric invariant coordinates help to achieve the state of the art in several correspondence problems, such as sparse feature matching, dense image registration, camera pose estimation, and consistency filtering. Further, the interpretable confidence map predicted by PCF-Net can also be leveraged to other novel applications from texture transfer to multi-homography classification. Weiyue Zhao, Hao Lu 0003, Zhiguo Cao 0001, Xin Li 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Point-and-Shoot All-in-Focus Photo Synthesis From Smartphone Camera PairabstractAll-in-Focus (AIF) photography is expected to be a commercial selling point for modern smartphones. Standard AIF synthesis requires manual, time-consuming operations such as focal stack compositing, which is unfriendly to ordinary people. To achieve point-and-shoot AIF photography with a smartphone, we expect that an AIF photo can be generated from one shot of the scene, instead of from multiple photos captured by the same camera. Benefiting from the multi-camera module in modern smartphones, we introduce a new task of AIF synthesis from main (wide) and ultra-wide cameras. The goal is to recover sharp details from defocused regions in the main-camera photo with the help of the ultra-wide-camera one. The camera setting poses new challenges such as parallax-induced occlusions and inconsistent color between cameras. To overcome the challenges, we introduce a predict-and-refine network to mitigate occlusions and propose dynamic frequency-domain alignment for color correction. To enable effective training and evaluation, we also build an AIF dataset with 2686 unique scenes. Each scene includes two photos captured by the main camera, one photo captured by the ultra-wide camera, and a synthesized AIF photo. Results show that our solution, termed EasyAIF, can produce high-quality AIF photos and outperforms strong baselines quantitatively and qualitatively. For the first time, we demonstrate point-and-shoot AIF photo synthesis successfully from main and ultra-wide cameras. Xianrui Luo, Juewen Peng, Weiyue Zhao, Ke Xian, Hao Lu 0003, Zhiguo Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | BokehMe: When Neural Rendering Meets Classical RenderingabstractWe propose BokehMe, a hybrid bokeh rendering framework that marries a neural renderer with a classical physically motivated renderer. Given a single image and a potentially imperfect disparity map, BokehMe generates high-resolution photo-realistic bokeh effects with adjustable blur size, focal plane, and aperture shape. To this end, we analyze the errors from the classical scattering-based method and derive a formulation to calculate an error map. Based on this formulation, we implement the classical renderer by a scattering-based method and propose a two-stage neural renderer to fix the erroneous areas from the classical renderer. The neural renderer employs a dynamic multi-scale scheme to efficiently handle arbitrary blur sizes, and it is trained to handle imperfect disparity input. Experiments show that our method compares favorably against previous methods on both synthetic image data and real image data with predicted disparity. A user study is further conducted to validate the advantage of our method. Juewen Peng, Zhiguo Cao 0001, Xianrui Luo, Hao Lu 0003, Ke Xian, Jianming Zhang 0001 |
CVPR | 4 |
| 2022 | Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingabstractClass-agnostic counting (CAC) aims to count all instances in a query image given few exemplars. A standard pipeline is to extract visual features from exemplars and match them with query images to infer object counts. Two essential components in this pipeline are feature representation and similarity metric. Existing methods either adopt a pretrained network to represent features or learn a new one, while applying a naive similarity metric with fixed inner product. We find this paradigm leads to noisy similarity matching and hence harms counting performance. In this work, we propose a similarity-aware CAC framework that jointly learns representation and similarity metric. We first instantiate our framework with a naive baseline called Bilinear Matching Network (BMNet), whose key component is a learnable bilinear similarity metric. To further embody the core of our framework, we extend BMNet to BMNet+ that models similarity from three aspects: 1) representing the instances via their self-similarity to enhance feature robustness against intra-class variations; 2) comparing the similarity dynamically to focus on the key patterns of each exemplar; 3) learning from a supervision signal to impose explicit constraints on matching results. Extensive experiments on a recent CAC dataset FSC147 show that our models significantly outperform state-of-the-art CAC approaches. In addition, we also validate the cross-dataset generality of BMNet and BMNet+ on a car counting dataset CARPK. Code is at tiny.one/BMNet Min Shi 0004, Hao Lu 0003, Chen Feng 0002, Zhiguo Cao 0001 |
CVPR | 2 |
| 2022 | FADE: Fusing the Assets of Decoder and Encoder for Task-Agnostic Upsampling
Hao Lu 0003, Wenze Liu, Hongtao Fu, Zhiguo Cao 0001 |
ECCV (27) | 1 |
| 2022 | Robust Object Detection with Inaccurate Bounding Boxes
Kewei Wang 0001, Hao Lu 0003, Zhiguo Cao 0001 |
ECCV (10) | 3 |
| 2022 | MPIB: An MPI-Based Bokeh Rendering Framework for Realistic Partial Occlusion Effects
Juewen Peng, Jianming Zhang 0001, Xianrui Luo, Hao Lu 0003, Ke Xian, Zhiguo Cao 0001 |
ECCV (6) | 4 |
| 2022 | 3D Instances as 1D Kernels
Yizheng Wu, Min Shi 0004, Shuaiyuan Du, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong |
ECCV (29) | 4 |
| 2022 | Hyprogan: Breaking the Dimensional wall From Human to AnimeabstractImage translation from human faces to anime ones brings a low-end, efficient way to create animation characters for animation industry. However, due to the significant inter-domain difference between anime images and human photos, existing image-to-image translation approaches cannot address this task well. To solve this dilemma, we propose HyProGAN, an exemplar-guided image-to-image translation model without paired data. The key contribution of HyPro-GAN is that it introduces a novel hybrid and progressive training strategy that expands the unidirectional translation between two domains into the bidirectional intra-domain and inter-domain translation. To enhance the consistency between input and output, we further propose a local masking loss to align the facial features between the human face and the generated anime face. Extensive experiments demonstrate the superiority of HyProGAN against state-of-the-art models. Yinpeng Chen, Zhiguo Cao 0001, Hao Lu 0003, Weicai Zhong |
ICIP | 4 |
| 2022 | Discriminate Clearer To Rank Better: Image Cropping By Amplifying View-Wise DifferencesabstractImage cropping aims to enhance the aesthetic quality of a given image by searching for the good cropping views. One common routine is to score and rank the candidate views by the neural network. The network is expected to discriminate the subtle view-wise differences. However, the image-wise differences and the ambiguity in the annotations render difficulties in discriminating the view-wise differences. To focus on the view-wise differences, we propose a feature spliter to build image-wise and view-wise feature and evaluate the candidate views only based on the view-wise feature. Then, we propose the ranking gain loss that alleviates the ambiguity in annotations to amplify the view-wise differences. The remarkable improvement compared with prior arts on public benchmarks illustrates that the view-wise differences matter in cropping view recommendation. Zhiguo Cao 0001, Ke Xian, Hao Lu 0003, Weicai Zhong |
ICIP | 4 |
| 2022 | Design What You Desire: Icon Generation from Orthogonal Application and Theme LabelsabstractGenerative adversarial networks,(GANs) have been trained to be professional artists able to create stunning artworks such as face generation and image style transfer. In this paper, we focus on a realistic business scenario: automated generation of customizable icons given desired mobile applications and theme styles. We first introduce a theme-application icon dataset, termed AppIcon, where each icon has two orthogonal theme and app labels. By investigating a strong baseline StyleGAN2, we observe mode collapse caused by the entanglement of the orthogonal labels. To solve this challenge, we propose IconGAN composed of a conditional generator and dual discriminators with orthogonal augmentations, and a contrastive feature disentanglement strategy is further designed to regularize the feature space of the two discriminators. Compared with other approaches, IconGAN indicates a superior advantage on the AppIcon benchmark. Further analysis also justifies the effectiveness of disentangling app and theme representations. Our project will be released at: https://github.com/architect-road/IconGAN. Yinpeng Chen, Min Shi 0004, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong |
ACM Multimedia | 4 |
| 2022 | DoF-NeRF: Depth-of-Field Meets Neural Radiance FieldsabstractNeural Radiance Field (NeRF) and its variants have exhibited great success on representing 3D scenes and synthesizing photo-realistic novel views. However, they are generally based on the pinhole camera model and assume all-in-focus inputs. This limits their applicability as images captured from the real world often have finite depth-of-field (DoF). To mitigate this issue, we introduce DoF-NeRF, a novel neural rendering approach that can deal with shallow DoF inputs and can simulate DoF effect. In particular, it extends NeRF to simulate the aperture of lens following the principles of geometric optics. Such a physical guarantee allows DoF-NeRF to operate views with different focus configurations. Benefiting from explicit aperture modeling, DoF-NeRF also enables direct manipulation of DoF effect by adjusting virtual aperture and focus parameters. It is plug-and-play and can be inserted into NeRF-based frameworks. Experiments on synthetic and real-world datasets show that, DoF-NeRF not only performs comparably with NeRF in the all-in-focus setting, but also can synthesize all-in-focus novel views conditioned on shallow DoF inputs. An interesting application of DoF-NeRF to DoF rendering is also demonstrated. The source code will be made available at: https://github.com/zijinwuzijin/DoF-NeRF. Zijin Wu, Xingyi Li 0005, Juewen Peng, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong |
ACM Multimedia | 4 |
| 2022 | SAPA: Similarity-Aware Point Affiliation for Feature UpsamplingabstractWe introduce point affiliation into feature upsampling, a notion that describes the affiliation of each upsampled point to a semantic cluster formed by local decoder feature points with semantic similarity. By rethinking point affiliation, we present a generic formulation for generating upsampling kernels. The kernels encourage not only semantic smoothness but also boundary sharpness in the upsampled feature maps. Such properties are particularly useful for some dense prediction tasks such as semantic segmentation. The key idea of our formulation is to generate similarity-aware kernels by comparing the similarity between each encoder feature point and the spatially associated local region of decoder features. In this way, the encoder feature point can function as a cue to inform the semantic cluster of upsampled feature points. To embody the formulation, we further instantiate a lightweight upsampling operator, termed Similarity-Aware Point Affiliation (SAPA), and investigate its variants. SAPA invites consistent performance improvements on a number of dense prediction tasks, including semantic segmentation, object detection, depth estimation, and image matting. Code is available at: https://github.com/poppinace/sapa Hao Lu 0003, Wenze Liu, Zixuan Ye, Hongtao Fu, Zhiguo Cao 0001 |
NeurIPS | 1 |
| 2022 | Multispectral Semantic Land Cover Segmentation From Aerial Imagery With Deep Encoder-Decoder NetworkabstractDeveloping accurate algorithms for agricultural pattern recognition from aerial imagery has become increasingly important due to the prevalence of unmanned aerial vehicles (UAVs). This letter introduces a deep encoder–decoder network for semantic land cover segmentation, where the goal is to classify six anomaly categories from multispectral aerial imagery. Since aerial imagery exhibits specific characteristics and visual challenges in this imaging domain, existing semantic segmentation models are not plug-and-play. Starting from a state-of-the-art segmentation model, we present a step-by-step analysis of key challenges and also reveal our observations in addressing these challenges. In particular, we investigate on how to exploit data prior knowledge, how to deal with sample imbalance, and how to encode global semantic and contextual information to improve segmentation. Experiments on a recent large-scale aerial land cover data set demonstrate that our method achieves compelling performance against other state-of-the-art approaches. Our results and insights can provide references for practitioners working in this field when dealing with similar segmentation problems. Shuaiyuan Du, Hao Lu 0003, Dehui Li, Zhiguo Cao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Index NetworksabstractWe show that existing upsampling operators in convolutional networks can be unified using the notion of the index function. This notion is inspired by an observation in the decoding process of deep image matting where indices-guided unpooling can often recover boundary details considerably better than other upsampling operators such as bilinear interpolation. By viewing the indices as a function of the feature map, we introduce the concept of 'learning to index', and present a novel index-guided encoder-decoder framework where indices are learned adaptively from data and are used to guide downsampling and upsampling stages, without extra training supervision. At the core of this framework is a new learnable module, termed Index Network (IndexNet), which dynamically generates indices conditioned on the feature map. IndexNet can be used as a plug-in, applicable to almost all convolutional networks that have coupled downsampling and upsampling stages, enabling the networks to dynamically capture variations of local patterns. In particular, we instantiate and investigate five families of IndexNet. We highlight their superiority in delivering spatial information over other upsampling operators with experiments on synthetic data, and demonstrate their effectiveness on four dense prediction tasks, including image matting, image denoising, semantic segmentation, and monocular depth estimation. Code and models are available at https://git.io/IndexNet. Hao Lu 0003, Yutong Dai 0001, Chunhua Shen, Songcen Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | TasselNetV3: Explainable Plant Counting With Guided Upsampling and Background SuppressionabstractFast and accurate plant counting tools affect revolution in modern agriculture. Agricultural practitioners, however, expect the output of the tools to be not only accurate but also explainable. Such explainability often refers to the ability to infer which instance is counted. One intuitive way is to generate a bounding box for each instance. Nevertheless, compared with counting by detection, plant counts can be inferred more directly in the local count framework, while one thing reproaching this paradigm is its poor explainability of output visualization. In particular, we find that the poor explainability becomes a bottleneck limiting the counting performance. To address this, we explore the idea of guided upsampling and background suppression where a novel upsampling operator is proposed to allow count redistribution, and segmentation decoders with different fusion strategies are investigated to suppress background, respectively. By integrating them into our previous counting model TasselNetV2, we introduce TasselNetV3 series: TasselNetV3-Lite and TasselNetV3-Seg. We validate the TasselNetV3 series on three public plant counting data sets and a new unmanned aircraft vehicle (UAV)-based data set, covering maize tassels counting, wheat ears counting, and rice plants counting. Extensive results show that guided upsampling and background suppression not only improve counting performance but also enable explainable visualization. Aside from state-of-the-art performance, we have several interesting observations: 1) a limited-receptive-field counter in most cases outperforms a large-receptive-field one; 2) it is sufficient to generate empirical segmentation masks from dotted annotations; 3) middle fusion is a good choice to integrate foreground–backgrounda prioriknowledge; and 4) decoupling the learning of counting and segmentation matters. Hao Lu 0003, Liang Liu 0001, Yanan Li 0006, Xiao-Ming Zhao, Xi-Qing Wang, Zhiguo Cao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | NSSNet: Scale-Aware Object Counting With Non-Scale SuppressionabstractIn object counting, objects often exhibit different sizes at different scales, even if they have similar physical sizes in reality. This is particularly true when targeting crowd counting and vehicle counting in intelligent transportation. Failing to model such variations leads to the mismatch between the object size and image scale. To address this problem, existing methods often extract multi-scale features, but they either still generate the single-scale prediction or lack an explicit suppression mechanism to eliminate predictions engendered by inappropriate scales. Our scale analysis manifests that, the single-scale estimation only works well for objects of certain sizes, and a suppression operator is required to isolate the estimation of a specific scale. In this work, we propose a scale-aware counting network termed NSSNet. NSSNet has two key features: it not only i) generates multi-scale predictions but also ii) applies a novel non-scale suppression (NSS) operator to suppress scale-mismatched estimations. NSS is inspired by the widely-used non-maximum suppression (NMS). In contrast to NMS that only reserves the maximum response, NSS filters out those clearly wrong predictions (the remaining predictions may still be from multiple scales). We evaluate NSSNet on four standard crowd and vehicle counting benchmarks and report state-of-the-art performance. We also show the scale adaptability of NSSNet through a controlled multi-scale experiment. Code and pretrained models are available athttps://git.io/nssnet. Liang Liu 0001, Zhiguo Cao 0001, Hao Lu 0003, Haipeng Xiong, Chunhua Shen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Learning Affinity-Aware Upsampling for Deep Image MattingabstractWe show that learning affinity in upsampling provides an effective and efficient approach to exploit pairwise interactions in deep networks. Second-order features are commonly used in dense prediction to build adjacent relations with a learnable module after upsampling such as non-local blocks. Since upsampling is essential, learning affinity in upsampling can avoid additional propagation layers, offering the potential for building compact models. By looking at existing upsampling operators from a unified mathematical perspective, we generalize them into a second-order form and introduce Affinity-Aware Upsampling (A2U) where upsampling kernels are generated using a light-weight lowrank bilinear model and are conditioned on second-order features. Our upsampling operator can also be extended to downsampling. We discuss alternative implementations of A2U and verify their effectiveness on two detail-sensitive tasks: image reconstruction on a toy dataset; and a largescale image matting task where affinity-based ideas constitute mainstream matting approaches. In particular, results on the Composition-1k matting dataset show that A2U achieves a 14% relative improvement in the SAD metric against a strong baseline with negligible increase of parameters (< 0.5%). Compared with the state-of-the-art matting network, we achieve 8% higher performance with only 40% model complexity. Yutong Dai 0001, Hao Lu 0003, Chunhua Shen |
CVPR | 2 |
| 2021 | Composing Photos Like a PhotographerabstractWe show that explicit modeling of composition rules benefits image cropping. Image cropping is considered a promising way to automate aesthetic composition in professional photography. Existing efforts, however, only model such professional knowledge implicitly, e.g., by ranking from comparative candidates. Inspired by the observation that natural composition traits always follow a specific rule, we propose to learn such rules in a discriminative manner, and more importantly, to incorporate learned composition clues explicitly in the model. To this end, we introduce the concept of the key composition map (KCM) to encode the composition rules. The KCM can reveal the common laws hidden behind different composition rules and can inform the cropping model of what is important in composition. With the KCM, we present a novel cropping-by-composition paradigm and instantiate a network to implement composition-aware image cropping. Extensive experiments on two benchmarks justify that our approach enables effective, interpretable, and fast image cropping. Chaoyi Hong, Shuaiyuan Du, Ke Xian, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong |
CVPR | 4 |
| 2021 | TransView: Inside, Outside, and Across the Cropping View BoundariesabstractWe show that relation modeling between visual elements matters in cropping view recommendation. Cropping view recommendation addresses the problem of image recomposition conditioned on the composition quality and the ranking of views (cropped sub-regions). This task is challenging because the visual difference is subtle when a visual element is reserved or removed. Existing methods represent visual elements by extracting region-based convolutional features inside and outside the cropping view boundaries, without probing a fundamental question: why some visual elements are of interest or of discard? In this work, we observe that the relation between different visual elements significantly affects their relative positions to the desired cropping view, and such relation can be characterized by the attraction inside/outside the cropping view boundaries and the repulsion across the boundaries. By instantiating a transformer-based solution that represents visual elements as visual words and that models the dependencies between visual words, we report not only state-of-the-art performance on public benchmarks, but also interesting visualizations that depict the attraction and repulsion between visual elements, which may shed light on what makes for effective cropping view recommendation. Zhiguo Cao 0001, Kewei Wang 0001, Hao Lu 0003, Weicai Zhong |
ICCV | 4 |
| 2021 | Robust Image Cropping by Filtering Composition Irrelevant Factors
Ke Xian, Hao Lu 0003, Zhiguo Cao 0001 |
ICIG (3) | 3 |
| 2021 | Image Cropping Assisted By Modeling Inter-Patch RelationsabstractImage cropping is a common way to enhance the aesthetic quality of images. Huge industrial demand and the tediousness of image cropping make automatic image cropping a prosperous task. Existing works, however, face two difficulties: objects are easily truncated and key components of images are discarded by the model. The key to solving this problem is to understand the relations between different components of an image. These relations break the limit of spatial distance and reflect the contextual information in images, which help the model decide whether to retain a component. Motivated by this, a patch-related graph module is proposed to model the relations between different patches of an image. The patch-related features are extracted by a graph convolution layer and then fused with the original local features by a proposed gated unit. Moreover, a gradient layer is designed to embed the edge information in the input. The edge-prior input helps the model read the contents of images and reserve the main objects completely. Experimental results show that our model grasps the inter-patch relations well and performs competitively with other state-of-the-art approaches. Tianpei Lian, Zhiguo Cao 0001, Hao Lu 0003, Zijin Wu, Weicai Zhong |
ICIP | 3 |
| 2021 | Towards Light-Weight Portrait Matting via Parameter SharingabstractAbstract Traditional portrait matting methods typically consist of a trimap estimation network and a matting network. Here, we propose a new light‐weight portrait matting approach, termed parameter‐sharing portrait matting (PSPM). Different from conventional portrait matting models where the encoder and decoder networks in two tasks are often separately designed, here a single encoder is employed for the two tasks in PSPM, while each task still has its task‐specific decoder. Thus, the role of the encoder is to extract semantic features and two decoders function as a bridge between low‐resolution feature maps generated by the encoder and high‐resolution feature maps for pixel‐wise classification/regression. In particular, three variants capable of implementing the parameter‐sharing portrait matting network are proposed and investigated, respectively. As demonstrated in our experiments, model capacity and computation costs can be reduced significantly, by up to and , respectively, with PSPM, whereas the matting accuracy only slightly deteriorates. In addition, qualitative and quantitative evaluations show that sharing the encoder is an effective way to achieve portrait matting with limited computational budgets, indicating a promising direction for applications of real‐time portrait matting on mobile devices. Yutong Dai 0001, Hao Lu 0003, Chunhua Shen |
Comput. Graph. Forum | 2 |
| 2021 | Decoupled Two-Stage Crowd Counting and BeyondabstractOne of appealing approaches to counting dense objects, such as crowd, is density map estimation. Density maps, however, present ambiguous appearance cues in congested scenes, rendering infeasibility in identifying individuals and difficulties in diagnosing errors. Inspired by an observation that counting can be interpreted as a two-stage process, i.e., identifying possible object regions and counting exact object numbers, we introduce a probabilistic intermediate representation termed the probability map that depicts the probability of each pixel being an object. This representation allows us to decouple counting into probability map regression (PMR) and count map regression (CMR). We therefore propose a novel decoupled two-stage counting (D2C) framework that sequentially regresses the probability map and learns a counter conditioned on the probability map. Given the probability map and the count map, a peak point detection algorithm is derived to localize each object with a point under the guidance of local counts. An advantage of D2C is that the counter can be learned reliably with additional synthesized probability maps. This addresses important data deficiency and sample imbalanced problems in counting. Our framework also enables easy diagnoses and analyses of error patterns. For instance, we find that, the counter per se is sufficiently accurate, while the bottleneck appears to be PMR. We further instantiate a network D2CNet in our framework and report state-of-the-art counting and localization performance across 6 crowd counting benchmarks. Since the probability map is a representation independent of visual appearance, D2CNet also exhibits remarkable cross-dataset transferability. Code and pretrained models are made available at: https://git.io/d2cnet. Jian Cheng 0001, Haipeng Xiong, Zhiguo Cao 0001, Hao Lu 0003 |
IEEE Trans. Image Process. | 4 |
| 2021 | SESV: Accurate Medical Image Segmentation by Predicting and Correcting ErrorsabstractMedical image segmentation is an essential task in computer-aided diagnosis. Despite their prevalence and success, deep convolutional neural networks (DCNNs) still need to be improved to produce accurate and robust enough segmentation results for clinical use. In this paper, we propose a novel and generic framework called Segmentation-Emendation-reSegmentation-Verification (SESV) to improve the accuracy of existing DCNNs in medical image segmentation, instead of designing a more accurate segmentation model. Our idea is to predict the segmentation errors produced by an existing model and then correct them. Since predicting segmentation errors is challenging, we design two ways to tolerate the mistakes in the error prediction. First, rather than using a predicted segmentation error map to correct the segmentation mask directly, we only treat the error map as the prior that indicates the locations where segmentation errors are prone to occur, and then concatenate the error map with the image and segmentation mask as the input of a re-segmentation network. Second, we introduce a verification network to determine whether to accept or reject the refined mask produced by the re-segmentation network on a region-by-region basis. The experimental results on the CRAG, ISIC, and IDRiD datasets suggest that using our SESV framework can improve the accuracy of DeepLabv3+ substantially and achieve advanced performance in the segmentation of gland cells, skin lesions, and retinal microaneurysms. Consistent conclusions can also be drawn when using PSPNet, U-Net, and FPN as the segmentation network, respectively. Therefore, our SESV framework is capable of improving the accuracy of different DCNNs on different medical image segmentation tasks. Yutong Xie 0001, Hao Lu 0003, Chunhua Shen, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Weighing Counts: Sequential Crowd Counting by Reinforcement Learning
Liang Liu 0001, Hao Lu 0003, Hongwei Zou, Haipeng Xiong, Zhiguo Cao 0001, Chunhua Shen |
ECCV (10) | 2 |
| 2020 | Multi - Direction Convolution for Semantic Segmentation
Dehui Li, Zhiguo Cao 0001, Ke Xian, Xinyuan Qi, Hao Lu 0003 |
ICPR | 6 |
| 2020 | Counting Objects by Blockwise ClassificationabstractIn this paper, we introduce the idea of blockwise classification to count objects. The current mainstream method for counting objects is to regress the density map or to regress the redundant count map via a deep convolutional neural network (CNN). However, these methods suffer from two critical issues: inaccurately generated regression targets and serious sample imbalances. First, the ground truth density map is generated by convolving the dot map using a Gaussian kernel. Because an inappropriate kernel can cover the background or uncover objects, this approach introduces a form of noise, and therefore results in ambiguities when training the networks. Second, inhomogeneously distributed objects often exist in images, which gives rise to a data collection bias. This leads to a long-tailed distribution of region counts, which is a typical characteristic that occurs with imbalanced samples; therefore, underestimations in high-density regions and overestimations in low-density regions are common. In this paper, we address these two issues within one framework-blockwise count level classification. The intuition behind this idea is that while it may not be possible to provide an exact count of pixels or patches, it is possible to provide a count of a region that falls within a certain interval with high confidence. Our method classifies the count levels of each block produced by nonlinearly quantizing the continuous counts, thus transforming the imbalance of sample patch counts into a class imbalance of count levels. Consequently, an information-entropy-inspired loss can be applied to alleviate this issue. Through ablative studies, we analyze the impact of imbalanced data, Gaussian kernel sizes, quantization errors, and the effectiveness of each module in our method. Without bells and whistles, our method outperforms or performs competitively with other state-of-the-art approaches on seven object-counting benchmarks, including four crowd-counting datasets from ShanghaiTech, WorldExpo'10, UCF-QNRF and UCF_CC_50, one vehicle-counting dataset (TRANCOS), one maize-tassel-counting dataset (MTC), and one challenging sonar fish-counting dataset that we constructed. The results suggest that our framework provides a strong and improved baseline for object counting. Liang Liu 0001, Hao Lu 0003, Haipeng Xiong, Ke Xian, Zhiguo Cao 0001, Chunhua Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Indices Matter: Learning to Index for Deep Image MattingabstractWe show that existing upsampling operators can be unified using the notion of the index function. This notion is inspired by an observation in the decoding process of deep image matting where indices-guided unpooling can often recover boundary details considerably better than other upsampling operators such as bilinear interpolation. By viewing the indices as a function of the feature map, we introduce the concept of 'learning to index', and present a novel index-guided encoder-decoder framework where indices are self-learned adaptively from data and are used to guide the pooling and upsampling operators, without extra training supervision. At the core of this framework is a flexible network module, termed IndexNet, which dynamically generates indices conditioned on the feature map. Due to its flexibility, IndexNet can be used as a plug-in applying to almost all off-the-shelf convolutional networks that have coupled downsampling and upsampling stages. We demonstrate the effectiveness of IndexNet on the task of natural image matting where the quality of learned indices can be visually observed from predicted alpha mattes. Results on the Composition-1k matting dataset show that our model built on MobileNetv2 exhibits at least 16.1% improvement over the seminal VGG-16 based deep matting baseline, with less training data and lower model capacity. Code and models have been made available at: https://tinyurl.com/IndexNetV1. Hao Lu 0003, Yutong Dai 0001, Chunhua Shen, Songcen Xu |
ICCV | 1 |
| 2019 | From Open Set to Closed Set: Counting Objects by Spatial Divide-and-ConquerabstractVisual counting, a task that predicts the number of objects from an image/video, is an open-set problem by nature, i.e., the number of population can vary in [0,+∞) in theory. However, the collected images and labeled count values are limited in reality, which means only a small closed set is observed. Existing methods typically model this task in a regression manner, while they are likely to suffer from an unseen scene with counts out of the scope of the closed set. In fact, counting is decomposable. A dense region can always be divided until the count values of sub-regions are within the previously observed closed set. Inspired by this idea, we propose a simple but effective approach, Spatial Divide-and-Conquer Network (S-DCNet). S-DCNet learns to classify closed-set counts and can generalize to open-set counts via S-DC. S-DCNet is also efficient. To avoid repeatedly computing sub-region convolutional features, S-DC is executed on the feature map instead of on the input image. S-DCNet achieves the state-of-the-art performance on three crowd counting datasets (ShanghaiTech, UCF_CC_50 and UCF-QNRF), a vehicle counting dataset (TRANCOS) and a plant counting dataset (MTC). Compared to the previous best methods, S-DCNet brings a 20.2% relative improvement on the ShanghaiTechPart B, 20.9% on the UCF-QNRF, 22.5% on the TRANCOS and 15.1% on the MTC. Code has been made available at: https://github.com/xhp-hust-2018-2011/S-DCNet. Haipeng Xiong, Hao Lu 0003, Liang Liu 0001, Zhiguo Cao 0001, Chunhua Shen |
ICCV | 2 |
| 2019 | Deep Segmentation-Emendation Model for Gland Instance Segmentation
Yutong Xie 0001, Hao Lu 0003, Chunhua Shen, Yong Xia 0001 |
MICCAI (1) | 2 |
| 2019 | Unsupervised Domain Adaptation Using Robust Class-Wise MatchingabstractUnsupervised domain adaptation (DA) enables a classifier trained on data from one domain to be applied to data from another without labels. Given that the key to transferring a classifier across domains is to mitigate the data distribution mismatch for each class, most previous works completely or partially focus on global distribution matching across domains. The global data space, however, can be complicated, which makes modeling the global distribution difficult. To mitigate this problem, we present a novel unsupervised DA framework where the DA problem is addressed by proposing a robust class-wise matching strategy. Specifically, through minimizing a maximum mean discrepancy-based class-wise fisher discriminant across domains, this framework jointly optimizes two modules: a transferable feature learning module that reduces the distribution discrepancy between the same classes as well as increasing the distribution discrepancy between different classes across domains by a linear projection, and a robust classifier that exploits both the supervised information in source domain and the unsupervised low-rank property of target domain. In experiments on three DA benchmark data sets, the proposed framework shows the state-of-the-art performance. Lei Zhang 0054, Peng Wang 0023, Wei Wei 0008, Hao Lu 0003, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Deep Attention-Based Classification Network for Robust Depth Prediction
Ruibo Li, Ke Xian, Chunhua Shen, Zhiguo Cao 0001, Hao Lu 0003, Lingxiao Hang |
ACCV (4) | 5 |
| 2018 | Monocular Relative Depth Perception With Web Stereo Data SupervisionabstractIn this paper we study the problem of monocular relative depth perception in the wild. We introduce a simple yet effective method to automatically generate dense relative depth annotations from web stereo images, and propose a new dataset that consists of diverse images as well as corresponding dense relative depth maps. Further, an improved ranking loss is introduced to deal with imbalanced ordinal relations, enforcing the network to focus on a set of hard pairs. Experimental results demonstrate that our proposed approach not only achieves state-of-the-art accuracy of relative depth perception in the wild, but also benefits other dense per-pixel prediction tasks, e.g., metric depth estimation and semantic segmentation. Ke Xian, Chunhua Shen, Zhiguo Cao 0001, Hao Lu 0003, Yang Xiao 0007, Ruibo Li, Zhenbo Luo |
CVPR | 4 |
| 2018 | Counting Fish in Sonar ImagesabstractThe goal of this paper is to estimate the population of fishes in sonar images. Compared to natural images, sonar images present substantially different visual characteristics. Fishes in sonar images exhibit unreliable appearance cues, expose under imaging noise, vary significantly in shape and size. These pose great challenges for counting even for a human expert. In Computer Vision, a possible solution to this task is object counting with deep networks. This paradigm is typically formulated as a regression problem. The regression, however, greatly suffers from the issue of sample imbalance caused by fish variations in size and density, leading to underestimates in high-density regions and over-estimates in low-density regions. To address this, we build upon a recent local counts regression network and propose two novel losses to regularize a modified l1 loss with slack constraints. In particular, a challenging sonar fish counting dataset with 537 images and manually labeled dotted annotations is constructed. Experimental results on the dataset justify the effectiveness of our proposition and show improved performance of our method over other state-of-the-art approaches. Liang Liu 0001, Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007 |
ICIP | 2 |
| 2018 | RGB-D Co-Segmentation on Indoor Scene with Geometric Prior and Hypothesis Filtering
Lingxiao Hang, Zhiguo Cao 0001, Yang Xiao 0007, Hao Lu 0003 |
PRCV (1) | 4 |
| 2018 | Toward Good Practices for Fine-Grained Maize Cultivar Identification With Filter-Specific Convolutional ActivationsabstractCrop cultivar identification is an important aspect in agricultural systems. Traditional solutions involve excessive human interventions, which is labor-intensive and timeconsuming. In addition, cultivar identification is a typical task of fine-grained visual categorization (FGVC). Compared with other common topics in FGVC, studies of this problem are somewhat lagging and limited. In this paper, targeting four Chinese maize cultivars of Jundan No.20, Wuyue No.3, Nongda No.108, and Zhengdan No.958, we first consider the problem of identifying the maize cultivar based on its tassel characteristics by computer vision. In particular, a novel fine-grained maize cultivar identification data set termed HUST-FG-MCI that contains 5000 images is first constructed. To better capture the textual differences in a weakly supervised manner, we proposed an effective deep convolutional neural network and Fisher vector (FV)based feature encoding mechanism. The mechanism tends to highlight subtle object patterns via filter-specific convolutional representations and thus provides strong discrimination for cultivar identification. Experimental results demonstrate that our method outperforms other state-of-the-art approaches. We show also that FV encoding can weaken the linear dependency between convolutional activations, redundant filters exist in the convolutional layer, and high accuracy can be maintained with relatively low-dimensional convolutional features and one or two Gaussian components in FV. Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Yanjun Zhu |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2018 | An Embarrassingly Simple Approach to Visual Domain AdaptationabstractWe show that it is possible to achieve high-quality domain adaptation without explicit adaptation. The nature of the classification problem means that when samples from the same class in different domains are sufficiently close, and samples from differing classes are separated by large enough margins, there is a high probability that each will be classified correctly. Inspired by this, we propose an embarrassingly simple yet effective approach to domain adaptation-only the class mean is used to learn class-specific linear projections. Learning these projections is naturally cast into a linear-discriminant-analysis-like framework, which gives an efficient, closed form solution. Furthermore, to enable to application of this approach to unsupervised learning, an iterative validation strategy is developed to infer target labels. Extensive experiments on cross-domain visual recognition demonstrate that, even with the simplest formulation, our approach outperforms existing non-deep adaptation methods and exhibits classification performance comparable with that of modern deep adaptation methods. An analysis of potential issues effecting the practical application of the method is also described, including robustness, convergence, and the impact of small sample sizes. Hao Lu 0003, Chunhua Shen, Zhiguo Cao 0001, Yang Xiao 0007, Anton van den Hengel |
IEEE Trans. Image Process. | 1 |
| 2017 | When Unsupervised Domain Adaptation Meets Tensor RepresentationsabstractDomain adaption (DA) allows machine learning methods trained on data sampled from one distribution to be applied to data sampled from another. It is thus of great practical importance to the application of such methods. Despite the fact that tensor representations are widely used in Computer Vision to capture multi-linear relationships that affect the data, most existing DA methods are applicable to vectors only. This renders them incapable of reflecting and preserving important structure in many problems. We thus propose here a learning-based method to adapt the source and target tensor representations directly, without vectorization. In particular, a set of alignment matrices is introduced to align the tensor representations from both domains into the invariant tensor subspace. These alignment matrices and the tensor subspace are modeled as a joint optimization problem and can be learned adaptively from the data using the proposed alternative minimization scheme. Extensive experiments show that our approach is capable of preserving the discriminative power of the source domain, of resisting the effects of label noise, and works effectively for small sample sizes, and even one-shot DA. We show that our method outperforms the state-of-the-art on the task of cross-domain visual recognition in both efficacy and efficiency, and particularly that it outperforms all comparators when applied to DA of the convolutional activations of deep convolutional networks. Hao Lu 0003, Lei Zhang 0054, Zhiguo Cao 0001, Wei Wei 0008, Ke Xian, Chunhua Shen, Anton van den Hengel |
ICCV | 1 |
| 2017 | Two-dimensional subspace alignment for convolutional activations adaptation
Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007, Yanjun Zhu |
Pattern Recognit. | 1 |
| 2016 | Fine-grained maize cultivar identification using filter-specific convolutional activationsabstractCultivar identification is an important aspect in agriculture and also a typical task of fine-grained visual categorization (FGVC). In comparison with other common topics in FGVC, studies on this problem are somewhat lagged and limited. In this paper, targeting four Chinese maize cultivars of Jundan No.20, Wuyue No.3, Nongda No.108, and Zhengdan No.958, we first consider the problem of identifying the maize cultivar based on its tassel characteristics. Technically, an effective convolutional neural network (CNN) based feature encoding pipeline that allows integration of deep CNN based column feature extraction, filter-specific Fisher vector (FV) encoding and mutual information (MI) based filter selection is proposed to better address this problem. In particular, a novel fine-grained maize cultivar identification dataset termed MCI-4000 that contains 4000 images is first constructed by our team. Experimental results demonstrate that our method outperforms other stat-of-the-art approaches by at least 5% in accuracy. We also show that, there exists redundant filters in the last convolutional layer, and high accuracy can be achieved with only relatively low-dimensional column features and a small number of Gaussian components in FV. Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Yanjun Zhu |
ICIP | 1 |
| 2016 | Exploiting Attribute Dependency for Attribute Assignment in Crowded ScenesabstractAttributes now play a vital role for characterizing a crowded scene. Compared to low-level visual features, processing informed by attributes can capture rich semantic information. However, to effectively assign attributes to a crowded scene still remains a challenging task. In this letter, inspired by a recently proposed zero-shot learning framework, a novel attribute assignment method that maps low-level features to predefined attributes is proposed. In particular, we propose to exploit the attribute dependency during the phase of attribute assignment, which can be regarded as our main contribution. In addition, to further enhance the performance, an effective low-level feature extraction mechanism is also proposed. More precisely, appearance and motion features are first simultaneously extracted from several sampled video frames and corresponding optical flow fields via deep convolutional neural network and then, respectively, aggregated by using Fisher vector encoding to form the low-level representation of crowded scenes. Experimental results on the challenging WWW dataset demonstrate that both the proposed attribute assignment method and the low-level feature extraction mechanism outperform the state of the art. Chunhua Deng, Zhiguo Cao 0001, Yang Xiao 0007, Hao Lu 0003, Ke Xian |
IEEE Signal Process. Lett. | 4 |
| 2015 | Blurred image recognition using domain adaptationabstractImage blurring significantly degrades the image recognition performance. In this paper, we novelly address the blurred image recognition task from the perspective of domain adaptation (DA). The scenario is that, the training set (source domain) only comprises of the labelled clear images, and the test set (target domain) is composed of the unlabelled blurred images. DA is executed to eliminate the domain shift by subspace alignment. In this way, the clear and blurred image domains are pushed closer in the feature space. The supervised LMDR metric learning method is employed by us to construct the source domain subspace for further performance enhancement, compared to the unsupervised one (i.e., PCA). The experimental results on two datasets demonstrate that, the proposed DA-based blurred image recognition mechanism can significantly enhance the performance of different kinds of visual descriptors, especially when the blurring degree is strong. Xiaokang Xie, Zhiguo Cao 0001, Yang Xiao 0007, Mengyu Zhu, Hao Lu 0003 |
ICIP | 5 |