EDBT 2026 Demo / reviewers in the wild / expert
Xin Zhao 0012
dblp:68/2766-12
· DBLP profile ↗
47ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-7660-9897ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 2 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | B3CT: Three-Branch Learning with Unlabeled Target Signals for Domain-Robust Semantic Segmentation
Xin Zhao 0012, Jian Jia, Junyan Wang 0001, Lijun Cao |
Int. J. Comput. Vis. | 2 |
| 2025 | IDEA-Bench: How Far are Generative Models from Professional Designing?abstractRecent advancements in image generation models enable the creation of high-quality images and targeted modifications based on textual instructions. Some models even support multimodal complex guidance and demonstrate robust task generalization capabilities. However, they still fall short of meeting the nuanced, professional demands of designers. To bridge this gap, we introduce IDEA-Bench, a comprehensive benchmark designed to advance image generation models toward applications with robust task generalization. IDEA-Bench comprises 100 professional image generation tasks and 275 specific cases, categorized into five major types based on the current capabilities of existing models. Furthermore, we provide a representative subset of 18 tasks with enhanced evaluation criteria to facilitate more nuanced and reliable evaluations using Multimodal Large Language Models (MLLMs). By assessing models’ ability to comprehend and execute novel, complex tasks, IDEA-Bench paves the way toward the development of generative models with autonomous and versatile visual generation capabilities. Lianghua Huang, Jingwu Fang, Huanzhang Dou, Wei Wang 0354, Zhi-Fan Wu, Yupeng Shi, Junge Zhang, Xin Zhao 0012, Yu Liu 0063 |
CVPR | 9 |
| 2025 | Joint Feature and Kernel Fusion for Improved Depth-Aware Panoptic SegmentationabstractDepth-aware Panoptic Segmentation, which combines panoptic segmentation and monocular depth estimation, is a challenging task that requires a comprehensive understanding of both scene geometry and object semantics. Recent multi-task learning approaches have leveraged dynamic kernel methods to tackle these tasks simultaneously. However, these methods often treat feature extraction and kernel generation for each task in isolation, failing to fully exploit the rich interdependencies between depth and semantic information. To address this, we propose a novel framework with Cross-Task Feature Fusion and Kernel Fusion mechanisms to enhance Depth-aware Panoptic Segmentation. Our approach enables deeper integration of features and kernels, promoting more effective information exchange and mutual reinforcement between tasks. Experiments show that our method brings significant improvements, demonstrating the potential of a more integrated multi-task learning strategy for Depth-aware Panoptic Segmentation. Shu Tian, Xin Zhao 0012, Xu-Cheng Yin |
ICASSP | 3 |
| 2025 | Distribution Optimization Under Gaussian Hypothesis for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to transfer a model, proficient in dense image classification, from a source domain to a target domain. While various transfer methods have been explored in previous studies, we argue that the modeling of categories within the model significantly affects its transferability. Building on the Gaussian Hypothesis, which posits that each category in the feature space adheres to a multidimensional Gaussian distribution, we propose a Class-Aware Variational Inference (CAVI) training method. This approach normalizes features of different categories into distinct multidimensional Gaussian distributions. To further learn domain-independent feature distributions, we optimize the feature space using a Gaussian-based alignment strategy and incorporate Gaussian-based contrastive learning. Experimental results demonstrate that our method achieves state-of-the-art performance on the GTAV → Cityscapes and Synthia → Cityscapes benchmarks. Xin Zhao 0012, Junyan Wang 0001, Lijun Cao, Junge Zhang |
WACV | 3 |
| 2024 | DDAE: Towards Deep Dynamic Vision BERT PretrainingabstractRecently, masked image modeling (MIM) has demonstrated promising prospects in self-supervised representation learning. However, existing MIM frameworks recover all masked patches equivalently, ignoring that the reconstruction difficulty of different patches can vary sharply due to their diverse distance from visible patches. In this paper, we propose a novel deep dynamic supervision to enable MIM methods to dynamically reconstruct patches with different degrees of difficulty at different pretraining phases and depths of the model. Our deep dynamic supervision helps to provide more locality inductive bias for ViTs especially in deep layers, which inherently makes up for the absence of local prior for self-attention mechanism. Built upon the deep dynamic supervision, we propose Deep Dynamic AutoEncoder (DDAE), a simple yet effective MIM framework that utilizes dynamic mechanisms for pixel regression and feature self-distillation simultaneously. Extensive experiments across a variety of vision tasks including ImageNet classification, semantic segmentation on ADE20K and object detection on COCO demonstrate the effectiveness of our approach. Xiangwen Kong, Xiangyu Zhang 0005, Xin Zhao 0012, Kaiqi Huang |
AAAI | 4 |
| 2024 | PeLK: Parameter-Efficient Large Kernel ConvNets with Peripheral ConvolutionabstractRecently, some large kernel convnets strike back with appealing performance and efficiency. However, given the square complexity of convolution, scaling up kernels can bring about an enormous amount of parameters and the proliferated parameters can induce severe optimization problem. Due to these issues, current CNNs compromise to scale up to 51 × 51 in the form of stripe convolution (i.e., 51 ×5 + 5 ×51) and start to saturate as the kernel size continues growing. In this paper, we delve into addressing these vital issues and explore whether we can continue scaling up kernels for more performance gains. Inspired by human vision, we propose a human-like peripheral convolution that efficiently reduces over 90% parameter count of dense grid convolution through parameter sharing, and manage to scale up kernel size to extremely large. Our peripheral convolution behaves highly similar to human, reducing the complexity of convolution from O(K2) to O(logK) without backfiring performance. Built on this, we propose Parameter-efficient Large Kernel Network (PeLK). Our PeLK outperforms modern vision Transformers and ConvNet architectures like Swin, ConvNeXt, RepLKNet and SLaK on various vision tasks including ImageNet classification, semantic segmentation on ADE20K and object detection on MS COCO. For the first time, we successfully scale up the kernel size of CNNs to an unprecedented 101 × 101 and demonstrate consistent improvements. Xiangxiang Chu, Xin Zhao 0012, Kaiqi Huang |
CVPR | 4 |
| 2024 | Robust Single-Particle Cryo-Em Image Denoising and RestorationabstractCryo-electron microscopy (cryo-EM) has achieved nearatomic level resolution of biomolecules by reconstructing 2D micrographs. However, the resolution and accuracy of the reconstructed particles are significantly reduced due to the extremely low signal-to-noise ratio (SNR) and complex noise structure of cryo-EM images. In this paper, we introduce a diffusion model with post-processing framework to effectively denoise and restore single particle cryo-EM images. Our method outperforms the state-of-the-art (SOTA) denoising methods by effectively removing structural noise that has not been addressed before. Additionally, more accurate and high-resolution three-dimensional reconstruction structures can be obtained from denoised cryo-EM images. Jing Zhang 0110, Tengfei Zhao, Xin Zhao 0012 |
ICASSP | 4 |
| 2024 | SOTVerse: A User-Defined Task Space of Single Object Tracking
Xin Zhao 0012, Kaiqi Huang |
Int. J. Comput. Vis. | 2 |
| 2024 | Correction: SOTVerse: A User-Defined Task Space of Single Object Tracking
Xin Zhao 0012, Kaiqi Huang |
Int. J. Comput. Vis. | 2 |
| 2024 | BioDrone: A Bionic Drone-Based Single Object Tracking Benchmark for Robust Vision
Xin Zhao 0012, Jing Zhang 0110, Yimin Hu, Rongshuai Liu, Haibin Ling, Yin Li 0003, Renshu Li, Jiadong Li |
Int. J. Comput. Vis. | 1 |
| 2024 | Global and local multi-modal feature mutual learning for retinal vessel segmentation
Xin Zhao 0012, Jing Zhang 0110, Qiaozhe Li, Tengfei Zhao, Zifeng Wu |
Pattern Recognit. | 1 |
| 2023 | Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision TransformersabstractVector-quantized image modeling has shown great potential in synthesizing high-quality images. However, generating high-resolution images remains a challenging task due to the quadratic computational overhead of the self-attention process. In this study, we seek to explore a more efficient two-stage framework for high-resolution image generation with improvements in the following three aspects. (1) Based on the observation that the first quantization stage has solid local property, we employ a local attention-based quantization model instead of the global attention mechanism used in previous methods, leading to better efficiency and reconstruction quality. (2) We emphasize the importance of multi-grained feature interaction during image generation and introduce an efficient attention mechanism that combines global attention (long-range semantic consistency within the whole image) and local attention (fined-grained details). This approach results in faster generation speed, higher generation fidelity, and improved resolution. (3) We propose a new generation pipeline incorporating autoencoding training and autoregressive generation strategy, demonstrating a better paradigm for image synthesis. Extensive experiments demonstrate the superiority of our approach in high-quality and high-resolution image reconstruction and generation. Shiyue Cao, Yueqin Yin, Lianghua Huang, Yu Liu 0063, Xin Zhao 0012, Deli Zhao, Kaiqi Huang |
ICCV | 5 |
| 2023 | A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and Causal RelationshipabstractTracking an arbitrary moving target in a video sequence is the foundation for high-level tasks like video understanding. Although existing visual-based trackers have demonstrated good tracking capabilities in short video sequences, they always perform poorly in complex environments, as represented by the recently proposed global instance tracking task, which consists of longer videos with more complicated narrative content. Recently, several works have introduced natural language into object tracking, desiring to address the limitations of relying only on a single visual modality. However, these selected videos are still short sequences with uncomplicated spatio-temporal and causal relationships, and the provided semantic descriptions are too simple to characterize video content.To address these issues, we (1) first propose a new multi-modal global instance tracking benchmark named MGIT. It consists of 150 long video sequences with a total of 2.03 million frames, aiming to fully represent the complex spatio-temporal and causal relationships coupled in longer narrative content. (2) Each video sequence is annotated with three semantic grains (i.e., action, activity, and story) to model the progressive process of human cognition. We expect this multi-granular annotation strategy can provide a favorable environment for multi-modal object tracking research and long video understanding. (3) Besides, we execute comparative experiments on existing multi-modal object tracking benchmarks, which not only explore the impact of different annotation methods, but also validate that our annotation method is a feasible solution for coupling human understanding into semantic labels. (4) Additionally, we conduct detailed experimental analyses on MGIT, and hope the explored performance bottlenecks of existing algorithms can support further research in multi-modal object tracking. The proposed benchmark, experimental results, and toolkit will be released gradually on http://videocube.aitestunion.com/. Dailing Zhang, Meiqi Wu, Xiaokun Feng, Xuchen Li 0001, Xin Zhao 0012, Kaiqi Huang |
NeurIPS | 6 |
| 2023 | Global Instance Tracking: Locating Target More Like HumansabstractTarget tracking, the essential ability of the human visual system, has been simulated by computer vision tasks. However, existing trackers perform well in austere experimental environments but fail in challenges like occlusion and fast motion. The massive gap indicates that researches only measure tracking performance rather than intelligence. How to scientifically judge the intelligence level of trackers? Distinct from decision-making problems, lacking three requirements (a challenging task, a fair environment, and a scientific evaluation procedure) makes it strenuous to answer the question. In this article, we first propose the global instance tracking (GIT) task, which is supposed to search an arbitrary user-specified instance in a video without any assumptions about camera or motion consistency, to model the human visual tracking ability. Whereafter, we construct a high-quality and large-scale benchmark VideoCube to create a challenging environment. Finally, we design a scientific evaluation procedure using human capabilities as the baseline to judge tracking intelligence. Additionally, we provide an online platform with toolkit and an updated leaderboard. Although the experimental results indicate a definite gap between trackers and humans, we expect to take a step forward to generate authentic human-like trackers. The database, toolkit, evaluation server, and baseline results are available at http://videocube.aitestunion.com. Xin Zhao 0012, Lianghua Huang, Kaiqi Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | QueryProp: Object Query Propagation for High-Performance Video Object DetectionabstractVideo object detection has been an important yet challenging topic in computer vision. Traditional methods mainly focus on designing the image-level or box-level feature propagation strategies to exploit temporal information. This paper argues that with a more effective and efficient feature propagation framework, video object detectors can gain improvement in terms of both accuracy and speed. For this purpose, this paper studies object-level feature propagation, and proposes an object query propagation (QueryProp) framework for high-performance video object detection. The proposed QueryProp contains two propagation strategies: 1) query propagation is performed from sparse key frames to dense non-key frames to reduce the redundant computation on non-key frames; 2) query propagation is performed from previous key frames to the current key frame to improve feature representation by temporal context modeling. To further facilitate query propagation, an adaptive propagation gate is designed to achieve flexible key frame selection. We conduct extensive experiments on the ImageNet VID dataset. QueryProp achieves comparable accuracy with state-of-the-art methods and strikes a decent accuracy/speed trade-off. Naiyu Gao, Jian Jia, Xin Zhao 0012, Kaiqi Huang |
AAAI | 4 |
| 2022 | PanopticDepth: A Unified Framework for Depth-aware Panoptic SegmentationabstractThis paper presents a unified framework for depth-aware panoptic segmentation (DPS), which aims to reconstruct 3D scene with instance-level semantics from one single image. Prior works address this problem by simply adding a dense depth regression head to panoptic segmentation (PS) networks, resulting in two independent task branches. This neglects the mutually-beneficial relations between these two tasks, thus failing to exploit handy instance-level semantic cues to boost depth accuracy while also producing sub-optimal depth maps. To overcome these limitations, we propose a unified framework for the DPS task by applying a dynamic convolution technique to both the PS and depth prediction tasks. Specifically, instead of predicting depth for all pixels at a time, we generate instance-specific kernels to predict depth and segmentation masks for each instance. Moreover, leveraging the instance-wise depth estimation scheme, we add additional instance-level depth cues to assist with supervising the depth learning via a new depth loss. Extensive experiments on Cityscapes-DPS and SemKITTI-DPS show the effectiveness and promise of our method. We hope our unified solution to DPS can lead a new paradigm in this area. Code is available at https://github.com/NaiyuGao/PanopticDepth. Naiyu Gao, Jian Jia, Yanhu Shan, Xin Zhao 0012, Kaiqi Huang |
CVPR | 6 |
| 2022 | InsPro: Propagating Instance Query and Proposal for Online Video Instance SegmentationabstractVideo instance segmentation (VIS) aims at segmenting and tracking objects in videos. Prior methods typically generate frame-level or clip-level object instances first and then associate them by either additional tracking heads or complex instance matching algorithms. This explicit instance association approach increases system complexity and fails to fully exploit temporal cues in videos. In this paper, we design a simple, fast and yet effective query-based framework for online VIS. Relying on an instance query and proposal propagation mechanism with several specially developed components, this framework can perform accurate instance association implicitly. Specifically, we generate frame-level object instances based on a set of instance query-proposal pairs propagated from previous frames. This instance query-proposal pair is learned to bind with one specific object across frames through conscientiously developed strategies. When using such a pair to predict an object instance on the current frame, not only the generated instance is automatically associated with its precursors on previous frames, but the model gets a good prior for predicting the same object. In this way, we naturally achieve implicit instance association in parallel with segmentation and elegantly take advantage of temporal clues in videos. To show the effectiveness of our method InsPro, we evaluate it on two popular VIS benchmarks, i.e., YouTube-VIS 2019 and YouTube-VIS 2021. Without bells-and-whistles, our InsPro with ResNet-50 backbone achieves 43.2 AP and 37.6 AP on these two benchmarks respectively, outperforming all other online VIS methods. Naiyu Gao, Jian Jia, Yanhu Shan, Xin Zhao 0012, Kaiqi Huang |
NeurIPS | 6 |
| 2022 | Revisiting instance search: A new benchmark using cycle self-training
Yuqi Zhang 0001, Chong Liu 0002, Fan Wang 0019, Hao Li 0030, Xin Zhao 0012 |
Neurocomputing | 8 |
| 2022 | Temporal-adaptive sparse feature aggregation for video object detection
Qiaozhe Li, Xin Zhao 0012, Kaiqi Huang |
Pattern Recognit. | 3 |
| 2021 | Can DNN Detectors Compete Against Human Vision in Object Detection Task?
Qiaozhe Li, Xin Zhao 0012, Kaiqi Huang |
PRCV (1) | 3 |
| 2021 | GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the WildabstractWe introduce here a large tracking database that offers an unprecedentedly wide coverage of common moving objects in the wild, called GOT-10k. Specifically, GOT-10k is built upon the backbone of WordNet structure [1] and it populates the majority of over 560 classes of moving objects and 87 motion patterns, magnitudes wider than the most recent similar-scale counterparts [19], [20], [23], [26]. By releasing the large high-diversity database, we aim to provide a unified training and evaluation platform for the development of class-agnostic, generic purposed short-term trackers. The features of GOT-10k and the contributions of this article are summarized in the following. (1) GOT-10k offers over 10,000 video segments with more than 1.5 million manually labeled bounding boxes, enabling unified training and stable evaluation of deep trackers. (2) GOT-10k is by far the first video trajectory dataset that uses the semantic hierarchy of WordNet to guide class population, which ensures a comprehensive and relatively unbiased coverage of diverse moving objects. (3) For the first time, GOT-10k introduces the one-shot protocol for tracker evaluation, where the training and test classes are zero-overlapped. The protocol avoids biased evaluation results towards familiar objects and it promotes generalization in tracker development. (4) GOT-10k offers additional labels such as motion classes and object visible ratios, facilitating the development of motion-aware and occlusion-aware trackers. (5) We conduct extensive tracking experiments with 39 typical tracking algorithms and their variants on GOT-10k and analyze their results in this paper. (6) Finally, we develop a comprehensive platform for the tracking community that offers full-featured evaluation toolkits, an online evaluation server, and a responsive leaderboard. The annotations of GOT-10k's test data are kept private to avoid tuning parameters on it. Lianghua Huang, Xin Zhao 0012, Kaiqi Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | SSAP: Single-Shot Instance Segmentation With Affinity PyramidabstractProposal-free instance segmentation methods mainly generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, previous methods mostly employ separate modules for these two sub-tasks and require multiple passes for inference. In addition to the lack of efficiency, previous methods also failed to perform as well as proposal-based approaches. To this end, this work proposes a single-shot proposal-free instance segmentation method that requires only one single pass for prediction. Our method is based on learning an affinity pyramid, which computes the probability that two pixels belong to the same instance in a hierarchical manner. Moreover, incorporating with the learned affinity pyramid, a novel cascaded graph partition (CGP) module is presented to fuse the two predictions and segment instances efficiently. As an additional contribution, we conduct an experiment to demonstrate the benefits of proposal-free methods in capturing detailed structures from finely annotated training examples. Our approach is evaluated on the Cityscapes and COCO datasets and achieves state-of-the-art performance. Naiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 0012, Kaiqi Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Learning Category- and Instance-Aware Pixel Embedding for Fast Panoptic SegmentationabstractPanoptic segmentation (PS) is a complex scene understanding task that requires providing high-quality segmentation for both thing objects and stuff regions. Previous methods handle these two classes with semantic and instance segmentation modules separately, following with heuristic fusion or additional modules to resolve the conflicts between the two outputs. This work simplifies this pipeline of PS by consistently modeling the two classes with a novel PS framework, which extends a detection model with an extra module to predict category- and instance-aware pixel embedding (CIAE). CIAE is a novel pixel-wise embedding feature that encodes both semantic-classification and instance-distinction information. At the inference process, PS results are simply derived by assigning each pixel to a detected instance or a stuff class according to the learned embedding. Our method not only demonstrates fast inference speed but also the first one-stage method to achieve comparable performance to two-stage methods on the challenging COCO benchmark. Naiyu Gao, Yanhu Shan, Xin Zhao 0012, Kaiqi Huang |
IEEE Trans. Image Process. | 3 |
| 2020 | Temporal Context Enhanced Feature Aggregation for Video Object DetectionabstractVideo object detection is a challenging task because of the presence of appearance deterioration in certain video frames. One typical solution is to aggregate neighboring features to enhance per-frame appearance features. However, such a method ignores the temporal relations between the aggregated frames, which is critical for improving video recognition accuracy. To handle the appearance deterioration problem, this paper proposes a temporal context enhanced network (TCENet) to exploit temporal context information by temporal aggregation for video object detection. To handle the displacement of the objects in videos, a novel DeformAlign module is proposed to align the spatial features from frame to frame. Instead of adopting a fixed-length window fusion strategy, a temporal stride predictor is proposed to adaptively select video frames for aggregation, which facilitates exploiting variable temporal information and requiring fewer video frames for aggregation to achieve better results. Our TCENet achieves state-of-the-art performance on the ImageNet VID dataset and has a faster runtime. Without bells-and-whistles, our TCENet achieves 80.3% mAP by only aggregating 3 frames. Naiyu Gao, Qiaozhe Li, Senyao Du, Xin Zhao 0012, Kaiqi Huang |
AAAI | 5 |
| 2020 | GlobalTrack: A Simple and Strong Baseline for Long-Term TrackingabstractA key capability of a long-term tracker is to search for targets in very large areas (typically the entire image) to handle possible target absences or tracking failures. However, currently there is a lack of such a strong baseline for global instance search. In this work, we aim to bridge this gap. Specifically, we propose GlobalTrack, a pure global instance search based tracker that makes no assumption on the temporal consistency of the target's positions and scales. GlobalTrack is developed based on two-stage object detectors, and it is able to perform full-image and multi-scale search of arbitrary instances with only a single query as the guide. We further propose a cross-query loss to improve the robustness of our approach against distractors. With no online learning, no punishment on position or scale changes, no scale smoothing and no trajectory refinement, our pure global instance search based tracker achieves comparable, sometimes much better performance on four large-scale tracking benchmarks (i.e., 52.1% AUC on LaSOT, 63.8% success rate on TLP, 60.3% MaxGM on OxUvA and 75.4% normalized precision on TrackingNet), compared to state-of-the-art approaches that typically require complex post-processing. More importantly, our tracker runs without cumulative errors, i.e., any type of temporary tracking failures will not affect its performance on future frames, making it ideal for long-term tracking. We hope this work will be a strong baseline for long-term tracking and will stimulate future works in this area. Lianghua Huang, Xin Zhao 0012, Kaiqi Huang |
AAAI | 2 |
| 2020 | TANet: Robust 3D Object Detection from Point Clouds with Triple AttentionabstractIn this paper, we focus on exploring the robustness of the 3D object detection in point clouds, which has been rarely discussed in existing approaches. We observe two crucial phenomena: 1) the detection accuracy of the hard objects, e.g., Pedestrians, is unsatisfactory, 2) when adding additional noise points, the performance of existing approaches decreases rapidly. To alleviate these problems, a novel TANet is introduced in this paper, which mainly contains a Triple Attention (TA) module, and a Coarse-to-Fine Regression (CFR) module. By considering the channel-wise, point-wise and voxel-wise attention jointly, the TA module enhances the crucial information of the target while suppresses the unstable cloud points. Besides, the novel stacked TA further exploits the multi-level feature attention. In addition, the CFR module boosts the accuracy of localization without excessive computation cost. Experimental results on the validation set of KITTI dataset demonstrate that, in the challenging noisy cases, i.e., adding additional random noisy points around each object, the presented approach goes far beyond state-of-the-art approaches. Furthermore, for the 3D object detection task of the KITTI benchmark, our approach ranks the first place on Pedestrian class, by using the point clouds as the only input. The running speed is around 29 frames per second. Zhe Liu 0033, Xin Zhao 0012, Tengteng Huang, Ruolan Hu, Yu Zhou 0016, Xiang Bai |
AAAI | 2 |
| 2020 | Recurrent Prediction With Spatio-Temporal Attention for Crowd Attribute RecognitionabstractCrowd attribute recognition is a challenging task for crowd video understanding because a crowd video often contains multiple attributes from various types. Traditional deep learning-based methods directly treat this recognition problem as a multiple binary classification problem and represent the video by vectorizing and fusing the separately learned spatial and temporal features in the fully connected layers. Therefore, the correlations between these attributes may not be well captured. In this paper, a bidirectional recurrent prediction model with a semantic-aware attention mechanism is proposed to explore the spatio-temporal and semantic relations between the attributes for more accurate recognition. The ConvLSTM is introduced for feature representation to capture the spatio-temporal structure of the crowd videos and facilitate the visual attention. The bidirectional recurrent attention module is proposed for sequential attribute prediction by associating each subcategory attributes to corresponding semantic-related regions iteratively. The experiments and evaluations on the challenging WWW crowd video dataset not only show that our approach significantly outperforms the state-of-the-art methods but also verify that our approach can effectively capture the spatio-temporal and semantic relations of the crowd attributes. Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Visual-Semantic Graph Reasoning for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition in surveillance is a challenging task due to poor image quality, significant appearance variations and diverse spatial distribution of different attributes. This paper treats pedestrian attribute recognition as a sequential attribute prediction problem and proposes a novel visual-semantic graph reasoning framework to address this problem. Our framework contains a spatial graph and a directed semantic graph. By performing reasoning using the Graph Convolutional Network (GCN), one graph captures spatial relations between regions and the other learns potential semantic relations between attributes. An end-to-end architecture is presented to perform mutual embedding between these two graphs to guide the relational learning for each other. We verify the proposed framework on three large scale pedestrian attribute datasets including PETA, RAP, and PA100k. Experiments show superiority of the proposed method over state-of-the-art methods and effectiveness of our joint GCN structures for sequential attribute prediction. Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang |
AAAI | 2 |
| 2019 | 3D Object Detection Using Scale Invariant and Feature Reweighting Networksabstract3D object detection plays an important role in a large number of real-world applications. It requires us to estimate the localizations and the orientations of 3D objects in real scenes. In this paper, we present a new network architecture which focuses on utilizing the front view images and frustum point clouds to generate 3D detection results. On the one hand, a PointSIFT module is utilized to improve the performance of 3D segmentation. It can capture the information from different orientations in space and the robustness to different scale shapes. On the other hand, our network obtains the useful features and suppresses the features with less information by a SENet module. This module reweights channel features and estimates the 3D bounding boxes more effectively. Our method is evaluated on both KITTI dataset for outdoor scenes and SUN-RGBD dataset for indoor scenes. The experimental results illustrate that our method achieves better performance than the state-of-the-art methods especially when point clouds are highly sparse. Xin Zhao 0012, Zhe Liu 0033, Ruolan Hu, Kaiqi Huang |
AAAI | 1 |
| 2019 | SSAP: Single-Shot Instance Segmentation With Affinity PyramidabstractRecently, proposal-free instance segmentation has received increasing attention due to its concise and efficient pipeline. Generally, proposal-free methods generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, previous methods mostly employ separate modules for these two sub-tasks and require multiple passes for inference. We argue that treating these two sub-tasks separately is suboptimal. In fact, employing multiple separate modules significantly reduces the potential for application. The mutual benefits between the two complementary sub-tasks are also unexplored. To this end, this work proposes a single-shot proposal-free instance segmentation method that requires only one single pass for prediction. Our method is based on a pixel-pair affinity pyramid, which computes the probability that two pixels belong to the same instance in a hierarchical manner. The affinity pyramid can also be jointly learned with the semantic class labeling and achieve mutual benefits. Moreover, incorporating with the learned affinity pyramid, a novel cascaded graph partition module is presented to sequentially generate instances from coarse to fine. Unlike previous time-consuming graph partition methods, this module achieves 5× speedup and 9% relative improvement on Average-Precision (AP). Our approach achieves new state of the art on the challenging Cityscapes dataset. Naiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 0012, Yinan Yu, Ming Yang 0007, Kaiqi Huang |
ICCV | 4 |
| 2019 | Bridging the Gap Between Detection and Tracking: A Unified ApproachabstractObject detection models have been a source of inspiration for many tracking-by-detection algorithms over the past decade. Recent deep trackers borrow designs or modules from the latest object detection methods, such as bounding box regression, RPN and ROI pooling, and can deliver impressive performance. In this paper, instead of redesigning a new tracking-by-detection algorithm, we aim to explore a general framework for building trackers directly upon almost any advanced object detector. To achieve this, three key gaps must be bridged: (1) Object detectors are class-specific, while trackers are class-agnostic. (2) Object detectors do not differentiate intra-class instances, while this is a critical capability of a tracker. (3) Temporal cues are important for stable long-term tracking while they are not considered in still-image detectors. To address the above issues, we first present a simple target-guidance module for guiding the detector to locate target-relevant objects. Then a meta-learner is adopted for the detector to fast learn and adapt a target-distractor classifier online. We further introduce an anchored updating strategy to alleviate the problem of overfitting. The framework is instantiated on SSD and FasterRCNN, the typical one- and two-stage detectors, respectively. Experiments on OTB, UAV123 and NfS have verified our framework and show that our trackers can benefit from deeper backbone networks, as opposed to many recent trackers. Lianghua Huang, Xin Zhao 0012, Kaiqi Huang |
ICCV | 2 |
| 2019 | Pedestrian Attribute Recognition by Joint Visual-semantic Reasoning and Knowledge DistillationabstractPedestrian attribute recognition in surveillance is a challenging task in computer vision due to significant pose variation, viewpoint change and poor image quality. To achieve effective recognition, this paper presents a graph-based global reasoning framework to jointly model potential visual-semantic relations of attributes and distill auxiliary human parsing knowledge to guide the relational learning. The reasoning framework models attribute groups on a graph and learns a projection function to adaptively assign local visual features to the nodes of the graph. After feature projection, graph convolution is utilized to perform global reasoning between the attribute groups to model their mutual dependencies. Then, the learned node features are projected back to visual space to facilitate knowledge transfer. An additional regularization term is proposed by distilling human parsing knowledge from a pre-trained teacher model to enhance feature representations. The proposed framework is verified on three large scale pedestrian attribute datasets including PETA, RAP, and PA-100k. Experiments show that our method achieves state-of-the-art results. Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang |
IJCAI | 2 |
| 2019 | Focal Boundary Guided Salient Object DetectionabstractThe performance of salient object segmentation has been significantly advanced by using deep convolutional networks. However, these networks often produce blob-like saliency maps without accurate object boundaries. This is caused by the limited spatial resolution of their feature maps after multiple pooling operations, and might hinder downstream applications that require precise object shapes. To address this issue, we propose a novel deep model-Focal Boundary Guided (Focal- BG) network. Our model is designed to jointly learn to segment salient object masks and detect salient object boundaries. Our key idea is that additional knowledge about object boundaries can help to precisely identify the shape of the object. Moreover, our model incorporates a refinement pathway to refine the mask prediction, and makes use of the focal loss to facilitate the learning of the hard boundary pixels. To evaluate our model, we conduct extensive experiments. Our Focal-BG network consistently outperforms state-of-the-art methods on five major benchmarks. We provide a detailed analysis of these results and demonstrate that our joint modeling of salient object boundary and mask helps to better capture shape details, especially in the vicinity of object boundaries. Yupei Wang, Xin Zhao 0012, Xuecai Hu, Yin Li 0003, Kaiqi Huang |
IEEE Trans. Image Process. | 2 |
| 2019 | Deep Crisp Boundaries: From Boundaries to Higher-Level TasksabstractEdge detection has made significant progress with the help of deep convolutional networks (ConvNet). These ConvNet-based edge detectors have approached human level performance on standard benchmarks. We provide a systematical study of these detectors' outputs. We show that the detection results did not accurately localize edge pixels, which can be adversarial for tasks that require crisp edge inputs. As a remedy, we propose a novel refinement architecture to address the challenging problem of learning a crisp edge detector using ConvNet. Our method leverages a top-down backward refinement pathway, and progressively increases the resolution of feature maps to generate crisp edges. Our results achieve superior performance, surpassing human accuracy when using standard criteria on BSDS500, and largely outperforming the state-of-the-art methods when using more strict criteria. More importantly, we demonstrate the benefit of crisp edge maps for several important applications in computer vision, including optical flow estimation, object proposal generation, and semantic segmentation. Yupei Wang, Xin Zhao 0012, Yin Li 0003, Kaiqi Huang |
IEEE Trans. Image Process. | 2 |
| 2018 | Densely Connected Single-Shot DetectorabstractOne-stage object detection approach which utilizes multi-scale feature maps to predict objects is currently the best real-time detector. However, in this approach, the high-resolution feature maps which are responsible for detecting small objects are harder to learn a proper abstraction of objects than the low-resolution feature maps. The problem is that these feature maps have to transform sufficient low-level information to the next layer while learning high-level abstraction. In this paper, we develop a transformation module which adopts the dense structure to simplify the learning problem of high-resolution feature maps. In addition, we utilize the inception module to enrich the representation power of high-resolution feature maps. Extensive experiments on most object detection datasets clearly demonstrate the effectiveness of our method. In particular, on PASCAL VOC 2007/2012, our method outperforms all the existing one-stage methods. Our model based on the VGG-16 network also achieves competitive result on MS COCO. Pei Xu 0003, Xin Zhao 0012, Kaiqi Huang |
ICPR | 2 |
| 2018 | Densely Cascaded Shadow Detection Network via Deeply Supervised Parallel FusionabstractShadow detection is an important and challenging problem in computer vision. Recently, single image shadow detection had achieved major progress with the development of deep convolutional networks. However, existing methods are still vulnerable to background clutters, and often fail to capture the global context of an input image. These global contextual and semantic cues are essential for accurately localizing the shadow regions. Moreover, rich spatial details are required to segment shadow regions with precise shape. To this end, this paper presents a novel model characterized by a deeply supervised parallel fusion (DSPF) network and a densely cascaded learning scheme. The DSPF network achieves a comprehensive fusion of global semantic cues and local spatial details by multiple stacked parallel fusion branches, which are learned in a deeply supervised manner. Moreover, the densely cascaded learning scheme is employed to refine the spatial details. Our method is evaluated on two widely used shadow detection benchmarks. Experimental results show that our method outperforms state-of-the-arts by a large margin. Yupei Wang, Xin Zhao 0012, Yin Li 0003, Xuecai Hu, Kaiqi Huang |
IJCAI | 2 |
| 2017 | Locality-Sensitive Deconvolution Networks with Gated Fusion for RGB-D Indoor Semantic SegmentationabstractThis paper focuses on indoor semantic segmentation using RGB-D data. Although the commonly used deconvolution networks (DeconvNet) have achieved impressive results on this task, we find there is still room for improvements in two aspects. One is about the boundary segmentation. DeconvNet aggregates large context to predict the label of each pixel, inherently limiting the segmentation precision of object boundaries. The other is about RGB-D fusion. Recent state-of-the-art methods generally fuse RGB and depth networks with equal-weight score fusion, regardless of the varying contributions of the two modalities on delineating different categories in different scenes. To address the two problems, we first propose a locality-sensitive DeconvNet (LS-DeconvNet) to refine the boundary segmentation over each modality. LS-DeconvNet incorporates locally visual and geometric cues from the raw RGB-D data into each DeconvNet, which is able to learn to upsample the coarse convolutional maps with large context whilst recovering sharp object boundaries. Towards RGB-D fusion, we introduce a gated fusion layer to effectively combine the two LS-DeconvNets. This layer can learn to adjust the contributions of RGB and depth over each pixel for high-performance object recognition. Experiments on the large-scale SUN RGB-D dataset and the popular NYU-Depth v2 dataset show that our approach achieves new state-of-the-art results for RGB-D indoor semantic segmentation. Yanhua Cheng, Rui Cai 0002, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang |
CVPR | 4 |
| 2017 | Deep Crisp BoundariesabstractEdge detection had made significant progress with the help of deep Convolutional Networks (ConvNet). ConvNet based edge detectors approached human level performance on standard benchmarks. We provide a systematical study of these detector outputs, and show that they failed to accurately localize edges, which can be adversarial for tasks that require crisp edge inputs. In addition, we propose a novel refinement architecture to address the challenging problem of learning a crisp edge detector using ConvNet. Our method leverages a top-down backward refinement pathway, and progressively increases the resolution of feature maps to generate crisp edges. Our results achieve promising performance on BSDS500, surpassing human accuracy when using standard criteria, and largely outperforming state-of-the-art methods when using more strict criteria. We further demonstrate the benefit of crisp edge maps for estimating optical flow and generating object proposals. Yupei Wang, Xin Zhao 0012, Kaiqi Huang |
CVPR | 2 |
| 2016 | Learning temporally correlated representations using lstms for visual trackingabstractIn this paper, we propose to learn object representations with inference from temporal correlation in videos to achieve effective visual tracking. Unlike traditional methods which perform feature learning either at image level or based on intuitive temporal constraint, we employ the recurrent network with Long Short Term Memory (LSTM) units to directly learn temporally correlated representations of the objects in long sequences. The recurrent network is pre-trained offline with auxiliary data and then online optimized to adapt to the target-specific object. A structured SVM is employed to account for the temporally correlated object appearance as well as distinguish the object from background distraction. Experiment results not only show that the appearance and dynamic patterns of the objects can be characterized via temporally correlated feature learning, but also demonstrate that the proposed tracking algorithm performs favorably against the state-of-the-art methods. Qiaozhe Li, Xin Zhao 0012, Kaiqi Huang |
ICIP | 2 |
| 2016 | Semi-Supervised Multimodal Deep Learning for RGB-D Object Recognition
Yanhua Cheng, Xin Zhao 0012, Rui Cai 0002, Zhiwei Li 0006, Kaiqi Huang, Yong Rui |
IJCAI | 2 |
| 2015 | Convolutional Fisher Kernels for RGB-D Object RecognitionabstractThis paper studies the problem of improving object recognition using the novel RGB-D data. To address the problem, a new convolutional Fisher Kernels (CFK) method is proposed to represent RGB-D objects powerfully yet efficiently. The core idea of our approach is to integrate the both advantages of the convolutional neural networks (CNN) and Fisher Kernel encoding (FK): CNN model is flexible to adapt to new data sources, but requires for large amounts of training data with significant computational resources for good generalization, In comparison, FK encoding is able to represent objects powerfully and efficiently with small training data, however, its success highly depends on the well-designed SIFT features in literature, which may not be suitable for the new depth data. CFK can be interpreted as a two-layer feature learning structure to bridge the two models. The first layer employs a single-layer CNN to learn low-level translation ally invariant features for both RGB and depth data efficiently. The second layer aggregates the convolutional responses by FK encoding. Here 2D and 3D spatial pyramids are applied to further improve the Fisher vector representation of each modality. Experiments on RGB-D object recognition benchmarks demonstrate that our approach can achieve the state-of-the-art results. Yanhua Cheng, Rui Cai 0002, Xin Zhao 0012, Kaiqi Huang |
3DV | 3 |
| 2015 | Query Adaptive Similarity Measure for RGB-D Object RecognitionabstractThis paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object pose and scale changes, as well as intra-class variations, and (2) effectively fusing RGB and depth cues is still an open problem. To address these problems, this paper first proposes a new similarity measure based on dense matching, through which objects in comparison are warped and aligned, to better tolerate variations. Towards RGB and depth fusion, we argue that a constant and golden weight doesn't exist. The two modalities have varying contributions when comparing objects from different categories. To capture such a dynamic characteristic, a group of matchers equipped with various fusion weights is constructed, to explore the responses of dense matching under different fusion configurations. All the response scores are finally merged following a learning-to-combination way, which provides quite good generalization ability in practice. The proposed approach win the best results on several public benchmarks, e.g., achieves 92.7% top-1 test accuracy on the Washington RGB-D object dataset, with a 5.1% improvement over the state-of-the-art. Yanhua Cheng, Rui Cai 0002, Chi Zhang 0069, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang, Yong Rui |
ICCV | 5 |
| 2015 | Semi-supervised learning and feature evaluation for RGB-D object recognition
Yanhua Cheng, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
Comput. Vis. Image Underst. | 2 |
| 2014 | Semi-supervised Learning for RGB-D Object RecognitionabstractConventional supervised object recognition methods have been investigated for many years. Despite their successes, there are still two suffering limitations: (1) various information of an object is represented by artificial features only derived from RGB images, (2) lots of manually labeled data is required by supervised learning. To address those limitations, we propose a new semi-supervised learning framework based on RGB and depth (RGB-D) images to improve object recognition. In particular, our framework has two modules: (1) RGB and depth images are represented by convolutional-recursive neural networks to construct high level features, respectively, (2) co-training is exploited to make full use of unlabeled RGB-D instances due to the existing two independent views. Experiments on the standard RGB-D object dataset demonstrate that our method can compete against with other state-of-the-art methods with only 20% labeled data. Yanhua Cheng, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
ICPR | 2 |
| 2012 | Local Hypersphere Coding Based on Edges between Visual Words
Weiqiang Ren, Yongzhen Huang, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
ACCV (1) | 3 |
| 2012 | Feature coding via vector difference for image classificationabstractAn effective image representation is important to an image classification task. The most popular image representation framework utilizes a feature coding algorithm to encode the extracted low-level feature descriptors into a vector representation. In this paper, we analyze the recently developed feature coding methods in a general way. According to their common characteristics, we propose a new coding scheme to perform feature coding based on the vector difference in a high-dimensional space which is obtained by explicit feature maps. As we illustrate, our method has promising results with small codebook sizes and generalizes most existing coding methods in a unified form. Xin Zhao 0012, Yinan Yu, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICIP | 1 |
| 2012 | Semantic windows mining in sliding window based object detection
Junge Zhang, Xin Zhao 0012, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICPR | 2 |