VLDB 2026 Research / reviewers in the wild / expert
Xiaobin Zhu 0001
dblp:37/3108-1 · also Xiao-Bin Zhu 0001
· DBLP profile ↗
75ranked-venue papers
10as first author
46since 2021 · last 2026
0000-0003-2702-4136ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 5 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 17 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Geometry Graph Network: Unifying Local and Global Priors for Few-Shot LearningabstractIn few-shot learning, utilizing local and global geometric priors to capture both subtle local class metrics and coarse global structures within the meta-task are important to obtain discriminative embeddings. However, existing graph-based and curvature-based few-shot approaches only focus on either one kind of geometric prior but neglect the other. To effectively utilize the pros of these two paradigms, we propose a novel Dual-Geometry Graph Network (DGGN) to adaptively integrate the local and global geometric priors via two key pathways. Specifically, the local-wise metric modeling pathway utilizes Ollivier-Ricci curvature to capture task-specific local class metrics among the instances, and the global-wise connectivity modeling pathway utilizes resistive embedding to capture global instance distributions and connectivity patterns of the entire meta-task. In addition, we introduce two new regularization loss functions to explicitly enhance the geometric representation ability of the local and global pathways respectively. We validate that DGGN's superior performance stems from its adaptively topological refinements by measuring the graph edit distance, demonstrating its ability to match the underlying data distribution. Extensive experiments show that DGGN sets a new state-of-the-art on standard, cross-domain, and semi-supervised few-shot benchmarks. Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
AAAI | 2 |
| 2026 | VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality EvaluationabstractVideo captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated significant potential in caption generation, existing benchmarks inadequately address fine-grained evaluation, particularly in capturing spatial-temporal details critical for video generation. To address this gap, we introduce the Fine-grained Video Caption Evaluation Benchmark (VCapsBench), the first large-scale fine-grained benchmark comprising 5,677 (5K+) videos and 109,796 (100K+) question-answer pairs. These QA-pairs are systematically annotated across 21 fine-grained dimensions (e.g., camera movement, and shot type) that are empirically proven critical for text-to-video generation. We further introduce three metrics (Accuracy (AR), Inconsistency Rate (IR), Coverage Rate (CR)), and an automated evaluation pipeline leveraging a large language model (LLM) to verify caption quality via contrastive QA-pairs analysis. Our benchmark can advance the development of robust text-to-video models by providing actionable insights for caption optimization. Shi-Xue Zhang, Hongfa Wang, Duojun Huang, Xiaobin Zhu 0001, Xu-Cheng Yin |
AAAI | 5 |
| 2026 | Beyond Optical Flow: Latent Micro-Motion as Visual Evidence in UAV VideoabstractOptical flow has long dominated motion representation in video by focusing on explicit pixel displacement caused by object or camera movement. In this work, we argue that video also contains a largely overlooked form of motion information, namely latent micro-motion, which arises from subtle, structure-constrained responses of rigid components to physical interaction with the environment. We study this phenomenon in UAV video, a physically grounded setting where onboard structures are continuously exposed to aerodynamic forces. Although such micro-motions are low in amplitude and are often treated as noise or residual vibration, we show that they form a consistent visual signal that becomes observable through structure-aware and temporally aggregated analysis, even when using simple segmentation and coarse motion descriptors. Through an exploratory analysis, we demonstrate that micro-motion patterns exhibit clear structure and respond systematically to changes in wind conditions, with particularly strong sensitivity to wind direction and weaker dependence on wind magnitude in the examined scenarios. These observations suggest that micro-motion constitutes a distinct regime of motion information in video, complementary to explicit displacement, and motivate a broader reconsideration of how motion is represented and exploited in physically grounded multimedia scenarios. Bowen Zhang 0011, Song-Lu Chen, Xiaobin Zhu 0001, Xu-Cheng Yin |
ICMR | 5 |
| 2026 | Local attention alignment fusion network for domain adaptive water body segmentation
Xiaobin Zhu 0001, Xu Qizhi, Yongjie Xia, Xu-Cheng Yin |
Expert Syst. Appl. | 2 |
| 2026 | Degradation decomposition learning for self-supervised blind image super-resolution
Xiaobin Zhu 0001, Liuling Chen, Jingyan Qin, Xu-Cheng Yin |
Pattern Recognit. | 2 |
| 2026 | Amplitude-Phase Reconstruction for Non-Stationary Time-Series ForecastingabstractReal-world time-series can be decomposed into multiple interacting frequency components whose amplitudes and phases co-evolve over time. Such coupled dynamics can deform spectral trajectories, leading to pronounced non-stationarity in frequency-domain representations and substantial degradation in long-horizon forecasting performance. However, most existing frequency-domain forecasting methods do not explicitly model amplitude–phase interactions and often implicitly assume globally consistent spectral structures, which limits their ability to capture drifting spectra under non-stationary dynamics. To address this challenge, we propose a novel Amplitude–Phase Reconstruction Network (APRNet), a frequency-domain time-series forecasting framework that models amplitude–phase interactions from temporal and channel-wise perspectives. Specifically, we propose an innovative Amplitude-Phase Global Correlation (APGC) module to capture spectrum-wide amplitude-phase dependencies and derive frequency-wise calibration factors under non-stationary dynamics. The calibrated spectral representations are further reconstructed into the time domain, forming a closed time-frequency reconstruction loop that suppresses spectral drift and mitigates non-stationarity in temporal features. In addition, we propose a novel Discrete Kolmogorov–Arnold Network (D-KAN) to further enhance fine-grained amplitude–phase modeling. By combining discretized nonlinear activations with locally supported B-spline basis functions, D-KAN enables frequency-adaptive piecewise nonlinear modeling, improving sensitivity to localized spectral variations and enhancing high-frequency expressiveness. Extensive experiments verify the superior performance of our APRNet. Our codes are available at:https://github.com/LH325/APRNet. Xiaobin Zhu 0001, Lirui Deng 0001, Xu-Cheng Yin |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid FrameworkabstractOptical flow estimation is essential for video processing tasks, such as restoration and action recognition. The quality of videos is constantly increasing, with current standards reaching 8K resolution. However, optical flow methods are usually designed for low resolution and do not generalize to large inputs due to their rigid architectures. They adopt downscaling or input tiling to reduce the input size, causing a loss of details and global information. There is also a lack of optical flow benchmarks to judge the actual performance of existing methods on high-resolution samples. Previous works only conducted qualitative high-resolution evaluations on hand-picked samples. This paper fills this gap in optical flow estimation in two ways. We propose DPFlow, an adaptive optical flow architecture capable of generalizing up to 8K resolution inputs while trained with only low-resolution samples. We also introduce Kubric-NK, a new benchmark for evaluating optical flow methods with input resolutions ranging from 1K to 8K. Our high-resolution evaluation pushes the boundaries of existing methods and reveals new insights about their generalization capabilities. Extensive experimental results show that DPFlow achieves state-of-the-art results on the MPI-Sintel, KITTI 2015, Spring, and other high-resolution benchmarks. The code and dataset are available at https://github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/dpflow. Henrique Morimitsu, Xiaobin Zhu 0001, Roberto Marcondes Cesar Junior, Xiangyang Ji, Xu-Cheng Yin |
CVPR | 2 |
| 2025 | T-LLaVA: An Effective Saliency-Aware Slicing Strategy for Text Recognition
Mengze Wei, Xiaobin Zhu 0001, Xu-Cheng Yin |
ICDAR (1) | 5 |
| 2025 | Aligning enhanced feature representation for generalized zero-shot learning
Zhiyu Fang, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
Sci. China Inf. Sci. | 2 |
| 2025 | Multi-Scale Texture Fusion for Reference-Based Image Super-Resolution: New Dataset and Solution
Xiaobin Zhu 0001, Jingyan Qin, Roberto Marcondes Cesar Junior, Xu-Cheng Yin |
Int. J. Comput. Vis. | 2 |
| 2025 | Multi-view optimization and refinement for high-fidelity 4D Gaussian splatting
Jinhui Lin, Zhenyang Wei, Silei Shen, Xiaobin Zhu 0001, Xu-Cheng Yin |
Neurocomputing | 5 |
| 2024 | Recurrent Partial Kernel Network for Efficient Optical Flow EstimationabstractOptical flow estimation is a challenging task consisting of predicting per-pixel motion vectors between images. Recent methods have employed larger and more complex models to improve the estimation accuracy. However, this impacts the widespread adoption of optical flow methods and makes it harder to train more general models since the optical flow data is hard to obtain. This paper proposes a small and efficient model for optical flow estimation. We design a new spatial recurrent encoder that extracts discriminative features at a significantly reduced size. Unlike standard recurrent units, we utilize Partial Kernel Convolution (PKConv) layers to produce variable multi-scale features with a single shared block. We also design efficient Separable Large Kernels (SLK) to capture large context information with low computational cost. Experiments on public benchmarks show that we achieve state-of-the-art generalization performance while requiring significantly fewer parameters and memory than competing methods. Our model ranks first in the Spring benchmark without finetuning, improving the results by over 10% while requiring an order of magnitude fewer FLOPs and over four times less memory than the following published method without finetuning. The code is available at github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/rpknet. Henrique Morimitsu, Xiaobin Zhu 0001, Xiangyang Ji, Xu-Cheng Yin |
AAAI | 2 |
| 2024 | Arbitrary Time Information Modeling via Polynomial Approximation for Temporal Knowledge Graph EmbeddingabstractDistinguished from traditional knowledge graphs (KGs), temporal knowledge graphs (TKGs) must explore and reason over temporally evolving facts adequately. However, existing TKG approaches still face two main challenges, i.e., the limited capability to model arbitrary timestamps continuously and the lack of rich inference patterns under temporal constraints. In this paper, we propose an innovative TKGE method (PTBox) via polynomial decomposition-based temporal representation and box embedding-based entity representation to tackle the above-mentioned problems. Specifically, we decompose time information by polynomials and then enhance the model’s capability to represent arbitrary timestamps flexibly by incorporating the learnable temporal basis tensor. In addition, we model every entity as a hyperrectangle box and define each relation as a transformation on the head and tail entity boxes. The entity boxes can capture complex geometric structures and learn robust representations, improving the model’s inductive capability for rich inference patterns. Theoretically, our PTBox can encode arbitrary time information or even unseen timestamps while capturing rich inference patterns and higher-arity relations of the knowledge base. Extensive experiments on real-world datasets demonstrate the effectiveness of our method. Zhiyu Fang, Jingyan Qin, Xiaobin Zhu 0001, Xu-Cheng Yin |
LREC/COLING | 3 |
| 2024 | LayoutFormer: Hierarchical Text Detection Towards Scene Text UnderstandingabstractExisting scene text detectors generally focus on accu-rately detecting single-level (i.e., word-level, line-level, or paragraph-level) text entities without exploring the relationships among different levels of text entities. To comprehensively understand scene texts, detecting multi-level texts while exploring their contextual information is criti-cal. To this end, we propose a unified framework (dubbed LayoutFormer) for hierarchical text detection, which simultaneously conducts multi-level text detection and predicts the geometric layouts for promoting scene text understanding. In LayoutFormer, WordDecoder, LineDecoder, and Pa- raDecoder are proposed to be responsible for word-level text prediction, line-level text prediction, and paragraph- level text prediction, respectively. Meanwhile, WordDe-coder and ParaDecoder adaptively learn word-line and line-paragraph relationships, respectively. In addition, we propose a Prior Location Sampler to be used on multi-scale features to adaptively select a few representative foreground features for updating text queries. It can improve hierar- chical detection performance while significantly reducing the computational cost. Comprehensive experiments verify that our method achieves state-of-the-art performance on single-level and hierarchical text detection. Jia-Wei Ma, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
CVPR | 3 |
| 2024 | RAPIDFlow: Recurrent Adaptable Pyramids with Iterative Decoding for Efficient Optical Flow EstimationabstractExtracting motion information from videos with optical flow estimation is vital in multiple practical robot applications. Current optical flow approaches show remarkable accuracy, but top-performing methods have high computational costs and are unsuitable for embedded devices. Although some previous works have focused on developing low-cost optical flow strategies, their estimation quality has a noticeable gap with more robust methods. In this paper, we develop a novel method to efficiently estimate high-quality optical flow in embedded devices. Our proposed RAPIDFlow model combines efficient NeXt1D convolution blocks with a fully recurrent structure based on feature pyramids to decrease computational costs without significantly impacting estimation accuracy. The adaptable recurrent encoder produces multi-scale features with a single shared block, which allows us to adjust the pyramid length at inference time and make it more robust to changes in input size. Also, it enables our model to offer multiple tradeoffs between accuracy and speed to suit different applications. Experiments using a Jetson Orin NX embedded system on the MPI-Sintel and KITTI public benchmarks show that RAPIDFlow outperforms previous approaches by significant margins at faster speeds. Our code is available at https://github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/rapidflow. Henrique Morimitsu, Xiaobin Zhu 0001, Roberto Marcondes Cesar Junior, Xiangyang Ji, Xu-Cheng Yin |
ICRA | 2 |
| 2024 | Exploring Stable Meta-Optimization Patterns via Differentiable Reinforcement Learning for Few-Shot ClassificationabstractExisting few-shot learning methods generally focus on designing exquisite structures of meta-learners for learning task-specific prior to improve the discriminative ability of global embeddings. However, they often ignore the importance of learning stability in meta-training, making it difficult to obtain a relatively optimal model. From this key observation, we propose an innovative generic differentiable Reinforcement Learning (RL) strategy for few-shot classification. It aims to explore stable meta-optimization patterns in meta-training by learning generalizable optimizations for producing task-adaptive embeddings. Accordingly, our differentiable RL strategy models the embedding procedure of feature transformation layers in meta-learner to optimize the gradient flow implicitly. Also, we propose a memory module to associate historical and current task states and actions for exploring inter-task similarity. Notably, our RL-based strategy can be easily extended to various backbones. In addition, we propose a novel task state encoder to encode task representation, which fully explores inner-task similarities between support set and query set. Extensive experiments verify that our approach can improve the performance of different backbones and achieve promising results against state-of-the-art methods in few-shot classification. Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
ACM Multimedia | 2 |
| 2024 | Transformer-based Reasoning for Learning Evolutionary Chain of Events on Temporal Knowledge GraphabstractTemporal Knowledge Graph (TKG) reasoning often involves completing missing factual elements along the timeline. Although existing methods can learn good embeddings for each factual element in quadruples by integrating temporal information, they often fail to infer the evolution of temporal facts. This is mainly because of (1) insufficiently exploring the internal structure and semantic relationships within individual quadruples and (2) inadequately learning a unified representation of the contextual and temporal correlations among different quadruples. To overcome these limitations, we propose a novel Transformer-based reasoning model (dubbed ECEformer) for TKG to learn the Evolutionary Chain of Events (ECE). Specifically, we unfold the neighborhood subgraph of an entity node in chronological order, forming an evolutionary chain of events as the input for our model. Subsequently, we utilize a Transformer encoder to learn the embeddings of intra-quadruples for ECE. We then craft a mixed-context reasoning module based on the multi-layer perceptron (MLP) to learn the unified representations of inter-quadruples for ECE while accomplishing temporal knowledge reasoning. In addition, to enhance the timeliness of the events, we devise an additional time prediction task to complete effective temporal information within the learned unified representation. Extensive experiments on six benchmark datasets verify the state-of-the-art performance and the effectiveness of our method. Zhiyu Fang, Shuai-Long Lei, Xiaobin Zhu 0001, Shi-Xue Zhang, Xu-Cheng Yin, Jingyan Qin |
SIGIR | 3 |
| 2024 | Semi-supervised domain adaptation via subspace explorationabstractAbstract Recent methods of learning latent representations in Domain Adaptation (DA) often entangle the learning of features and exploration of latent space into a unified process. However, these methods can cause a false alignment problem and do not generalise well to the alignment of distributions with large discrepancy. In this study, the authors propose to explore a robust subspace for Semi‐Supervised Domain Adaptation (SSDA) explicitly. To be concrete, for disentangling the intricate relationship between feature learning and subspace exploration, the authors iterate and optimise them in two steps: in the first step, the authors aim to learn well‐clustered latent representations by aggregating the target feature around the estimated class‐wise prototypes; in the second step, the authors adaptively explore a subspace of an autoencoder for robust SSDA. Specially, a novel denoising strategy via class‐agnostic disturbance to improve the discriminative ability of subspace is adopted. Extensive experiments on publicly available datasets verify the promising and competitive performance of our approach against state‐of‐the‐art methods. Xiaobin Zhu 0001, Zhiyu Fang, Jingyan Qin, Xu-Cheng Yin |
IET Comput. Vis. | 2 |
| 2024 | Inverse-Like Antagonistic Scene Text Spotting via Reading-Order Estimation and Dynamic SamplingabstractScene text spotting is a challenging task, especially for inverse-like scene text, which has complex layouts,e.g., mirrored, symmetrical, or retro-flexed. In this paper, we propose a unified end-to-end trainable inverse-like antagonistic text spotting framework dubbed IATS, which can effectively spot inverse-like scene texts without sacrificing general ones. Specifically, we propose an innovative reading-order estimation module (REM) that extracts reading-order information from the initial text boundary generated by an initial boundary module (IBM). To optimize and train REM, we propose a joint reading-order estimation loss (LRE) consisting of a classification loss, an orthogonality loss, and a distribution loss. With the help of IBM, we can divide the initial text boundary into two symmetric control points and iteratively refine the new text boundary using a lightweight boundary refinement module (BRM) for adapting to various shapes and scales. To alleviate the incompatibility between text detection and recognition, we propose a dynamic sampling module (DSM) with a thin-plate spline that can dynamically sample appropriate features for recognition in the detected text region. Without extra supervision, the DSM can proactively learn to sample appropriate features for text recognition through the gradient returned by the recognition module. Extensive experiments on both challenging scene text and inverse-like scene text datasets demonstrate that our method achieves superior performance both on irregular and inverse-like text spotting. Shi-Xue Zhang, Xiaobin Zhu 0001, Hongfa Wang, Xu-Cheng Yin |
IEEE Trans. Image Process. | 3 |
| 2024 | Arbitrary Shape Text Detection via Boundary TransformerabstractIn arbitrary shape text detection, locating accurate text boundaries is challenging and non-trivial. Existing methods often suffer from indirect text boundary modeling or complex post-processing. In this article, we systematically present a unified coarse-to-fine framework via boundary learning for arbitrary shape text detection, which can accurately and efficiently locate text boundaries without post-processing. In our method, we explicitly model the text boundary via an innovative iterative boundary transformer in a coarse-to-fine manner. In this way, our method can directly gain accurate text boundaries and abandon complex post-processing to improve efficiency. Specifically, our method mainly consists of a feature extraction backbone, a boundary proposal module, and an iteratively optimized boundary transformer module. The boundary proposal module consisting of multi-layer dilated convolutions will predict important prior information (including classification map, distance field, and direction field) for generating coarse boundary proposals while guiding the boundary transformer's optimization. The boundary transformer module adopts an encoder-decoder structure, in which the encoder is constructed by multi-layer transformer blocks with residual connection while the decoder is a simple multi-layer perceptron network (MLP). Under the guidance of prior information, the boundary transformer module will gradually refine the coarse boundary proposals via iterative boundary deformation. Furthermore, we propose a novel boundary energy loss (BEL) that introduces an energy minimization constraint and an energy monotonically decreasing constraint to further optimize and stabilize the learning of boundary refinement. Extensive experiments on publicly available and challenging datasets demonstrate the state-of-the-art performance and promising efficiency of our method. Shi-Xue Zhang, Xiaobin Zhu 0001, Xu-Cheng Yin |
IEEE Trans. Multim. | 3 |
| 2023 | Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-ResolutionabstractAlthough existing image deep learning super-resolution (SR) methods achieve promising performance on benchmark datasets, they still suffer from severe performance drops when the degradation of the low-resolution (LR) input is not covered in training. To address the problem, we propose an innovative unsupervised method of Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-Resolution. Highly inspired by the generalized sampling theory, our method aims to enhance the strength of off-the-shelf SR methods trained on known degradations and adapt to unknown complex degradations to generate improved results. Specifically, we first conduct degradation estimation for each local image region by learning the internal distribution in an unsupervised manner via GAN. Instead of assuming degradation are spatially invariant across the whole image, we learn correction filters to adjust degradations to known degradations in a spatially variant way by a novel linearly-assembled pixel degradation-adaptive regression module (DARM). DARM is lightweight and easy to optimize on a dictionary of multiple pre-defined filter bases. Extensive experiments on synthetic and real-world datasets verify the effectiveness of our method both qualitatively and quantitatively. Code can be available at: https://github.com/edbca/DARSR. Xiaobin Zhu 0001, Jianqing Zhu, Shi-Xue Zhang, Jingyan Qin, Xu-Cheng Yin |
ICCV | 2 |
| 2023 | Graph fusion network for multi-oriented object detection
Shi-Xue Zhang, Xiaobin Zhu 0001, Jie-Bo Hou, Xu-Cheng Yin |
Appl. Intell. | 2 |
| 2023 | Attribute-Image Person Re-identification via Modal-Consistent Metric Learning
Jianqing Zhu, Liu Liu 0014, Yibing Zhan, Xiaobin Zhu 0001, Huanqiang Zeng, Dacheng Tao |
Int. J. Comput. Vis. | 4 |
| 2023 | Visible-infrared person re-identification using high utilization mismatch amending triplet loss
Jianqing Zhu, Hanxiao Wu, Huanqiang Zeng, Xiaobin Zhu 0001, Jingchang Huang, Canhui Cai |
Image Vis. Comput. | 5 |
| 2023 | Arbitrary Shape Text Detection via Segmentation With Probability MapsabstractArbitrary shape text detection is a challenging task due to the significantly varied sizes and aspect ratios, arbitrary orientations or shapes, inaccurate annotations, etc. Due to the scalability of pixel-level prediction, segmentation-based methods can adapt to various shape texts and hence attracted considerable attention recently. However, accurate pixel-level annotations of texts are formidable, and the existing datasets for scene text detection only provide coarse-grained boundary annotations. Consequently, numerous misclassified text pixels or background pixels inside annotations always exist, degrading the performance of segmentation-based text detection methods. Generally speaking, whether a pixel belongs to text or not is highly related to the distance with the adjacent annotation boundary. With this observation, in this paper, we propose an innovative and robust segmentation-based detection method via probability maps for accurately detecting text instances. To be concrete, we adopt a Sigmoid Alpha Function (SAF) to transfer the distances between boundaries and their inside pixels to a probability map. However, one probability map can not cover complex probability distributions well because of the uncertainty of coarse-grained text boundary annotations. Therefore, we adopt a group of probability maps computed by a series of Sigmoid Alpha Functions to describe the possible probability distributions. In addition, we propose an iterative model to learn to predict and assimilate probability maps for providing enough information to reconstruct text instances. Finally, simple region growth algorithms are adopted to aggregate probability maps to complete text instances. Experimental results demonstrate that our method achieves state-of-the-art performance in terms of detection accuracy on several benchmarks. Notably, our method with Watershed Algorithm as post-processing achieves the best F-measure on Total-Text (88.79%), CTW1500 (85.75%), and MSRA-TD500 (88.93%). Besides, our method achieves promising performance on multi-oriented datasets (ICDAR2015) and multilingual datasets (ICDAR2017-MLT). Code is available at: https://github.com/GXYM/TextPMs. Shi-Xue Zhang, Xiaobin Zhu 0001, Lei Chen 0069, Jie-Bo Hou, Xu-Cheng Yin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Towards open-set text recognition via label-to-prototype learning
Chang Liu 0083, Haibo Qin, Xiaobin Zhu 0001, Cheng-Lin Liu 0001, Xu-Cheng Yin |
Pattern Recognit. | 4 |
| 2023 | Hypersphere guided embedding for masked face recognition
Xiaobin Zhu 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin, Lei Chen 0069 |
Pattern Recognit. Lett. | 2 |
| 2023 | GiT: Graph Interactive Transformer for Vehicle Re-IdentificationabstractTransformers are more and more popular in computer vision, which treat an image as a sequence of patches and learn robust global features from the sequence. However, pure transformers are not entirely suitable for vehicle re-identification because vehicle re-identification requires both robust global features and discriminative local features. For that, a graph interactive transformer (GiT) is proposed in this paper. In the macro view, a list of GiT blocks are stacked to build a vehicle re-identification model, in where graphs are to extract discriminative local features within patches and transformers are to extract robust global features among patches. In the micro view, graphs and transformers are in an interactive status, bringing effective cooperation between local and global features. Specifically, one current graph is embedded after the former level's graph and transformer, while the current transform is embedded after the current graph and the former level's transformer. In addition to the interaction between graphs and transforms, the graph is a newly-designed local correction graph, which learns discriminative local features within a patch by exploring nodes' relationships. Extensive experiments on three large-scale vehicle re-identification datasets demonstrate that our GiT method is superior to state-of-the-art vehicle re-identification approaches. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Huanqiang Zeng |
IEEE Trans. Image Process. | 4 |
| 2023 | HFENet: Hybrid Feature Enhancement Network for Detecting Texts in Scenes and Traffic PanelsabstractText detection in complex scene images is a challenging task for intelligent transportation. Existing scene text detection methods often adopt multi-scale feature learning strategies to extract informative feature representations for covering objects of various sizes. However, the sampling operation inherent in multi-scale feature generation can easily impair high-frequency details (e.g., textures and boundaries), which are critical for text detection. In this work, we propose an innovative Hybrid Feature Enhancement Network (dubbed HFENet) to explicitly improve the quality of high-frequency information for detecting texts in scenes and traffic panels. To be concrete, we propose a simple yet effective self-guided feature enhancement module (SFEM) for globally lifting feature representations to highly discriminative and high-frequency abundant ones. Notably, our SFEM is pluggable and will be removed after training without introducing extra computational costs. In addition, due to the challenge and importance of accurately predicting boundaries for text detection, we propose a novel boundary enhancement module (BEM) to explicitly strengthen local feature representations in the guidance of boundary annotation for accurate localization. Extensive experiments on multiple publicly available datasets (i.e., MSRA-TD500, CTW1500, Total-Text, Traffic Guide Panel Dataset, Chinese Road Plate Dataset, and ASAYAR_TXT) verify the state-of-the-art performance of our method. Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | AdapNet: Adaptability Decomposing Encoder-Decoder Network for Weakly Supervised Action Recognition and LocalizationabstractThe point process is a solid framework to model sequential data, such as videos, by exploring the underlying relevance. As a challenging problem for high-level video understanding, weakly supervised action recognition and localization in untrimmed videos have attracted intensive research attention. Knowledge transfer by leveraging the publicly available trimmed videos as external guidance is a promising attempt to make up for the coarse-grained video-level annotation and improve the generalization performance. However, unconstrained knowledge transfer may bring about irrelevant noise and jeopardize the learning model. This article proposes a novel adaptability decomposing encoder-decoder network to transfer reliable knowledge between the trimmed and untrimmed videos for action recognition and localization by bidirectional point process modeling, given only video-level annotations. By decomposing the original features into the domain-adaptable and domain-specific ones based on their adaptability, trimmed-untrimmed knowledge transfer can be safely confined within a more coherent subspace. An encoder-decoder-based structure is carefully designed and jointly optimized to facilitate effective action classification and temporal localization. Extensive experiments are conducted on two benchmark data sets (i.e., THUMOS14 and ActivityNet1.3), and the experimental results clearly corroborate the efficacy of our method. Xiaoyu Zhang 0002, Haichao Shi, Xiaobin Zhu 0001, Peng Li 0035, Jing Dong 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Kernel Proposal Network for Arbitrary Shape Text DetectionabstractSegmentation-based methods have achieved great success for arbitrary shape text detection. However, separating neighboring text instances is still one of the most challenging problems due to the complexity of texts in scene images. In this article, we propose an innovative kernel proposal network (dubbed KPN) for arbitrary shape text detection. The proposed KPN can separate neighboring text instances by classifying different texts into instance-independent feature maps, meanwhile avoiding the complex aggregation process existing in segmentation-based arbitrary shape text detection methods. To be concrete, our KPN will predict a Gaussian center map for each text image, which will be used to extract a series of candidate kernel proposals (i.e., dynamic convolution kernel) from the embedding feature maps according to their corresponding keypoint positions. To enforce the independence between kernel proposals, we propose a novel orthogonal learning loss (OLL) via orthogonal constraints. Specifically, our kernel proposals contain important self-information learned by network and location information by position embedding. Finally, kernel proposals will individually convolve all embedding feature maps for generating individual embedded maps of text instances. In this way, our KPN can effectively separate neighboring text instances and improve the robustness against unclear boundaries. To the best of our knowledge, our work is the first to introduce the dynamic convolution kernel strategy to efficiently and effectively tackle the adhesion problem of neighboring text instances in text detection. Experimental results on challenging datasets verify the impressive performance and efficiency of our method. The code and model are available at https://github.com/GXYM/KPN. Shi-Xue Zhang, Xiaobin Zhu 0001, Jie-Bo Hou, Xu-Cheng Yin |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Learning Aligned Cross-Modal Representation for Generalized Zero-Shot ClassificationabstractLearning a common latent embedding by aligning the latent spaces of cross-modal autoencoders is an effective strategy for Generalized Zero-Shot Classification (GZSC). However, due to the lack of fine-grained instance-wise annotations, it still easily suffer from the domain shift problem for the discrepancy between the visual representation of diversified images and the semantic representation of fixed attributes. In this paper, we propose an innovative autoencoder network by learning Aligned Cross-Modal Representations (dubbed ACMR) for GZSC. Specifically, we propose a novel Vision-Semantic Alignment (VSA) method to strengthen the alignment of cross-modal latent features on the latent subspaces guided by a learned classifier. In addition, we propose a novel Information Enhancement Module (IEM) to reduce the possibility of latent variables collapse meanwhile encouraging the discriminative ability of latent variables. Extensive experiments on publicly available datasets demonstrate the state-of-the-art performance of our method. Zhiyu Fang, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
AAAI | 2 |
| 2022 | SD-GAN: Semantic Decomposition for Face Image Synthesis with Discrete AttributeabstractManipulating latent code in generative adversarial networks (GANs) for facial image synthesis mainly focuses on continuous attribute synthesis (e.g., age, pose and emotion), while discrete attribute synthesis (like face mask and eyeglasses) receives less attention. Directly applying existing works to facial discrete attributes may cause inaccurate results. In this work, we propose an innovative framework to tackle challenging facial discrete attribute synthesis via semantic decomposing, dubbed SD-GAN. To be concrete, we explicitly decompose the discrete attribute representation into two components, i.e. the semantic prior basis and offset latent representation. The semantic prior basis shows an initializing direction for manipulating face representation in the latent space. The offset latent presentation obtained by 3D-aware semantic fusion network is proposed to adjust prior basis. In addition, the fusion network integrates 3D embedding for better identity preservation and discrete attribute synthesis. The combination of prior basis and offset latent representation enable our method to synthesize photo-realistic face images with discrete attributes. Notably, we construct a large and valuable dataset MEGN (Face Mask and Eyeglasses images crawled from Google and Naver) for completing the lack of discrete attributes in the existing dataset. Extensive qualitative and quantitative experiments demonstrate the state-of-the-art performance of our method. Our code is available at an anonymous website: https://github.com/MontaEllis/SD-GAN. Kangneng Zhou, Xiaobin Zhu 0001, Daiheng Gao, Kai Lee, Xinjie Li 0002, Xu-Cheng Yin |
ACM Multimedia | 2 |
| 2022 | A sample-proxy dual triplet loss function for object re-identificationabstractAbstract Object re‐identification, such as vehicle re‐identification or pedestrian re‐identification, plays a significant role in intelligent video surveillance systems for public security. Due to viewpoint variations and appearance changes, both pedestrians and vehicles usually have complex intra‐class variations. However, most existing object re‐identification methods often use a sample‐level triplet loss function cooperating with a single‐proxy softmax loss function, which could not handle complex intra‐class variations well. In this paper, a sample‐proxy dual triplet (SPDT) loss function is proposed, which works with a multi‐proxy softmax (MPS) loss function. The MPS loss function is in charge of learning multiple proxies to represent a class. The SPDT loss function is responsible for enlarging inter‐class distances as well as shrinking intra‐class distances on both sample and proxy levels. Therefore, the method not only handles multi‐proxy intra‐class variations but also fully learns discrimination on samples and proxies. Experiments on two large datasets, that is, VeRi776 and DukeMTMC‐reID, demonstrate that the method is superior to state‐of‐the‐art object re‐identification approaches. Hanxiao Wu, Fei Shen 0004, Jianqing Zhu, Huanqiang Zeng, Xiaobin Zhu 0001, Zhen Lei 0001 |
IET Image Process. | 5 |
| 2022 | Scene text detection via decoupled feature pyramid networks
Jie-Bo Hou, Xiaobin Zhu 0001, Jingyan Qin |
Int. J. Document Anal. Recognit. | 3 |
| 2022 | An Efficient Multiresolution Network for Vehicle ReidentificationabstractIn general, vehicle images have varying resolutions due to vehicles’ movements and different camera settings. However, most existing vehicle reidentification models are single-resolution deep networks trained with preuniformly resizing vehicle images, which underestimate adverse effects of varying resolutions and lead to unsatisfactory performance. A straightforward solution for dealing with varying resolutions is to train multiple vehicle reidentification models. Each model is independently trained with images of a specific resolution. However, this straightforward solution requires significant overhead and ignores intrinsic associations among different resolution images. For that, an efficient multiresolution network (EMRN) is proposed for vehicle reidentification in this article. First, EMRN embeds a newly designed multiresolution feature dimension uniform module (MR-FDUM) behind a traditional backbone network (i.e., ResNet-50). As a result, the whole model can extract fixed dimensional features from different resolution images so that it can be trained with one loss function of fixed dimensional parameters rather than training multiple models. Second, a multiresolution image randomly feeding strategy is designed to train EMRN, making each minibatch data of a random resolution during the training process. Consequently, EMRN can implicitly learn collaborative multiresolution features via only a unitary deep network. The experiments on three large-scale data sets, i.e., VeRi776, VehicleID, and VRIC, demonstrate that EMRN is superior to state-of-the-art vehicle reidentification methods. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Jingchang Huang, Huanqiang Zeng, Zhen Lei 0001, Canhui Cai |
IEEE Internet Things J. | 3 |
| 2022 | Two-Stage Copy-Move Forgery Detection With Self Deep Matching and Proposal SuperGlueabstractCopy-move forgery detection identifies a tampered image by detecting pasted and source regions in the same image. In this paper, we propose a novel two-stage framework specially for copy-move forgery detection. The first stage is a backbone self deep matching network, and the second stage is named as Proposal SuperGlue. In the first stage, atrous convolution and skip matching are incorporated to enrich spatial information and leverage hierarchical features. Spatial attention is built on self-correlation to reinforce the ability to find appearance similar regions. In the second stage, Proposal SuperGlue is proposed to remove false-alarmed regions and remedy incomplete regions. Specifically, a proposal selection strategy is designed to enclose highly suspected regions based on proposal generation and backbone score maps. Then, pairwise matching is conducted among candidate proposals by deep learning based keypoint extraction and matching, i.e., SuperPoint and SuperGlue. Integrated score map generation and refinement methods are designed to integrate results of both stages and obtain optimized results. Our two-stage framework unifies end-to-end deep matching and keypoint matching by obtaining highly suspected proposals, and opens a new gate for deep learning research in copy-move forgery detection. Experiments on publicly available datasets demonstrate the effectiveness of our two-stage framework. Xiaobin Zhu 0001, Shengwei Xu |
IEEE Trans. Image Process. | 3 |
| 2022 | Exploring Spatial Significance via Hybrid Pyramidal Graph Network for Vehicle Re-IdentificationabstractExisting vehicle re-identification methods commonly use spatial pooling operations to aggregate feature maps extracted via off-the-shelf backbone networks, such as visual geometry group network (VGGNet), Google network (GoogLeNet) and residual network (ResNet). They ignore exploring the spatial significance of feature maps, eventually degrading the vehicle re-identification performance. In this paper, firstly, an innovative spatial graph network (SGN) is proposed to elaborately explore the spatial significance of feature maps. The SGN stacks multiple spatial graphs (SGs). Each SG assigns feature map’s elements as nodes and utilizes spatial neighborhood relationships to determine edges among nodes. During the SGN’s propagation, each node and its spatial neighbors on an SG are aggregated to the next SG. On the next SG, each aggregated node is re-weighted with a learnable parameter to find the significance at the corresponding location. Secondly, a novel pyramidal graph network (PGN) is designed to comprehensively explore the spatial significance of feature maps at multiple scales. The PGN organizes multiple SGNs in a pyramidal manner and makes each SGN handles feature maps of a specific scale. Finally, a hybrid pyramidal graph network (HPGN) is developed by embedding the PGN behind a ResNet-50 based backbone network. Extensive experiments on three large scale vehicle databases (i.e., VeRi776, VehicleID, and VeRi-Wild) demonstrate that the proposed HPGN is superior to state-of-the-art vehicle re-identification approaches in terms of accuracy, parameter cost, and computation cost. In addition, experiments show that the proposed PGN is universal to various backbone networks. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Jingchang Huang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Adaptive Boundary Proposal Network for Arbitrary Shape Text DetectionabstractArbitrary shape text detection is a challenging task due to the high complexity and variety of scene texts. In this work, we propose a novel adaptive boundary proposal network for arbitrary shape text detection, which can learn to directly produce accurate boundary for arbitrary shape text without any post-processing. Our method mainly consists of a boundary proposal model and an innovative adaptive boundary deformation model. The boundary proposal model constructed by multi-layer dilated convolutions is adopted to produce prior information (including classification map, distance field, and direction field) and coarse boundary proposals. The adaptive boundary deformation model is an encoder-decoder network, in which the encoder mainly consists of a Graph Convolutional Network (GCN) and a Recurrent Neural Network (RNN). It aims to perform boundary deformation in an iterative way for obtaining text instance shape guided by prior information from the boundary proposal model. In this way, our method can directly and efficiently generate accurate text boundaries without complex post-processing. Extensive experiments on publicly available datasets demonstrate the state-of-the-art performance of our method. Code is available at the website: https://github.com/GXYM/TextBPN. Shi-Xue Zhang, Xiaobin Zhu 0001, Hongfa Wang, Xu-Cheng Yin |
ICCV | 2 |
| 2021 | Dynamic Receptive Field Adaptation for Attention-Based Text Recognition
Haibo Qin, Xiaobin Zhu 0001, Xu-Cheng Yin |
ICDAR (2) | 3 |
| 2021 | Real-World Image Super-Resolution Via Spatio-Temporal Correlation NetworkabstractSuper-resolving real-world image is very challenging due to the degradations in real-world low-resolution images are highly complicated. In this paper, we propose a novel Spatio-temporal Correlation Network (STCN) for real-world single image super-resolution. Specifically, we adopt a very deep network which consists of several attention groups. Each attention group (AG) contains a series of residual channel attention blocks (RCABs) and one spatio-temporal correlation block (STCB). Notably, STCB mainly consists of a residual 3D convolution, and aims to fully explore the local spatial and temporal correlations between channels of feature maps generated by RCABs for selectively capturing more informative features. In addition, we propose an innovative dual restriction (DR) through a simple degradation model to reduce the possible space of mapping functions in super-resolution. Experiments conducted on two public available real-world datasets demonstrate the superior performance of our method. Xiaobin Zhu 0001, Xu-Cheng Yin |
ICME | 2 |
| 2021 | A spatial structural similarity triplet loss for auxiliary vehicle re-identification
Jianqing Zhu, Liu Liu 0014, Xiaobin Zhu 0001, Huanqiang Zeng |
Sci. China Inf. Sci. | 3 |
| 2021 | Multi-orientation scene text detection with scale-guided regression
Jie-Bo Hou, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
Neurocomputing | 3 |
| 2021 | GCCNet: Grouped channel composition network for scene text detection
Chang Liu 0083, Jie-Bo Hou, Long-Huang Wu, Xiaobin Zhu 0001, Lei Xiao 0001, Xu-Cheng Yin |
Neurocomputing | 5 |
| 2021 | Video super-resolution based on a spatio-temporal matching network
Xiaobin Zhu 0001, Zhuangzi Li, Jungang Lou, Qing Shen 0005 |
Pattern Recognit. | 1 |
| 2021 | Detecting Text in Scene and Traffic Guide Panels With Attention Anchor MechanismabstractText detection in complex scene images is a challenging task for intelligent transportation. Recently, anchor mechanisms are widely utilized in scene text detection tasks. However, in existing methods, anchors are generally predefined empirically, degrading robustness to complex scenarios with various sizes and orientation variations. In this paper, we propose a novel Attention Anchor Mechanism (AAM), especially targeting at predicting appropriate anchors for each pixel. To be concrete, we regard a series of predefined anchors as basic anchors and utilize an attention model to predict weights corresponding to basic anchors. Consequently, the weighted sum of basic anchors in each pixel can obtain a predicted anchor. In this way, the gap between the predicted anchors and the corresponding ground truth boxes could be narrowed, making the network easier to regress. For facilitating the design of basic anchors, we adopt a dimension-decomposition mechanism to predict width, height, and angle of anchors, respectively. Extensive experiments on several public datasets demonstrate that our method achieves state-of-the-art performance. Jie-Bo Hou, Xiaobin Zhu 0001, Chang Liu 0083, Long-Huang Wu, Hongfa Wang, Xu-Cheng Yin |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2020 | Deep Relational Reasoning Graph Network for Arbitrary Shape Text DetectionabstractArbitrary shape text detection is a challenging task due to the high variety and complexity of scenes texts. In this paper, we propose a novel unified relational reasoning graph network for arbitrary shape text detection. In our method, an innovative local graph bridges a text proposal model via Convolutional Neural Network (CNN) and a deep relational reasoning network via Graph Convolutional Network (GCN), making our network end-to-end trainable. To be concrete, every text instance will be divided into a series of small rectangular components, and the geometry attributes (e.g., height, width, and orientation) of the small components will be estimated by our text proposal model. Given the geometry attributes, the local graph construction model can roughly establish linkages between different text components. For further reasoning and deducing the likelihood of linkages between the component and its neighbors, we adopt a graph-based network to perform deep relational reasoning on local graphs. Experiments on public available datasets demonstrate the state-of-the-art performance of our method. Code is available at https://github.com/GXYM/DRRG. Shi-Xue Zhang, Xiaobin Zhu 0001, Jie-Bo Hou, Chang Liu 0083, Hongfa Wang, Xu-Cheng Yin |
CVPR | 2 |
| 2020 | Global Context-Based Network with Transformer for Image2latexabstractImage2latex usually means converts mathematical formulas in images into latex markup. It is a very challenging job due to the complex two-dimensional structure, variant scales of input, and very long representation sequence. Many researchers use encoder-decoder based model to solve this task and achieved good results. However, these methods don't make full use of the structure and position information of the formula. To solve this problem, we propose a global context-based network with transformer that can (1) learn a more powerful and robust intermediate representation via aggregating global features and (2) encode position information explicitly and (3) learn latent dependencies between symbols by using self-attention mechanism. The experimental results on the dataset IM2LATEX-100K demonstrate the effectiveness of our method. Nuo Pang, Xiaobin Zhu 0001, Jixuan Li, Xu-Cheng Yin |
ICPR | 3 |
| 2020 | Body Symmetry and Part-Locality-Guided Direct Nonparametric Deep Feature Enhancement for Person ReidentificationabstractIn recent years, deep learning (DL) has been successfully and widely applied in the person reidentification (Re-ID). However, the DL-based person Re-ID methods face a bottleneck that the scales of most existing person Re-ID databases are not large enough for training very deep models. To address this problem, a body symmetry and part-locality-guided direct nonparametric deep feature enhancement (DNDFE) method is proposed in this article. Based on the observation that the body symmetry and part locality are two important appearance properties inherited in the upright walking persons, the proposed method designs two nonparametric layers, namely, the body symmetry average pooling and local normalization layers, to construct a DNDFE module to well explore the body symmetry and part locality properties. The proposed DNDFE module could be directly embedded between the traditional deep feature learning module and similarity learning module to enhance the DL features so as to improve the person Re-ID performance. The experimental results have shown that the proposed DNDFE method is superior to multiple state-of-the-art person Re-ID methods in terms of accuracy and efficiency. Jianqing Zhu, Huanqiang Zeng, Jingchang Huang, Xiaobin Zhu 0001, Zhen Lei 0001, Canhui Cai, Lixin Zheng |
IEEE Internet Things J. | 4 |
| 2020 | Attention-aware perceptual enhancement nets for low-resolution image classification
Xiaobin Zhu 0001, Zhuangzi Li, Xianbo Li |
Inf. Sci. | 1 |
| 2020 | HAM: Hidden Anchor Mechanism for Scene Text DetectionabstractDirect regression and anchor are the two mainly effective and prevailing mechanisms in the paradigm of scene text detection. However, the use of direct regression-based methods may be challenging during optimization without the help of anchors as references. Unfortunately, the anchor-based methods always suffer from the careful design of the anchors, degrading the robustness to complex scenes. To address the above-mentioned problems, we propose a novel hidden anchor mechanism (HAM) especially for scene text detection. The predictions of anchors are innovatively regarded as hidden layers, and the weighted sum of the predictions is integrated into a direct regression-based network. Hence, the architecture of our HAM still has the characteristic of simplicity as with direct regression-based methods. Moreover, it is easier to optimize anchors as references with this type of method than with direct regression-based methods. In this way, our network can take advantage of both direct regression and anchor mechanisms. In addition, we decouple three kinds of one-dimensional anchors from three-dimensional anchors, greatly reducing the number of anchors in text bounding box matching without performance degradation. We also propose a post-processing technique for long text detection, named iterative regression box (IRB), which takes a few additional computational costs and can be easily generalized to other methods. Experiments on several public datasets demonstrate that the proposed method achieves state-of-the-art performance. Code is available athttps://github.com/hjbplayer/HAM. Jie-Bo Hou, Xiaobin Zhu 0001, Chang Liu 0083, Kekai Sheng, Long-Huang Wu, Hongfa Wang, Xu-Cheng Yin |
IEEE Trans. Image Process. | 2 |
| 2019 | Learning Transferable Self-Attentive Representations for Action Recognition in Untrimmed Videos with Weak SupervisionabstractAction recognition in videos has attracted a lot of attention in the past decade. In order to learn robust models, previous methods usually assume videos are trimmed as short sequences and require ground-truth annotations of each video frame/sequence, which is quite costly and time-consuming. In this paper, given only video-level annotations, we propose a novel weakly supervised framework to simultaneously locate action frames as well as recognize actions in untrimmed videos. Our proposed framework consists of two major components. First, for action frame localization, we take advantage of the self-attention mechanism to weight each frame, such that the influence of background frames can be effectively eliminated. Second, considering that there are trimmed videos publicly available and also they contain useful information to leverage, we present an additional module to transfer the knowledge from trimmed videos for improving the classification performance in untrimmed ones. Extensive experiments are conducted on two benchmark datasets (i.e., THUMOS14 and ActivityNet1.3), and experimental results clearly corroborate the efficacy of our method. Xiaoyu Zhang 0002, Haichao Shi, Kai Zheng 0001, Xiaobin Zhu 0001, Lixin Duan |
AAAI | 5 |
| 2019 | Residual Invertible Spatio-Temporal Network for Video Super-ResolutionabstractVideo super-resolution is a challenging task, which has attracted great attention in research and industry communities. In this paper, we propose a novel end-to-end architecture, called Residual Invertible Spatio-Temporal Network (RISTN) for video super-resolution. The RISTN can sufficiently exploit the spatial information from low-resolution to high-resolution, and effectively models the temporal consistency from consecutive video frames. Compared with existing recurrent convolutional network based approaches, RISTN is much deeper but more efficient. It consists of three major components: In the spatial component, a lightweight residual invertible block is designed to reduce information loss during feature transformation and provide robust feature representations. In the temporal component, a novel recurrent convolutional model with residual dense connections is proposed to construct deeper network and avoid feature degradation. In the reconstruction component, a new fusion method based on the sparse strategy is proposed to integrate the spatial and temporal features. Experiments on public benchmark datasets demonstrate that RISTN outperforms the state-ofthe-art methods. Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002 |
AAAI | 1 |
| 2019 | Detecting Text in News Images with Similarity Embedded ProposalsabstractText extraction plays an important role in news images analysis tasks. However, the conglutination of subtitles and station logos makes text detection challenging. In this paper, we develop an effective news text detection framework by introducing a novel similarity embedded proposal mechanism. The main idea is to predict similarity for each fine-scale coarse proposal to help construct text bounding boxes. Specifically, a CNN and bi-directional LSTM based network is used to produce vectors embedded in coarse proposals provided by Connectionist Text Proposal Network (CTPN). Notably, similarity embedded proposal mechanism can be generalized to other sub-text level text detection models. Comparing to the state-of-the-art method (CTPN), our framework improves F-measure by 25.2% on our Private News Dataset and 8.9% on ICDAR 2013 benchmarks, respectively. Miaotong Jiang, Jie-Bo Hou, Xiaobin Zhu 0001, Xu-Cheng Yin |
ICDAR | 4 |
| 2019 | Deep Super-Resolution Hashing Network for Low-Resolution Image Retrieval
Zhuangzi Li, Naiguang Zhang, Xiaobin Zhu 0001, Peng Li 0035 |
ICIG (3) | 5 |
| 2019 | Active semi-supervised learning based on self-expressive correlation with generative adversarial networks
Xiaoyu Zhang 0002, Haichao Shi, Xiaobin Zhu 0001, Peng Li 0035 |
Neurocomputing | 3 |
| 2019 | A novel framework for semantic segmentation with generative adversarial network
Xiaobin Zhu 0001, Xiaoyu Zhang 0002, Lei Wang 0101 |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Deep convolutional representations and kernel extreme learning machines for image classification
Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002, Peng Li 0035, Lei Wang 0101 |
Multim. Tools Appl. | 1 |
| 2019 | Hash Code Reconstruction for Fast Similarity SearchabstractLearning to hash is a popular technique for fast similarity search on a large-scale image database. However, many hashing methods do not achieve satisfactory results because of the quantization loss in the straightforward binary code generation procedure. In order to address this problem, we propose a novel hash code reconstruction framework for existing unsupervised hashing methods. In our proposed approach, the hash codes are generated through reconstructing the original images with relaxed hamming vector representation, such that the final learned codes will be more approximate to characterize the intrinsic image structure. Moreover, our proposed hash code reconstruction algorithm is very efficient for computing, which can be generalized to various hashing methods. Extensive experiments are conducted on four public image datasets by incorporating our proposed scheme with different hashing methods, and the comparison results have shown that significant performance improvements can be achieved with minor additional time cost for fast similarity search task. Peng Li 0035, Xiaobin Zhu 0001, Xiaoyu Zhang 0002, Peng Ren 0001, Lei Wang 0101 |
IEEE Signal Process. Lett. | 2 |
| 2019 | Adversarial Learning for Constrained Image Splicing Detection and Localization Based on Atrous ConvolutionabstractConstrained image splicing detection and localization (CISDL), which investigates two input suspected images and identifies whether one image has suspected regions pasted from the other, is a newly proposed challenging task for image forensics. In this paper, we propose a novel adversarial learning framework to learn a deep matching network for CISDL. Our framework mainly consists of three building blocks. First, a deep matching network based on atrous convolution (DMAC) aims to generate two high-quality candidate masks, which indicate suspected regions of the two input images. In DMAC, atrous convolution is adopted to extract features with rich spatial information, a correlation layer based on a skip architecture is proposed to capture hierarchical features, and atrous spatial pyramid pooling is constructed to localize tampered regions at multiple scales. Second, a detection network is designed to rectify inconsistencies between the two corresponding candidate masks. Finally, a discriminative network drives the DMAC network to produce masks that are hard to distinguish from ground-truth ones. The detection network and the discriminative network collaboratively supervise the training of DMAC in an adversarial way. Besides, a sliding window-based matching strategy is investigated for high-resolution images matching. Extensive experiments, conducted on five groups of datasets, demonstrate the effectiveness of the proposed framework and the superior performance of DMAC. Xiaobin Zhu 0001, Xianfeng Zhao, Yun Cao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Image Classification Using Convolutional Neural Networks and Kernel Extreme Learning MachinesabstractWe know that convolutional neural networks are good at learning invariant features, but not always optimal for classification. Contrarily, Kernel Extreme Learning Machines (KELMs) are good at approximating any target continuous function with extremely fast speed, but cannot learn complicated invariances. In this paper, we propose a novel image classification framework, in which KELM instead of Softmax function is adopted as a classifier in the convolutional neural network (CNN) architecture for promoting the performance of image classification. Experiments conducted on the publicly available datasets demonstrate the superior performance of the proposed method. Zhuangzi Li, Xiaobin Zhu 0001, Lei Wang 0101, Peiyu Guo |
ICIP | 2 |
| 2018 | Generative Adversarial Image Super-Resolution Through Deep Dense Skip ConnectionsabstractAbstract Recently, image super‐resolution works based on Convolutional Neural Networks (CNNs) and Generative Adversarial Nets (GANs) have shown promising performance. However, these methods tend to generate blurry and over‐smoothed super‐resolved (SR) images, due to the incomplete loss function and powerless architectures of networks. In this paper, a novel generative adversarial image super‐resolution through deep dense skip connections (GSR‐DDNet), is proposed to solve the above‐mentioned problems. It aims to take advantage of GAN's ability of modeling data distributions, so that GSR‐DDNet can select informative feature representation and model the mapping across the low‐quality and high‐quality images in an adversarial way. The pipeline of the proposed method consists of three main components: 1) The generator of a novel dense skip connection network with the deep structure for learning robust mapping function is proposed to generate SR images from low‐resolution images; 2) The feature extraction network based on VGG‐19 is adopted to capture high frequency feature maps for content loss; and 3) The discriminator with Wasserstein distance is adopted to identify the overall style of SR and ground‐truth images. Experiments conducted on four publicly available datasets demonstrate the superiority against the state‐of‐the‐art methods. Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002, Hai-Sheng Li 0002, Lei Wang 0101 |
Comput. Graph. Forum | 1 |
| 2017 | ListNet-based object proposals ranking
Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Qingxiao Guan, Xianfeng Zhao |
Neurocomputing | 3 |
| 2016 | Spatially Regularized Streaming Sensor SelectionabstractSensor selection has become an active topic aimed at energy saving, information overload prevention, and communication cost planning in sensor networks. In many real applications, often the sensors' observation regions have overlaps and thus the sensor network is inherently redundant. Therefore it is important to select proper sensors to avoid data redundancy. This paper focuses on how to incrementally select a subset of sensors in a streaming scenario to minimize information redundancy, and meanwhile meet the power consumption constraint. We propose to perform sensor selection in a multi-variate interpolation framework, such that the data sampled by the selected sensors can well predict those of the inactive sensors. Importantly, we incorporate sensors' spatial information as two regularizers, which leads to significantly better prediction performance. We also define a statistical variable to store sufficient information for incremental learning, and introduce a forgetting factor to track sensor streams' evolvement. Experiments on both synthetic and real datasets validate the effectiveness of the proposed method. Moreover, our method is over 10 times faster than the state-of-the-art sensor selection algorithm. Weishan Dong, Xiangfeng Wang 0001, Junchi Yan, Xiaobin Zhu 0001, Qingshan Liu 0001, Xin Zhang 0008 |
AAAI | 6 |
| 2016 | Weighted hierarchical geographic information description model for social relation estimation
Kai Zhang 0079, Xiao-chun Yun, Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Chao Li 0062 |
Neurocomputing | 4 |
| 2016 | Socio-mobile landmark recognition using local features with adaptive region selection
Chunjie Zhang 0001, Yifan Zhang 0001, Xiaobin Zhu 0001, Zhe Xue, Qingming Huang, Qi Tian 0001 |
Neurocomputing | 3 |
| 2016 | Boosted random contextual semantic space based representation for visual recognition
Chunjie Zhang 0001, Zhe Xue, Xiaobin Zhu 0001, Huanian Wang, Qingming Huang, Qi Tian 0001 |
Inf. Sci. | 3 |
| 2015 | Saliency detection using two-stage scoringabstractIn this paper, we propose a novel saliency detection approach, which is robust to images with complex background. In our algorithm, an intuitive and straightforward pre-treatment method is formulated for conducting over-segmentation adaptively. To detect saliency effectively, a two-stage scoring method is adopted, in which both background prior and foreground cues are considered. In the first stage, we conduct random walk on absorbing Markov chain with background prior. And in the second stage, we use the saliency scores computed by the first stage scoring as foreground cues for manifold ranking. Experimental results on publicly available datasets demonstrate that our method outperforms the state-of-the-art methods in detecting salient objects. Qiang Cai 0001, Xiaobin Zhu 0001, Jian Cao 0003, Hai-Sheng Li 0002 |
ICIP | 3 |
| 2015 | Context-aware local abnormality detection in crowded scene
Xiaobin Zhu 0001, Xin Jin 0015, Xiaoyu Zhang 0002, Fugang He, Lei Wang 0101 |
Sci. China Inf. Sci. | 1 |
| 2015 | Update vs. upgrade: Modeling with indeterminate multi-class active learning
Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Xiao-chun Yun, Guangjun Wu, Yipeng Wang 0001 |
Neurocomputing | 3 |
| 2015 | Joint image representation and classification in random semantic spaces
Chunjie Zhang 0001, Xiaobin Zhu 0001, Liang Li 0003, Yifan Zhang 0001, Jing Liu 0001, Qingming Huang, Qi Tian 0001 |
Neurocomputing | 2 |
| 2015 | Human Age Estimation Based on Locality and Ordinal InformationabstractIn this paper, we propose a novel feature selection-based method for facial age estimation. The face aging is a typical temporal process, and facial images should have certain ordinal patterns in the aging feature space. From the geometrical perspective, a facial image can be usually seen as sampled from a low-dimensional manifold embedded in the original high-dimensional feature space. Thus, we first measure the energy of each feature in preserving the underlying local structure information and the ordinal information of the facial images, respectively, and then we intend to learn a low-dimensional aging representation that can maximally preserve both kinds of information. To further improve the performance, we try to eliminate the redundant local information and ordinal information as much as possible by minimizing nonlinear correlation and rank correlation among features. Finally, we formulate all these issues into a unified optimization problem, which is similar to linear discriminant analysis in format. Since it is expensive to collect the labeled facial aging images in practice, we extend the proposed supervised method to a semi-supervised learning mode including the semi-supervised feature selection method and the semi-supervised age prediction algorithm. Extensive experiments are conducted on the FACES dataset, the Images of Groups dataset, and the FG-NET aging dataset to show the power of the proposed algorithms, compared to the state-of-the-arts. Qingshan Liu 0001, Weishan Dong, Xiaobin Zhu 0001, Jing Liu 0001, Hanqing Lu |
IEEE Trans. Cybern. | 4 |
| 2014 | Key observation selection-based effective video synopsis for camera network
Xiaobin Zhu 0001, Jing Liu 0001, Jinqiao Wang, Hanqing Lu |
Mach. Vis. Appl. | 1 |
| 2014 | Sparse representation for robust abnormality detection in crowded scenes
Xiaobin Zhu 0001, Jing Liu 0001, Jinqiao Wang, Hanqing Lu |
Pattern Recognit. | 1 |
| 2012 | Weighted Interaction Force Estimation for Abnormality Detection in Crowd Scenes
Xiaobin Zhu 0001, Jing Liu 0001, Jinqiao Wang, Hanqing Lu |
ACCV (3) | 1 |