Xu-Cheng Yin

dblp:70/4187 · also Xucheng Yin · DBLP profile ↗
← Back
167ranked-venue papers
17as first author
102since 2021 · last 2026
0000-0003-0023-0220ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 102 · 15 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 74 · 2 first-author · 57 since 2021Databases, data management, data science and information retrieval · 25 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Dual-Geometry Graph Network: Unifying Local and Global Priors for Few-Shot Learning
abstract
In few-shot learning, utilizing local and global geometric priors to capture both subtle local class metrics and coarse global structures within the meta-task are important to obtain discriminative embeddings. However, existing graph-based and curvature-based few-shot approaches only focus on either one kind of geometric prior but neglect the other. To effectively utilize the pros of these two paradigms, we propose a novel Dual-Geometry Graph Network (DGGN) to adaptively integrate the local and global geometric priors via two key pathways. Specifically, the local-wise metric modeling pathway utilizes Ollivier-Ricci curvature to capture task-specific local class metrics among the instances, and the global-wise connectivity modeling pathway utilizes resistive embedding to capture global instance distributions and connectivity patterns of the entire meta-task. In addition, we introduce two new regularization loss functions to explicitly enhance the geometric representation ability of the local and global pathways respectively. We validate that DGGN's superior performance stems from its adaptively topological refinements by measuring the graph edit distance, demonstrating its ability to match the underlying data distribution. Extensive experiments show that DGGN sets a new state-of-the-art on standard, cross-domain, and semi-supervised few-shot benchmarks.
Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin
AAAI5
2026 VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
abstract
Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated significant potential in caption generation, existing benchmarks inadequately address fine-grained evaluation, particularly in capturing spatial-temporal details critical for video generation. To address this gap, we introduce the Fine-grained Video Caption Evaluation Benchmark (VCapsBench), the first large-scale fine-grained benchmark comprising 5,677 (5K+) videos and 109,796 (100K+) question-answer pairs. These QA-pairs are systematically annotated across 21 fine-grained dimensions (e.g., camera movement, and shot type) that are empirically proven critical for text-to-video generation. We further introduce three metrics (Accuracy (AR), Inconsistency Rate (IR), Coverage Rate (CR)), and an automated evaluation pipeline leveraging a large language model (LLM) to verify caption quality via contrastive QA-pairs analysis. Our benchmark can advance the development of robust text-to-video models by providing actionable insights for caption optimization.
Shi-Xue Zhang, Hongfa Wang, Duojun Huang, Xiaobin Zhu 0001, Xu-Cheng Yin
AAAI6
2026 Predicting Emergent Tool Use in LLMs Before It Emerges: A Proxy Perspective
abstract
Tool-use capabilities fundamentally transform large language models (LLMs) from passive language generators into active agents with real-world utility, drawing intense research focus. Yet, their emergent nature renders traditional scaling laws ineffective for early-stage prediction, obstructing principled model design and efficient training. In this work, we propose a proxy-task perspective that predicts tool-use capabilities by measuring early model performance on selected non-emergent proxy tasks. Our method quantifies two properties of each proxy task: alignment, which reflects how well it captures tool-use trajectories, and stability, which indicates how consistently it behaves across training conditions. These properties are used to weight predictive signals. Theoretically, we formalize how these weighted signals approximate emergent tool use through bounded extrapolation under relaxed assumptions. Empirically, we validate our approach across training checkpoints, model scales, and data setups. Results show that a carefully weighted ensemble of proxy tasks can accurately rank downstream tool-use ability long before it arises. Our findings provide new theoretical foundations and practical tools for efficient training and capability planning, and advance the understanding of how complex abilities arise in LLMs.
Bowen Zhang 0011, Yan Yan 0004, Guang Liu 0006, Xu-Cheng Yin
AAAI4
2026 Querying Historical $k$-Dense Subgraphs on Temporal Graphs
Yalong Zhang, Xu-Cheng Yin, Guoren Wang
ICDE4
2026 Beyond Optical Flow: Latent Micro-Motion as Visual Evidence in UAV Video
abstract
Optical flow has long dominated motion representation in video by focusing on explicit pixel displacement caused by object or camera movement. In this work, we argue that video also contains a largely overlooked form of motion information, namely latent micro-motion, which arises from subtle, structure-constrained responses of rigid components to physical interaction with the environment. We study this phenomenon in UAV video, a physically grounded setting where onboard structures are continuously exposed to aerodynamic forces. Although such micro-motions are low in amplitude and are often treated as noise or residual vibration, we show that they form a consistent visual signal that becomes observable through structure-aware and temporally aggregated analysis, even when using simple segmentation and coarse motion descriptors. Through an exploratory analysis, we demonstrate that micro-motion patterns exhibit clear structure and respond systematically to changes in wind conditions, with particularly strong sensitivity to wind direction and weaker dependence on wind magnitude in the examined scenarios. These observations suggest that micro-motion constitutes a distinct regime of motion information in video, complementary to explicit displacement, and motivate a broader reconsideration of how motion is represented and exploited in physically grounded multimedia scenarios.
Bowen Zhang 0011, Song-Lu Chen, Xiaobin Zhu 0001, Xu-Cheng Yin
ICMR6
2026 Local attention alignment fusion network for domain adaptive water body segmentation
Xiaobin Zhu 0001, Xu Qizhi, Yongjie Xia, Xu-Cheng Yin
Expert Syst. Appl.7
2026 Object re-identification with cross-view heterogeneous contrastive learning
Xu-Cheng Yin
J. Vis. Commun. Image Represent.2
2026 Robust scene text understanding with OCR token and word alignment for Text-VQA and text-caption
Zanxia Jin, Pinle Qin, Suzhen Lin, Shuangjiao Zhai, Jianchao Zeng 0001, Xu-Cheng Yin
Pattern Recognit.7
2026 Degradation decomposition learning for self-supervised blind image super-resolution
Xiaobin Zhu 0001, Liuling Chen, Jingyan Qin, Xu-Cheng Yin
Pattern Recognit.6
2026 Character recognition on continuous casting slabs via rotated detection and corner regression
Zhongjie Hu, Song-Lu Chen, Xiuxin Ge, Xu-Cheng Yin
Pattern Recognit. Lett.6
2026 Towards Efficient and Reliable Training Assurance of Untrusted Federated Learning Participants Under Hardware Non-Determinism
abstract
Federated learning (FL) is a popular privacy-preserving machine learning paradigm, enabling collaborative training across participants without exposing local data. Since FL loses direct control over participants' training executions, a fundamental requirement is to verify that participants faithfully perform the assigned training tasks. In this paper, we present TrustFL+, an efficient, scalable, and reliable verification scheme that ensures the training correctness of federated learning participants by leveraging both Trusted Execution Environments (TEEs) and GPUs. Essentially, it pushes all local training on high-performance but untrusted GPUs, while the TEE replicates the random parts for tunable levels of assurance. A key challenge is that hardware non-determinism can cause the same floating-point operations to yield different results between GPUs and TEEs, leading to false positives when participants behave honestly. TrustFL+ builds on deterministic training by recording rounding directions of intermediate operations during GPU-side model training and reusing them in TEE-based verification. It especially introduces adaptive rounding precisions to practically control non-determinism while maintaining global model performance in federated learning systems with lots of heterogeneous GPUs and iterative training. We prototype TrustFL+ using a range of NVIDIA GPUs covering multiple hardware architectures, along with Intel SGX, and evaluate its performance across convolutional neural networks and transformer-based networks. The experimental results demonstrate that TrustFL+ delivers up to an order of magnitude speedup compared to naive SGX-based training. Furthermore, all models trained with TrustFL+ on different GPU architectures successfully pass verification within SGX, resulting in 0 false positives.
Xiaoli Zhang 0003, Jiaqing Cheng, Wenmao Liu, Xiaohu Ye, Ke Xu 0002, Qi Li 0002, Xu-Cheng Yin
IEEE Trans. Dependable Secur. Comput.8
2026 Amplitude-Phase Reconstruction for Non-Stationary Time-Series Forecasting
abstract
Real-world time-series can be decomposed into multiple interacting frequency components whose amplitudes and phases co-evolve over time. Such coupled dynamics can deform spectral trajectories, leading to pronounced non-stationarity in frequency-domain representations and substantial degradation in long-horizon forecasting performance. However, most existing frequency-domain forecasting methods do not explicitly model amplitude–phase interactions and often implicitly assume globally consistent spectral structures, which limits their ability to capture drifting spectra under non-stationary dynamics. To address this challenge, we propose a novel Amplitude–Phase Reconstruction Network (APRNet), a frequency-domain time-series forecasting framework that models amplitude–phase interactions from temporal and channel-wise perspectives. Specifically, we propose an innovative Amplitude-Phase Global Correlation (APGC) module to capture spectrum-wide amplitude-phase dependencies and derive frequency-wise calibration factors under non-stationary dynamics. The calibrated spectral representations are further reconstructed into the time domain, forming a closed time-frequency reconstruction loop that suppresses spectral drift and mitigates non-stationarity in temporal features. In addition, we propose a novel Discrete Kolmogorov–Arnold Network (D-KAN) to further enhance fine-grained amplitude–phase modeling. By combining discretized nonlinear activations with locally supported B-spline basis functions, D-KAN enables frequency-adaptive piecewise nonlinear modeling, improving sensitivity to localized spectral variations and enhancing high-frequency expressiveness. Extensive experiments verify the superior performance of our APRNet. Our codes are available at:https://github.com/LH325/APRNet.
Xiaobin Zhu 0001, Lirui Deng 0001, Xu-Cheng Yin
IEEE Trans. Knowl. Data Eng.5
2025 AtomNet: Designing Tiny Models from Operators Under Extreme MCU Constraints
abstract
Tiny machine learning (TinyML) has attracted heightened attention for its ability to provide low-cost and instantaneous performance on edge devices. Particularly, the commonly used microcontroller unit (MCU) imposes extreme constraints on peak memory (SRAM) and storage (Flash). Existing TinyML methods often rely on a customized and hard-to-obtain inference libraries, as well as necessitate a time-consuming search for a deployable architecture using advanced Neural Architecture Search (NAS) algorithms. To solve these problems, we fully exploit the resources on MCU and deduce hardware-oriented guidelines for designing models under extreme MCU constraints. In detail, we delve into thorough information about the atom operators by collecting the runtime data of Flash, SRAM, and latency to build a dataset named AtomDB. Based on AtomDB, several critical operator guidelines are established to fully utilize limited Flash and SRAM, while minimizing latency. By transferring the guidelines to analyze blocks, we propose a hybrid pattern that organizes appropriate blocks at different network stages to form the AtomNet, a more hardware-oriented architecture, to handle the former SRAM bottleneck and the latter Flash bottleneck. Extensive experiments demonstrate the effectiveness of the exploitation of the hardware characteristics. Remarkably, AtomNet pioneeringly achieve 3.5% accuracy enhancement and more than 15% latency reduction on 320KB MCU using readily available official inference libraries for ImageNet tasks, surpassing the current state-of-the-art method.
Zhiwei Dong, Mingzhu Shen, Shihao Bai, Xiuying Wei, Jinyang Guo 0002, Ruihao Gong, Song-Lu Chen, Xianglong Liu 0001, Xu-Cheng Yin
AAAI9
2025 FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
abstract
Humans can perceive speakers’ characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech ( TTS ) scholars grounded their investigations on real-person faces, thereby restricting effective speech synthesis from applying to vast potential usage scenarios with diverse characters and image styles. To solve this issue, we introduce a novel FaceSpeak approach. It extracts salient identity characteristics and emotional representations from a wide variety of image styles. Meanwhile, it mitigates the extraneous information (e.g., background, clothing, and hair color, etc.), resulting in synthesized speech closely aligned with a character’s persona. Furthermore, to overcome the scarcity of multi-modal TTS data, we have devised an innovative dataset, namely Expressive Multi-Modal TTS ( EM2TTS), which is diligently curated and annotated to facilitate research in this domain. The experimental results demonstrate our proposed FaceSpeak can generate portrait-aligned voice with satisfactory naturalness and quality.
Tian-Hao Zhang, Xinyuan Qian 0001, Xu-Cheng Yin
AAAI5
2025 DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid Framework
abstract
Optical flow estimation is essential for video processing tasks, such as restoration and action recognition. The quality of videos is constantly increasing, with current standards reaching 8K resolution. However, optical flow methods are usually designed for low resolution and do not generalize to large inputs due to their rigid architectures. They adopt downscaling or input tiling to reduce the input size, causing a loss of details and global information. There is also a lack of optical flow benchmarks to judge the actual performance of existing methods on high-resolution samples. Previous works only conducted qualitative high-resolution evaluations on hand-picked samples. This paper fills this gap in optical flow estimation in two ways. We propose DPFlow, an adaptive optical flow architecture capable of generalizing up to 8K resolution inputs while trained with only low-resolution samples. We also introduce Kubric-NK, a new benchmark for evaluating optical flow methods with input resolutions ranging from 1K to 8K. Our high-resolution evaluation pushes the boundaries of existing methods and reveals new insights about their generalization capabilities. Extensive experimental results show that DPFlow achieves state-of-the-art results on the MPI-Sintel, KITTI 2015, Spring, and other high-resolution benchmarks. The code and dataset are available at https://github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/dpflow.
Henrique Morimitsu, Xiaobin Zhu 0001, Roberto Marcondes Cesar Junior, Xiangyang Ji, Xu-Cheng Yin
CVPR5
2025 Joint Feature and Kernel Fusion for Improved Depth-Aware Panoptic Segmentation
abstract
Depth-aware Panoptic Segmentation, which combines panoptic segmentation and monocular depth estimation, is a challenging task that requires a comprehensive understanding of both scene geometry and object semantics. Recent multi-task learning approaches have leveraged dynamic kernel methods to tackle these tasks simultaneously. However, these methods often treat feature extraction and kernel generation for each task in isolation, failing to fully exploit the rich interdependencies between depth and semantic information. To address this, we propose a novel framework with Cross-Task Feature Fusion and Kernel Fusion mechanisms to enhance Depth-aware Panoptic Segmentation. Our approach enables deeper integration of features and kernels, promoting more effective information exchange and mutual reinforcement between tasks. Experiments show that our method brings significant improvements, demonstrating the potential of a more integrated multi-task learning strategy for Depth-aware Panoptic Segmentation.
Shu Tian, Xin Zhao 0012, Xu-Cheng Yin
ICASSP4
2025 Tool Playgrounds: A Comprehensive and Analyzable Benchmark for LLM Tool Invocation
abstract
The rapid advancement of large language models (LLMs) has paved the way for their use in solving real-world problems, which in turn has significantly driven the development of tool-assisted LLMs. This progress necessitates thorough evaluation methods. However, existing benchmarks typically only provide end-to-end scores but lack in-depth analysis and often suffer from issues such as instability. To address this gap, we have meticulously designed the Tool Playgrounds framework, a comprehensive, analyzable, and extensible benchmark. This framework evaluates boundary dimensions such as parameter missing interaction, parameter correction, tool failover, and leveraging internal knowledge. Our findings indicate that even the most advanced commercial models frequently overlook these essential aspects and face challenges in managing complex tool usage. To foster further research and development, we have made our code, dataset, and leaderboard publicly available on https://github.com/zhiwei-dong/ToolPlaygrounds.
Zhiwei Dong, Ruihao Gong, Yang Yong, Yongqiang Yao, Song-Lu Chen, Xu-Cheng Yin
ICASSP7
2025 Data-Free Post-Training Quantization with Block-wise Enhanced Sample Generation
abstract
Data-free quantization is known for quantizing a pre-trained deep neural network without access to any training data, which applies to many real-world scenarios in that the training data is unavailable due to security, user privacy, or proprietary concerns. Most of the existing data-free quantization methods adopt a generator-quantization framework, which generator network to synthesize fake samples and Quantization-Aware Training (QAT) to quantize model. While the combination of the generator network and QAT can result in good accuracy for quantized models, the diversity of generated samples is lacking and quantizing a single model may take over 10 hours, which contrasts with Post-Training Quantization (PTQ)’s time-saving potential but poor accuracy. In order to address these issues, we have made improvements to the data generation and quantization process. In detail, 1) We propose Generator Exploration Enhancement (GEE) for utilizing the batch normalization statistics and adversarial sample exploration to enhance the quality and diversity of synthetic samples; 2) We introduce Block-wise Sample Generation (BSG) to collectively optimize individual blocks and the generator, leveraging PTQ as a foundation to boost workflow efficiency. Experiment results show that our proposed method improves both the performance and the efficiency of the data-free quantization compared to that of existing methods. Significantly, BSG achieves an 18% accuracy improvement and reduces quantization time by over 50% for 3-bit ResNet-18 in ImageNet tasks, surpassing the current state-of-the-art QAT method.
Ruiyao Zhang, Zhiwei Dong, Shutong Ti, Song-Lu Chen, Xu-Cheng Yin
ICASSP6
2025 Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
abstract
Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weighted Finite-State Transducer (WFST) is extensively used to integrate acoustic and language models, leveraging its capacity to implicitly fuse language models within static graphs, thereby ensuring robust recognition while also facilitating rapid error correction. However, WFST necessitates a frame-by-frame search of CTC posterior probabilities through autoregression, which significantly hampers inference speed. In this work, we thoroughly investigate the spike property of CTC outputs and further propose the conjecture that adjacent frames to non-blank spikes carry semantic information beneficial to the model. Building on this, we propose the Spike Window Decoding algorithm, which greatly improves the inference speed by making the number of frames decoded in WFST linearly related to the number of spiking frames in the CTC output, while guaranteeing the recognition performance. Our method achieves SOTA recognition accuracy with significantly accelerates decoding speed, proven across both AISHELL-1 and large-scale In-House datasets, establishing a pioneering approach for integrating CTC output with WFST.
Tian-Hao Zhang, Xinyuan Qian 0001, Xu-Cheng Yin
ICASSP7
2025 T-LLaVA: An Effective Saliency-Aware Slicing Strategy for Text Recognition
Mengze Wei, Xiaobin Zhu 0001, Xu-Cheng Yin
ICDAR (1)6
2025 SPAN: A Salient Patch-Clue Aware Network for Cross-Domain Face Anti-Spoofing
abstract
Face anti-spoofing (FAS) plays a critical role in ensuring the security of face recognition system from different kinds of presentation attacks. Most existing FAS research faces several limitations: 1) insufficient consideration of the role of local fine-grained information, 2) the assumption that spoofing patterns are uniformly distributed across the entire image, neglecting the uneven distribution of spoofing clues, and 3) an overemphasis on intra-domain scenarios, leading to limited generalization capabilities for unseen domains. In this paper, we propose a Salient Patch-Clue Aware Network (SPAN) for cross-domain face anti-spoofing to tackle the aforementioned issues. Specifically, we use all patches cropped from the complete image as input to the FAS network, enabling the network to focus on local information while avoiding information loss. Additionally, we propose a patch perception mechanism to extract key regions containing salient spoofing clues, such as reflections and edges, thereby reducing interference from irrelevant information. Furthermore, we introduce a pixel perception mechanism to capture finer-grained details. Based on these two mechanisms, we design a Salient Clue Perception Module (SCPM). We conduct cross-domain experiments on CASIA-FASD, Idiap Replay-Attack, MSU-MFSD, and OULU-NPU datasets. Our method achieves state-of-the-art HTER on seven protocols, especially excelling on M&I to C and M&I to O, surpassing the second place by 9.79% and 6.19%, showcasing strong generalization capability. The codes are available at https://github.com/SPAN2025/SPAN.
Liangfeng Zhang, Lei Chen 0069, Jinhui Lin, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
IJCNN8
2025 Can the Perceived Capability of Your Virtual Avatar Enhance Exercise Performance?
abstract
The rise of Virtual Reality (VR) sports has been driven by evolving work patterns, limited access to physical exercise spaces, and a growing focus on health and wellness. Beyond the enjoyment and convenience offered by VR exercise, we aim to enhance users' athletic performance through this medium. Prior research on the Proteus Effect in VR sports has demonstrated the potential of customized, stronger, or younger avatars to improve exercise outcomes. However, such effects are limited for users who already perceive themselves as strong and youthful. To address this, we propose a more general approach: representing perceived avatar capability through facial expressions and vocal cues to influence user performance. In this study, we examined how manipulating the perceived capability level of virtual avatars (high, neutral, low) during dumbbell lateral raises in a VR gym affected exercise performance. Results indicated that avatars with high-capability expressions significantly enhanced performance and motivation compared to low-capability avatars. These findings underscore the promise of using facial and auditory cues to represent perceived capability, offering new directions for designing emotionally intelligent fitness applications.
Sen-Zhe Xu 0001, Bo-Sheng Huang, Zian Zhou, Run-Yu Li, Song-Hai Zhang, Xu-Cheng Yin
ISMAR6
2025 Detecting and Adapting to Stealthy Label-Inversion Drifts via Conditional Distribution Inference
abstract
Deep learning (DL) based malicious traffic detectors have been widely developed to detect diverse network attacks, yet they are suffering from significant performance degradation due to concept drift. Existing anti-concept drift arts focus on combating the drifting traffic whose features significantly diverge from training traffic. However, they neglect a stealthy yet common situation where the testing traffic has similar features to the training traffic but with opposite ground truth labels. As a result, the DL-based detectors would always make incorrect predictions for the stealthy drifting traffic, insufficient to perform long-term real-world intrusion detection. In this paper, we propose Chameleon, a novel active learning framework that combats stealthy drifting traffic by inferring the conditional distribution of the testing traffic with small manual labeling overhead. Specifically, Chameleon measures the fine-grained correlations between the high-dimensional and heterogeneous testing traffic and selects a small number of highly representative testing traffic samples for manual labeling, to accurately infer other testing samples’ labels. With the inferred labels, Chameleon checks the conditional distribution shift from the training to testing traffic to detect concept drift and incrementally trains the DL-based detectors to make them effectively adapt to the shifted distribution. Extensive experiments with six supervised and unsupervised DL-based detectors on three public and four synthetic datasets show that, under stealthy drifting traffic, Chameleon improves the AUT of the DL-based detectors by a range of $18.53 \%$ to $23.89 \%$, while the improvement of SOTA baselines is only between $0.06 \%$ and $1.86 \%$.
Xiaoli Zhang 0003, Qilei Yin, Jianrong Zhang, Ke Xu 0002, Qi Li 0002, Xu-Cheng Yin
RAID9
2025 Aligning enhanced feature representation for generalized zero-shot learning
Zhiyu Fang, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin
Sci. China Inf. Sci.6
2025 Multi-Scale Texture Fusion for Reference-Based Image Super-Resolution: New Dataset and Solution
Xiaobin Zhu 0001, Jingyan Qin, Roberto Marcondes Cesar Junior, Xu-Cheng Yin
Int. J. Comput. Vis.6
2025 Special issue on advanced topics in document analysis (2025 ICDAR-IJDAR journal track)
Daniel P. Lopresti, Dimosthenis Karatzas, Xu-Cheng Yin
Int. J. Document Anal. Recognit.3
2025 Multi-view optimization and refinement for high-fidelity 4D Gaussian splatting
Jinhui Lin, Zhenyang Wei, Silei Shen, Xiaobin Zhu 0001, Xu-Cheng Yin
Neurocomputing6
2025 Decoupling and Interaction: task coordination in single-stage object detection
Jia-Wei Ma, Shu Tian, Haixia Man, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin
Multim. Tools Appl.6
2025 OSS-OCL: Occlusion Scenario Simulation and Occluded-edge Concentrated Learning for pedestrian detection
Keqi Lu, Chao Zhu 0003, Mengyin Liu, Xu-Cheng Yin
Pattern Recognit. Lett.4
2025 HA-FGOVD: Highlighting Fine-Grained Attributes via Explicit Linear Composition for Open-Vocabulary Object Detection
abstract
Open-vocabulary object detection (OVD) models are considered to be Large Multi-modal Models (LMM), due to their extensive training data and a large number of parameters. Mainstream OVD models prioritize object coarse-grained category rather than focus on their fine-grained attributes, e.g., colors or materials, thus failed to identify objects specified with certain attributes. Despite being pretrained on large-scale image-text pairs with rich attribute information, their latent feature space does not highlight these fine-grained attributes. In this paper, we introduce HA-FGOVD, a universal and explicit method that enhances the attribute-level detection capabilities of frozen OVD models by highlighting fine-grained attributes in explicit linear space. Our approach uses a LLM to extract attribute words in input text as a zero-shot task. Then, token attention masks are adjusted to guide text encoders in extracting both global and attribute-specific features, which are explicitly composited as two vectors in linear space to form a new attribute-highlighted feature for detection tasks. The composition weight scalars can be learned or transferred across different OVD models, showcasing the universality of our method. Experimental results show that HA-FGOVD achieves state-of-the-art performance on the FG-OVD benchmark and demonstrates promising generalization on the OVDEval benchmark, suggesting that our method addresses significant limitations in fine-grained attribute detection and has potential for broader fine-grained detection applications.
Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
IEEE Trans. Multim.4
2024 Recurrent Partial Kernel Network for Efficient Optical Flow Estimation
abstract
Optical flow estimation is a challenging task consisting of predicting per-pixel motion vectors between images. Recent methods have employed larger and more complex models to improve the estimation accuracy. However, this impacts the widespread adoption of optical flow methods and makes it harder to train more general models since the optical flow data is hard to obtain. This paper proposes a small and efficient model for optical flow estimation. We design a new spatial recurrent encoder that extracts discriminative features at a significantly reduced size. Unlike standard recurrent units, we utilize Partial Kernel Convolution (PKConv) layers to produce variable multi-scale features with a single shared block. We also design efficient Separable Large Kernels (SLK) to capture large context information with low computational cost. Experiments on public benchmarks show that we achieve state-of-the-art generalization performance while requiring significantly fewer parameters and memory than competing methods. Our model ranks first in the Spring benchmark without finetuning, improving the results by over 10% while requiring an order of magnitude fewer FLOPs and over four times less memory than the following published method without finetuning. The code is available at github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/rpknet.
Henrique Morimitsu, Xiaobin Zhu 0001, Xiangyang Ji, Xu-Cheng Yin
AAAI4
2024 Arbitrary Time Information Modeling via Polynomial Approximation for Temporal Knowledge Graph Embedding
abstract
Distinguished from traditional knowledge graphs (KGs), temporal knowledge graphs (TKGs) must explore and reason over temporally evolving facts adequately. However, existing TKG approaches still face two main challenges, i.e., the limited capability to model arbitrary timestamps continuously and the lack of rich inference patterns under temporal constraints. In this paper, we propose an innovative TKGE method (PTBox) via polynomial decomposition-based temporal representation and box embedding-based entity representation to tackle the above-mentioned problems. Specifically, we decompose time information by polynomials and then enhance the model’s capability to represent arbitrary timestamps flexibly by incorporating the learnable temporal basis tensor. In addition, we model every entity as a hyperrectangle box and define each relation as a transformation on the head and tail entity boxes. The entity boxes can capture complex geometric structures and learn robust representations, improving the model’s inductive capability for rich inference patterns. Theoretically, our PTBox can encode arbitrary time information or even unseen timestamps while capturing rich inference patterns and higher-arity relations of the knowledge base. Extensive experiments on real-world datasets demonstrate the effectiveness of our method.
Zhiyu Fang, Jingyan Qin, Xiaobin Zhu 0001, Xu-Cheng Yin
LREC/COLING5
2024 LayoutFormer: Hierarchical Text Detection Towards Scene Text Understanding
abstract
Existing scene text detectors generally focus on accu-rately detecting single-level (i.e., word-level, line-level, or paragraph-level) text entities without exploring the relationships among different levels of text entities. To comprehensively understand scene texts, detecting multi-level texts while exploring their contextual information is criti-cal. To this end, we propose a unified framework (dubbed LayoutFormer) for hierarchical text detection, which simultaneously conducts multi-level text detection and predicts the geometric layouts for promoting scene text understanding. In LayoutFormer, WordDecoder, LineDecoder, and Pa- raDecoder are proposed to be responsible for word-level text prediction, line-level text prediction, and paragraph- level text prediction, respectively. Meanwhile, WordDe-coder and ParaDecoder adaptively learn word-line and line-paragraph relationships, respectively. In addition, we propose a Prior Location Sampler to be used on multi-scale features to adaptively select a few representative foreground features for updating text queries. It can improve hierar- chical detection performance while significantly reducing the computational cost. Comprehensive experiments verify that our method achieves state-of-the-art performance on single-level and hierarchical text detection.
Jia-Wei Ma, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin
CVPR5
2024 HardMo: A Large-Scale Hardcase Dataset for Motion Capture
abstract
Recent years have witnessed rapid progress in monoc-ular human mesh recovery. Despite their impressive performance on public benchmarks, existing methods are vulnerable to unusual poses, which prevents them from deploying to challenging scenarios such as dance and martial arts. This issue is mainly attributed to the domain gap induced by the data scarcity in relevant cases. Most existing datasets are captured in constrained scenarios and lack samples of such complex movements. For this reason, we propose a data collection pipeline comprising automatic crawling, precise annotation, and hardcase mining. Based on this pipeline, we establish a large dataset in a short time. The dataset, named HardMo, contains 7M images along with precise annotations covering 15 categories of dance and 14 categories of martial arts. Empirically, we find that the prediction failure in dance and martial arts is mainly characterized by the misalignment of hand-wrist and foot-ankle. To dig deeper into the two hardcases, we leverage the proposed automatic pipeline to filter collected data and construct two subsets named HardMo-Hand and HardMo-Foot. Extensive experiments demonstrate the effectiveness of the annotation pipeline and the data-driven solution to failure cases. Specifically, after being trained on HardMo, HMR, an early pioneering method, can even outperform the current state of the art, 4DHumans, on our benchmarks. Dataset will be publicly available at https://ljqnb.github.io/HardMo.github.io.
Jiaqi Liao, Chuanchen Luo, Yinuo Du, Yuxi Wang 0001, Xu-Cheng Yin, Man Zhang 0005, Zhaoxiang Zhang 0001, Junran Peng
CVPR5
2024 Attention Decoupling for Query-Based Object Detection
abstract
Benefiting from attention mechanisms, query-based detectors have a strong model capacity. They predict classification and regression by utilizing their shared queries and features in the decoder. Inter-task biases cause multi-directional gradients that disturb each other to limit model optimization. In this work, we introduce an attention decoupling (AD) for query-based detectors to explicitly align multi-task features. Specifically, AD consists of a Dense-to-Sparse Query Generator (DSQG) and a Split Cross-Attention (SCA), enabling query and feature decoupling respectively in decoding phase. Then, we propose a task consistency loss (TCL) which integrates a novel task alignment metric to classification loss to further improve task consistency across multiple decoding stages. Thus, AD effectively mitigates query-based detectors’ task misalignment problem and inspires subsequent multi-task paradigms. Moreover, extensive experiments on COCO dataset demonstrate that the proposed AD can enhance a variety of representative detectors. Remarkably, AD-DINO achieves the state-of-the-art performance.
Jia-Wei Ma, Haixia Man, Shu Tian, Jingyan Qin, Xu-Cheng Yin
ICASSP6
2024 Multi-task Learning for License Plate Recognition in Unconstrained Scenarios
Zhen-Lun Mo, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
ICDAR (1)5
2024 HQOD: Harmonious Quantization for Object Detection
abstract
Task inharmony problem commonly occurs in modern object detectors, leading to inconsistent qualities between classification and regression tasks. The predicted boxes with high classification scores but poor localization positions or low classification scores but accurate localization positions will worsen the performance of detectors after Non-Maximum Suppression. Furthermore, when object detectors collaborate with Quantization- Aware Training (QAT), we observe that the task inharmony problem will be further exacerbated, which is considered one of the main causes of the performance degradation of quantized detectors. To tackle this issue, we propose the Harmonious Quantization for Object Detection (HQOD) framework, which consists of two components. Firstly, we propose a task-correlated loss to encourage detectors to focus on improving samples with lower task harmony quality during QAT. Secondly, a harmonious Intersection over Union (IoU) loss is incorporated to balance the optimization of the regression branch across different IoU levels. The proposed HQOD can be easily integrated into different QAT algorithms and detectors. Remarkably, on the MS COCO dataset, our 4-bit ATSS with ResNet-50 backbone achieves a state-of-the- art mAP of 39.6%, even surpassing the full-precision one. Codes are available at https://github.com/Menace-Dragon/VP-QOD.
Zhiwei Dong, Song-Lu Chen, Ruiyao Zhang, Shutong Ti, Feng Chen 0040, Xu-Cheng Yin
ICME7
2024 Towards Low-resource License Plate Recognition via Feature Shuffling
abstract
Manual annotation is costly and limits the availability of sufficient annotated license plates for training recognition models. Small-scale license plate datasets (i.e., low-resource) often exhibit a long-tailed distribution in character classes at some character positions, primarily due to their limited variation in character permutations. Previous methods tend to prioritize head classes with high occurrence probability when applied to small-scale datasets. To solve this problem, we propose feature shuffling to balance the occurrence distribution across various character classes, thereby improving the recognition of tail classes with low occurrence probability. Moreover, we introduce global perception to holistically understand the overall character layout for effective feature shuffling. Extensive experiments on the small-scale UFPR and SSIG-SegPlate datasets demonstrate that our method achieves state-of-the-art results, with an average improvement of 43.70% over the baseline. Experiments on RodoSol and CCPD prove our method achieves state-of-the-art performance on large-scale datasets, verifying its generality.
Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
ICME5
2024 RAPIDFlow: Recurrent Adaptable Pyramids with Iterative Decoding for Efficient Optical Flow Estimation
abstract
Extracting motion information from videos with optical flow estimation is vital in multiple practical robot applications. Current optical flow approaches show remarkable accuracy, but top-performing methods have high computational costs and are unsuitable for embedded devices. Although some previous works have focused on developing low-cost optical flow strategies, their estimation quality has a noticeable gap with more robust methods. In this paper, we develop a novel method to efficiently estimate high-quality optical flow in embedded devices. Our proposed RAPIDFlow model combines efficient NeXt1D convolution blocks with a fully recurrent structure based on feature pyramids to decrease computational costs without significantly impacting estimation accuracy. The adaptable recurrent encoder produces multi-scale features with a single shared block, which allows us to adjust the pyramid length at inference time and make it more robust to changes in input size. Also, it enables our model to offer multiple tradeoffs between accuracy and speed to suit different applications. Experiments using a Jetson Orin NX embedded system on the MPI-Sintel and KITTI public benchmarks show that RAPIDFlow outperforms previous approaches by significant margins at faster speeds. Our code is available at https://github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/rapidflow.
Henrique Morimitsu, Xiaobin Zhu 0001, Roberto Marcondes Cesar Junior, Xiangyang Ji, Xu-Cheng Yin
ICRA5
2024 Transmitted and Aggregated Self-Attention for Automatic Speech Recognition
Tian-Hao Zhang, Xinyuan Qian 0001, Feng Chen 0040, Xu-Cheng Yin
INTERSPEECH4
2024 Exploring Stable Meta-Optimization Patterns via Differentiable Reinforcement Learning for Few-Shot Classification
abstract
Existing few-shot learning methods generally focus on designing exquisite structures of meta-learners for learning task-specific prior to improve the discriminative ability of global embeddings. However, they often ignore the importance of learning stability in meta-training, making it difficult to obtain a relatively optimal model. From this key observation, we propose an innovative generic differentiable Reinforcement Learning (RL) strategy for few-shot classification. It aims to explore stable meta-optimization patterns in meta-training by learning generalizable optimizations for producing task-adaptive embeddings. Accordingly, our differentiable RL strategy models the embedding procedure of feature transformation layers in meta-learner to optimize the gradient flow implicitly. Also, we propose a memory module to associate historical and current task states and actions for exploring inter-task similarity. Notably, our RL-based strategy can be easily extended to various backbones. In addition, we propose a novel task state encoder to encode task representation, which fully explores inner-task similarities between support set and query set. Extensive experiments verify that our approach can improve the performance of different backbones and achieve promising results against state-of-the-art methods in few-shot classification.
Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin
ACM Multimedia6
2024 MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets
abstract
Driven by powerful image diffusion models, recent research has achieved the automatic creation of 3D objects from textual or visual guidance. By performing score distillation sampling (SDS) iteratively across different views, these methods succeed in lifting 2D generative prior to the 3D space. However, such a 2D generative image prior bakes the effect of illumination and shadow into the texture. As a result, material maps optimized by SDS inevitably involve spurious correlated components. The absence of precise material definition makes it infeasible to relight the generated assets reasonably in novel scenes, which limits their application in downstream scenarios. In contrast, humans can effortlessly circumvent this ambiguity by deducing the material of the object from its appearance and semantics. Motivated by this insight, we propose MaterialSeg3D, a 3D asset material generation framework to infer underlying material from the 2D semantic prior. Based on such a prior model, we devise a mechanism to parse material in 3D space. We maintain a UV stack, each map of which is unprojected from a specific viewpoint. After traversing all viewpoints, we fuse the stack through a weighted voting scheme and then employ region unification to ensure the coherence of the object parts. To fuel the learning of semantics prior, we collect a material dataset, named Materialized Individual Objects (MIO), which features abundant images, diverse categories, and accurate annotations. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our method.
Ruitong Gan, Chuanchen Luo, Yuxi Wang 0001, Qing Li 0001, Xu-Cheng Yin, Man Zhang 0005, Zhaoxiang Zhang 0001, Junran Peng
ACM Multimedia8
2024 Unsupervised Multi-view Pedestrian Detection
abstract
With the prosperity of the intelligent surveillance, multiple cameras have been applied to localize pedestrians more accurately. However, previous methods rely on laborious annotations of pedestrians in every frame and camera view. Therefore, we propose in this paper an Unsupervised Multi-view Pedestrian Detection approach (UMPD) to learn an annotation-free detector via vision-language models and 2D-3D cross-modal mapping: 1) Firstly, Semantic-aware Iterative Segmentation (SIS) is proposed to extract unsupervised representations of multi-view images, which are converted into 2D masks as pseudo labels, via our proposed iterative PCA and zero-shot semantic classes from vision-language models; 2) Secondly, we propose Geometry-aware Volume-based Detector (GVD) to end-to-end encode multi-view 2D images into a 3D volume to predict voxel-wise density and color via 2D-to-3D geometric projection, trained by 3D-to-2D rendering losses with SIS pseudo labels; 3) Thirdly, for better detection results, i.e., the 3D density projected on Birds-Eye-View, we propose Vertical-aware BEV Regularization (VBR) to constrain pedestrians to be vertical like the natural poses. Extensive experiments on popular multi-view pedestrian detection benchmarks Wildtrack, Terrace, and MultiviewX, show that our proposed UMPD, as the first fully-unsupervised method to our best knowledge, performs competitively to the previous state-of-the-art supervised methods. Code is available at https://github.com/lmy98129/UMPD.
Mengyin Liu, Chao Zhu 0003, Shiqi Ren, Xu-Cheng Yin
ACM Multimedia4
2024 Improving Small License Plate Detection with Bidirectional Vehicle-Plate Relation
Songkang Dai, Song-Lu Chen, Qi Liu 0041, Chao Zhu 0003, Feng Chen 0040, Xu-Cheng Yin
MMM (2)7
2024 Irregular License Plate Recognition via Global Information Integration
Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
MMM (2)5
2024 Integrated Recognition of Arbitrary-Oriented Multi-line Billet Number
Zhongjie Hu, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
PRCV (7)6
2024 Transformer-based Reasoning for Learning Evolutionary Chain of Events on Temporal Knowledge Graph
abstract
Temporal Knowledge Graph (TKG) reasoning often involves completing missing factual elements along the timeline. Although existing methods can learn good embeddings for each factual element in quadruples by integrating temporal information, they often fail to infer the evolution of temporal facts. This is mainly because of (1) insufficiently exploring the internal structure and semantic relationships within individual quadruples and (2) inadequately learning a unified representation of the contextual and temporal correlations among different quadruples. To overcome these limitations, we propose a novel Transformer-based reasoning model (dubbed ECEformer) for TKG to learn the Evolutionary Chain of Events (ECE). Specifically, we unfold the neighborhood subgraph of an entity node in chronological order, forming an evolutionary chain of events as the input for our model. Subsequently, we utilize a Transformer encoder to learn the embeddings of intra-quadruples for ECE. We then craft a mixed-context reasoning module based on the multi-layer perceptron (MLP) to learn the unified representations of inter-quadruples for ECE while accomplishing temporal knowledge reasoning. In addition, to enhance the timeliness of the events, we devise an additional time prediction task to complete effective temporal information within the learned unified representation. Extensive experiments on six benchmark datasets verify the state-of-the-art performance and the effectiveness of our method.
Zhiyu Fang, Shuai-Long Lei, Xiaobin Zhu 0001, Shi-Xue Zhang, Xu-Cheng Yin, Jingyan Qin
SIGIR6
2024 OCRBench: on the hidden mystery of OCR in large multimodal models
Mingxin Huang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Xiang Bai
Sci. China Inf. Sci.7
2024 Semi-supervised domain adaptation via subspace exploration
abstract
Abstract Recent methods of learning latent representations in Domain Adaptation (DA) often entangle the learning of features and exploration of latent space into a unified process. However, these methods can cause a false alignment problem and do not generalise well to the alignment of distributions with large discrepancy. In this study, the authors propose to explore a robust subspace for Semi‐Supervised Domain Adaptation (SSDA) explicitly. To be concrete, for disentangling the intricate relationship between feature learning and subspace exploration, the authors iterate and optimise them in two steps: in the first step, the authors aim to learn well‐clustered latent representations by aggregating the target feature around the estimated class‐wise prototypes; in the second step, the authors adaptively explore a subspace of an autoencoder for robust SSDA. Specially, a novel denoising strategy via class‐agnostic disturbance to improve the discriminative ability of subspace is adopted. Extensive experiments on publicly available datasets verify the promising and competitive performance of our approach against state‐of‐the‐art methods.
Xiaobin Zhu 0001, Zhiyu Fang, Jingyan Qin, Xu-Cheng Yin
IET Comput. Vis.6
2024 Improving license plate recognition via diverse stylistic plate generation
Qi Liu 0041, Song-Lu Chen, Yu-Xiang Chen, Xu-Cheng Yin
Pattern Recognit. Lett.4
2024 M3TTS: Multi-modal text-to-speech of multi-scale style control for dubbing
Li-Fang Wei, Xinyuan Qian 0001, Tian-Hao Zhang, Song-Lu Chen, Xu-Cheng Yin
Pattern Recognit. Lett.6
2024 CFOR: Character-First Open-Set Text Recognition via Context-Free Learning
abstract
The open-set text recognition task is a generalized form of the (close-set) text recognition task, where the model is further challenged to spot and incrementally recognize novel characters not covered by the training data. Novel characters also indicate that the language model of the training set is biased from the "real-world". In this work, we alleviate the confounding effect of such biases by learning from individual character representations isolated from their context. Specifically, we propose a Character-First Open-Set Text Recognition framework that cotrains the feature extractor with two context-free learning tasks. First, a Context Isolation Learning task is proposed to wipe the context for each character from the input image, utilizing a character mask learned in a weak supervision manner. Second, the framework adopts an Individual Character Learning task, which is a single-character classification task with synthetic samples. After training on English and simplified Chinese data, our framework can adapt to recognize unseen characters in Japanese, Korean, Greek, and other scripts without retraining, and can reliably spot unseen characters in Japanese with an F1-score over 64%. The framework also shows 91.5% line accuracy on IIIT5k and a speed of over 69 FPS single-batched, making it a feasible universal lightweight OCR solution that works well for both open-set and close-set use cases.
Chang Liu 0083, Zhiyu Fang, Haibo Qin, Xu-Cheng Yin
IEEE Trans. Image Process.5
2024 Inverse-Like Antagonistic Scene Text Spotting via Reading-Order Estimation and Dynamic Sampling
abstract
Scene text spotting is a challenging task, especially for inverse-like scene text, which has complex layouts,e.g., mirrored, symmetrical, or retro-flexed. In this paper, we propose a unified end-to-end trainable inverse-like antagonistic text spotting framework dubbed IATS, which can effectively spot inverse-like scene texts without sacrificing general ones. Specifically, we propose an innovative reading-order estimation module (REM) that extracts reading-order information from the initial text boundary generated by an initial boundary module (IBM). To optimize and train REM, we propose a joint reading-order estimation loss (LRE) consisting of a classification loss, an orthogonality loss, and a distribution loss. With the help of IBM, we can divide the initial text boundary into two symmetric control points and iteratively refine the new text boundary using a lightweight boundary refinement module (BRM) for adapting to various shapes and scales. To alleviate the incompatibility between text detection and recognition, we propose a dynamic sampling module (DSM) with a thin-plate spline that can dynamically sample appropriate features for recognition in the detected text region. Without extra supervision, the DSM can proactively learn to sample appropriate features for text recognition through the gradient returned by the recognition module. Extensive experiments on both challenging scene text and inverse-like scene text datasets demonstrate that our method achieves superior performance both on irregular and inverse-like text spotting.
Shi-Xue Zhang, Xiaobin Zhu 0001, Hongfa Wang, Xu-Cheng Yin
IEEE Trans. Image Process.6
2024 Improving Multi-Type License Plate Recognition via Learning Globally and Contrastively
abstract
Previous license plate recognition (LPR) methods have achieved impressive performance on single-type license plates. However, multi-type license plate recognition is still challenging due to various character layouts and fonts. There are two main problems: one is that recognition models are prone to incorrectly perceive the location of characters due to diverse character layouts, and the other is that characters of different categories may have similar glyphs due to various fonts, causing character misidentification. Therefore, to solve the above problems, we propose two plug-and-play modules based on an attention-based framework for multi-type license plate recognition. First, we propose a global modeling module to integrate character layout information to precisely perceive the location of characters, thus generating accurate predictions. Second, a position-aware contrastive learning module is proposed to enhance the robustness and discriminability of features to alleviate character misidentification of similar glyphs. Finally, to verify the effectiveness and generality, we apply the proposed modules to six baseline models, and the results demonstrate that the proposed method can achieve state-of-the-art performance on three multi-type license plate datasets. Moreover, extensive experiments prove that our proposed modules can significantly improve performance by 6.8% on RODOSOL-ALPR with a small parameter increase.
Qi Liu 0041, Song-Lu Chen, Tian-Hao Zhang, Feng Chen 0040, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.6
2024 Sample Weighting with Hierarchical Equalization Loss for Dense Object Detection
abstract
Label assignment (LA) is one of the essential phases in the object detection paradigm and aims to classify samples as foreground or background. Current LA strategies generally discriminate samples by explicit thresholds and then calculate weighted losses based on their significances. However, existing methods mostly neglect to consider the importance of samples comprehensively due to the uneven distribution of objects and the limitations of detector structures. In this paper, we propose a hierarchical equalization loss (HEL) by reconsidering the underlying factors affecting sample weights. First, we mitigate sample imbalance at three progressive levels. (1) Task level. We propose task-reconciled weights (TRW) to overcome the effects caused by inter-task inconsistencies (i.e., the inherent differences of classification and localization). (2) Instance level. We propose instance-aware normalization (IAN) for reconstructing the distribution of sample weights within an instance to suppress environmental noise. (3) Pyramid level. We propose hierarchical modulation (HM) to alleviate the unbalanced distribution of multi-scale objects on feature pyramids. Then, we stack the above three mechanisms and formulate the effective weighted loss. Moreover, we propose a staggered candidate bag construction (SCBC) mechanism to further improve the robustness of our method. Without adding any extra overhead, HEL can improve the performance of representative detectors by an impressive margin. Equipped with HEL, a single “ResNet-50+FPN+Head” detector can achieve a performance of 41.9 AP on COCO under 1× schedule, outperforming other existing LA methods. Extensive experiments conducted on multiple backbones and datasets demonstrate the effectiveness of our method.
Jia-Wei Ma, Lei Chen 0069, Shu Tian, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin
IEEE Trans. Multim.7
2024 Arbitrary Shape Text Detection via Boundary Transformer
abstract
In arbitrary shape text detection, locating accurate text boundaries is challenging and non-trivial. Existing methods often suffer from indirect text boundary modeling or complex post-processing. In this article, we systematically present a unified coarse-to-fine framework via boundary learning for arbitrary shape text detection, which can accurately and efficiently locate text boundaries without post-processing. In our method, we explicitly model the text boundary via an innovative iterative boundary transformer in a coarse-to-fine manner. In this way, our method can directly gain accurate text boundaries and abandon complex post-processing to improve efficiency. Specifically, our method mainly consists of a feature extraction backbone, a boundary proposal module, and an iteratively optimized boundary transformer module. The boundary proposal module consisting of multi-layer dilated convolutions will predict important prior information (including classification map, distance field, and direction field) for generating coarse boundary proposals while guiding the boundary transformer's optimization. The boundary transformer module adopts an encoder-decoder structure, in which the encoder is constructed by multi-layer transformer blocks with residual connection while the decoder is a simple multi-layer perceptron network (MLP). Under the guidance of prior information, the boundary transformer module will gradually refine the coarse boundary proposals via iterative boundary deformation. Furthermore, we propose a novel boundary energy loss (BEL) that introduces an energy minimization constraint and an energy monotonically decreasing constraint to further optimize and stabilize the learning of boundary refinement. Extensive experiments on publicly available and challenging datasets demonstrate the state-of-the-art performance and promising efficiency of our method.
Shi-Xue Zhang, Xiaobin Zhu 0001, Xu-Cheng Yin
IEEE Trans. Multim.4
2023 VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-Supervision
abstract
Detecting pedestrians accurately in urban scenes is significant for realistic applications like autonomous driving or video surveillance. However, confusing human-like objects often lead to wrong detections, and small scale or heavily occluded pedestrians are easily missed due to their unusual appearances. To address these challenges, only object regions are inadequate, thus how to fully utilize more explicit and semantic contexts becomes a key problem. Meanwhile, previous context-aware pedestrian detectors either only learn latent contexts with visual clues, or need laborious annotations to obtain explicit and semantic contexts. Therefore, we propose in this paper a novel approach via Vision-Language semantic self-supervision for context-aware Pedestrian Detection (VLPD) to model explicitly semantic contexts without any extra annotations. Firstly, we propose a self-supervised Vision-Language Semantic (VLS) segmentation method, which learns both fully-supervised pedestrian detection and contextual segmentation via self-generated explicit labels of semantic classes by vision-language models. Furthermore, a self-supervised Prototypical Semantic Contrastive (PSC) learning method is proposed to better discriminate pedestrians and other classes, based on more explicit and semantic contexts obtained from VLS. Extensive experiments on popular benchmarks show that our proposed VLPD achieves superior performances over the previous state-of-the-arts, particularly under challenging circumstances like small scale and heavy occlusion. Code is available at https://github.com/lmy98129/VLPD.
Mengyin Liu, Jie Jiang 0015, Chao Zhu 0003, Xu-Cheng Yin
CVPR4
2023 Self-Convolution for Automatic Speech Recognition
abstract
Self-attention plays a significant role in recent automatic speech recognition (ASR) models with promising results. However, it suffers from high computational complexity and weak capability in modeling local information. In contrast, the convolutional neural network (CNN) is computationally effective and superior in learning local information. Whereas it fails in self-interaction and capturing long-range dependence among input tokens. Accordingly, we take their complementary advantages and propose a new module, namely self-convolution, to compensate for each individual limitations. Specifically, self-convolution generates convolution kernels at each token (to model local information) which are then used to convolve itself (for self-interaction). Moreover, we bring in global information during the generation of convolution kernel to enhance the learning of long-range dependencies. In this way, the advantages of self-attention and CNN are both utilized. We conduct rigorous experiments on LibriSpeech, Tedlium2, and AIShell1 datasets and demonstrate that our proposed self-convolution can achieve superior ASR performance than self-attention with less computational cost.
Qi Liu 0041, Xinyuan Qian 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICASSP6
2023 Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-Resolution
abstract
Although existing image deep learning super-resolution (SR) methods achieve promising performance on benchmark datasets, they still suffer from severe performance drops when the degradation of the low-resolution (LR) input is not covered in training. To address the problem, we propose an innovative unsupervised method of Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-Resolution. Highly inspired by the generalized sampling theory, our method aims to enhance the strength of off-the-shelf SR methods trained on known degradations and adapt to unknown complex degradations to generate improved results. Specifically, we first conduct degradation estimation for each local image region by learning the internal distribution in an unsupervised manner via GAN. Instead of assuming degradation are spatially invariant across the whole image, we learn correction filters to adjust degradations to known degradations in a spatially variant way by a novel linearly-assembled pixel degradation-adaptive regression module (DARM). DARM is lightweight and easy to optimize on a dictionary of multiple pre-defined filter bases. Extensive experiments on synthetic and real-world datasets verify the effectiveness of our method both qualitatively and quantitatively. Code can be available at: https://github.com/edbca/DARSR.
Xiaobin Zhu 0001, Jianqing Zhu, Shi-Xue Zhang, Jingyan Qin, Xu-Cheng Yin
ICCV7
2023 End-to-End Multi-line License Plate Recognition with Cascaded Perception
Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
ICDAR (5)4
2023 Open-Set Text Recognition via Shape-Awareness Visual Reconstruction
Chang Liu 0083, Xu-Cheng Yin
ICDAR (6)3
2023 Complex Glyph Enhancement for License Plate Generation
Yu-Xiang Chen, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICIG (1)7
2023 Towards Discriminative Semantic Relationship for Fine-grained Crowd Counting
abstract
As an extended task of crowd counting, fine-grained crowd counting aims to estimate the number of people in each semantic category instead of the whole in an image, and faces challenges including 1) inter-category crowd appearance similarity, 2) intra-category crowd appearance variations, and 3) frequent scene changes. In this paper, we propose a new fine-grained crowd counting approach named DSR to tackle these challenges by modeling Discriminative Semantic Relationship, which consists of two key components: Word Vector Module (WVM) and Adaptive Kernel Module (AKM). The WVM introduces more explicit semantic relationship information to better distinguish people of different semantic groups with similar appearance. The AKM dynamically adjusts kernel weights according to the features from different crowd appearance and scenes. The proposed DSR achieves superior results over state-of-the-art on the standard dataset. Our approach can serve as a new solid baseline and facilitate future research for the task of fine-grained crowd counting.
Shiqi Ren, Chao Zhu 0003, Mengyin Liu, Xu-Cheng Yin
ICME4
2023 InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition
Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu 0041, Xinyuan Qian 0001, Li-Fang Wei, Feng Chen 0040, Song-Lu Chen, Xu-Cheng Yin
INTERSPEECH8
2023 Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding
Tian-Hao Zhang, Haibo Qin, Zhi-Hao Lai, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xinyuan Qian 0001, Xu-Cheng Yin
INTERSPEECH8
2023 LiteHandNet: A Lightweight Hand Pose Estimation Network via Structural Feature Enhancement
Zhi-Yong Huang, Song-Lu Chen, Qi Liu 0041, Chong-Jian Zhang, Feng Chen 0040, Xu-Cheng Yin
MMM (1)6
2023 Feature Enhancement and Reconstruction for Small Object Detection
Chong-Jian Zhang, Song-Lu Chen, Qi Liu 0041, Zhi-Yong Huang, Feng Chen 0040, Xu-Cheng Yin
MMM (1)6
2023 Feature Implicit Enhancement via Super-Resolution for Small Object Detection
Zhehao Xu, Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
PRCV (12)5
2023 Graph fusion network for multi-oriented object detection
Shi-Xue Zhang, Xiaobin Zhu 0001, Jie-Bo Hou, Xu-Cheng Yin
Appl. Intell.4
2023 Arbitrary Shape Text Detection via Segmentation With Probability Maps
abstract
Arbitrary shape text detection is a challenging task due to the significantly varied sizes and aspect ratios, arbitrary orientations or shapes, inaccurate annotations, etc. Due to the scalability of pixel-level prediction, segmentation-based methods can adapt to various shape texts and hence attracted considerable attention recently. However, accurate pixel-level annotations of texts are formidable, and the existing datasets for scene text detection only provide coarse-grained boundary annotations. Consequently, numerous misclassified text pixels or background pixels inside annotations always exist, degrading the performance of segmentation-based text detection methods. Generally speaking, whether a pixel belongs to text or not is highly related to the distance with the adjacent annotation boundary. With this observation, in this paper, we propose an innovative and robust segmentation-based detection method via probability maps for accurately detecting text instances. To be concrete, we adopt a Sigmoid Alpha Function (SAF) to transfer the distances between boundaries and their inside pixels to a probability map. However, one probability map can not cover complex probability distributions well because of the uncertainty of coarse-grained text boundary annotations. Therefore, we adopt a group of probability maps computed by a series of Sigmoid Alpha Functions to describe the possible probability distributions. In addition, we propose an iterative model to learn to predict and assimilate probability maps for providing enough information to reconstruct text instances. Finally, simple region growth algorithms are adopted to aggregate probability maps to complete text instances. Experimental results demonstrate that our method achieves state-of-the-art performance in terms of detection accuracy on several benchmarks. Notably, our method with Watershed Algorithm as post-processing achieves the best F-measure on Total-Text (88.79%), CTW1500 (85.75%), and MSRA-TD500 (88.93%). Besides, our method achieves promising performance on multi-oriented datasets (ICDAR2015) and multilingual datasets (ICDAR2017-MLT). Code is available at: https://github.com/GXYM/TextPMs.
Shi-Xue Zhang, Xiaobin Zhu 0001, Lei Chen 0069, Jie-Bo Hou, Xu-Cheng Yin
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Towards open-set text recognition via label-to-prototype learning
Chang Liu 0083, Haibo Qin, Xiaobin Zhu 0001, Cheng-Lin Liu 0001, Xu-Cheng Yin
Pattern Recognit.6
2023 Beyond OCR + VQA: Towards end-to-end reading and reasoning for robust and accurate textvqa
abstract
Text-based visual question answering (TextVQA), which answers a visual question by considering both visual contents and scene texts, has attracted increasing attention recently. Most existing methods employ an optical character recognition (OCR) module as a pre-processor to read texts, then combine it with a visual question answering (VQA) framework. However, inaccurate OCR results may lead to cumulative error propagation , and the correlation between text reading and text-based reasoning is not fully exploited. In this work, we integrate OCR into the flow of TextVQA, targeting the mutual reinforcement of OCR and VQA tasks. Specifically, a visually enhanced text embedding module is proposed to predict semantic features from the visual information of texts, by which texts can be reasonably understood even without accurate recognition. Further, two elaborate schemes are developed to leverage contextual information in VQA to modify OCR results. The first scheme is a reading modification module that adaptively selects the answer results according to the contexts. Second, we propose an efficient end-to-end text reading and reasoning network, where the downstream VQA signal contributes to the optimization of text reading. Extensive experiments show that our method outperforms existing alternatives in terms of accuracy and robustness, whether ground truth OCR annotations are used or not.
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015, Weiping Wang 0005, Xu-Cheng Yin
Pattern Recognit.8
2023 Hypersphere guided embedding for masked face recognition
Xiaobin Zhu 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin, Lei Chen 0069
Pattern Recognit. Lett.5
2023 Self-supervised contrastive speaker verification with nearest neighbor positive instances
Li-Fang Wei, Chuan-Fei Zhang, Tian-Hao Zhang, Song-Lu Chen, Xu-Cheng Yin
Pattern Recognit. Lett.6
2023 HFENet: Hybrid Feature Enhancement Network for Detecting Texts in Scenes and Traffic Panels
abstract
Text detection in complex scene images is a challenging task for intelligent transportation. Existing scene text detection methods often adopt multi-scale feature learning strategies to extract informative feature representations for covering objects of various sizes. However, the sampling operation inherent in multi-scale feature generation can easily impair high-frequency details (e.g., textures and boundaries), which are critical for text detection. In this work, we propose an innovative Hybrid Feature Enhancement Network (dubbed HFENet) to explicitly improve the quality of high-frequency information for detecting texts in scenes and traffic panels. To be concrete, we propose a simple yet effective self-guided feature enhancement module (SFEM) for globally lifting feature representations to highly discriminative and high-frequency abundant ones. Notably, our SFEM is pluggable and will be removed after training without introducing extra computational costs. In addition, due to the challenge and importance of accurately predicting boundaries for text detection, we propose a novel boundary enhancement module (BEM) to explicitly strengthen local feature representations in the guidance of boundary annotation for accurate localization. Extensive experiments on multiple publicly available datasets (i.e., MSRA-TD500, CTW1500, Total-Text, Traffic Guide Panel Dataset, Chinese Road Plate Dataset, and ASAYAR_TXT) verify the state-of-the-art performance of our method.
Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.5
2023 RUArt: A Novel Text-Centered Solution for Text-Based Visual Question Answering
abstract
Text-based visual question answering (VQA) requires to read and understand text in an image to correctly answer a given question. However, most current methods simply add optical character recognition (OCR) tokens extracted from the image into the VQA model without considering contextual information of OCR tokens and mining the relationships between OCR tokens and scene objects. In this paper, we propose a novel text-centered method called RUArt (Reading, Understanding and Answering the Related Text) for text-based VQA. Taking an image and a question as input, RUArt first reads the image and obtains text and scene objects. Then, it understands the question, OCRed text and objects in the context of the scene, and further mines the relationships among them. Finally, it answers the related text for the given question through text semantic matching and reasoning. We evaluate our RUArt on two text-based VQA benchmarks (ST-VQA and TextVQA) and conduct extensive ablation studies for exploring the reasons behind RUArt’s effectiveness. Experimental results demonstrate that our method can effectively explore the contextual information of the text and mine the stable relationships between the text and objects.
Zanxia Jin, Heran Wu, Jingyan Qin, Lei Xiao 0001, Xu-Cheng Yin
IEEE Trans. Multim.7
2023 Kernel Proposal Network for Arbitrary Shape Text Detection
abstract
Segmentation-based methods have achieved great success for arbitrary shape text detection. However, separating neighboring text instances is still one of the most challenging problems due to the complexity of texts in scene images. In this article, we propose an innovative kernel proposal network (dubbed KPN) for arbitrary shape text detection. The proposed KPN can separate neighboring text instances by classifying different texts into instance-independent feature maps, meanwhile avoiding the complex aggregation process existing in segmentation-based arbitrary shape text detection methods. To be concrete, our KPN will predict a Gaussian center map for each text image, which will be used to extract a series of candidate kernel proposals (i.e., dynamic convolution kernel) from the embedding feature maps according to their corresponding keypoint positions. To enforce the independence between kernel proposals, we propose a novel orthogonal learning loss (OLL) via orthogonal constraints. Specifically, our kernel proposals contain important self-information learned by network and location information by position embedding. Finally, kernel proposals will individually convolve all embedding feature maps for generating individual embedded maps of text instances. In this way, our KPN can effectively separate neighboring text instances and improve the robustness against unclear boundaries. To the best of our knowledge, our work is the first to introduce the dynamic convolution kernel strategy to efficiently and effectively tackle the adhesion problem of neighboring text instances in text detection. Experimental results on challenging datasets verify the impressive performance and efficiency of our method. The code and model are available at https://github.com/GXYM/KPN.
Shi-Xue Zhang, Xiaobin Zhu 0001, Jie-Bo Hou, Xu-Cheng Yin
IEEE Trans. Neural Networks Learn. Syst.5
2022 Learning Aligned Cross-Modal Representation for Generalized Zero-Shot Classification
abstract
Learning a common latent embedding by aligning the latent spaces of cross-modal autoencoders is an effective strategy for Generalized Zero-Shot Classification (GZSC). However, due to the lack of fine-grained instance-wise annotations, it still easily suffer from the domain shift problem for the discrepancy between the visual representation of diversified images and the semantic representation of fixed attributes. In this paper, we propose an innovative autoencoder network by learning Aligned Cross-Modal Representations (dubbed ACMR) for GZSC. Specifically, we propose a novel Vision-Semantic Alignment (VSA) method to strengthen the alignment of cross-modal latent features on the latent subspaces guided by a learned classifier. In addition, we propose a novel Information Enhancement Module (IEM) to reduce the possibility of latent variables collapse meanwhile encouraging the discriminative ability of latent variables. Extensive experiments on publicly available datasets demonstrate the state-of-the-art performance of our method.
Zhiyu Fang, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin
AAAI6
2022 Open-Set Text Recognition via Character-Context Decoupling
abstract
The open-set text recognition task is an emerging chal-lenge that requires an extra capability to cognize novel characters during evaluation. We argue that a major cause of the limited performance for current methods is the con-founding effect of contextual information over the visual information of individual characters. Under open-set sce-narios, the intractable bias in contextual information can be passed down to visual information, consequently im-pairing the classification performance. In this paper, a Character-Context Decoupling framework is proposed to alleviate this problem by separating contextual information and character-visual information. Contextual information can be decomposed into temporal information and lin-guistic information. Here, temporal information that mod-els character order and word length is isolated with a de-tached temporal attention module. Linguistic information that models n- gram and other linguistic statistics is sepa-rated with a decoupled context anchor mechanism. A va-riety of quantitative and qualitative experiments show that our method achieves promising performance on open-set, zero-shot, and close-set text recognition datasets.
Chang Liu 0083, Xu-Cheng Yin
CVPR3
2022 Non-Autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition
abstract
Non-autoregressive (NAR) transformer models have been studied intensively in automatic speech recognition (ASR), and many NAR transformer models is to use the causal mask to limit token dependencies. However, the causal mask is designed for the left-to-right decoding process of the non-parallel autoregressive (AR) transformer, which is inappropriate for the parallel NAR transformer since it ignores the right-to-left contexts. Some methods are proposed to utilize right-to-left contexts with an extra decoder, but these methods increase the model complexity. To tackle the above problems, we propose a new non-autoregressive transformer with a unified bidirectional decoder (NAT-UBD), which can simultaneously utilize left-to-right and right-to-left contexts for ASR. However, direct use of bidirectional contexts will cause information leakage, which means the decoder output can be affected by the character information of the input in the same position. To avoid information leakage, we propose a novel attention mask and modify vanilla queries, keys, and values matrices for NAT-UBD. Experimental results verify that NAT-UBD can achieve character error rates (CERs) of 5.0%/5.5% on the Aishell-1 dev/test sets, outperforming all previous NAR transformer models. Moreover, NAT-UBD can run 49.8× faster than the AR transformer baseline when decoding in a single step.
Chuan-Fei Zhang, Tian-Hao Zhang, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICASSP6
2022 Adaptive Rounding Compensation for Post-training Quantization
Jinhui Lin, Song-Lu Chen, Ruiyao Zhang, Zhiwei Dong, Feng Chen 0040, Xu-Cheng Yin
ICONIP (5)8
2022 Semi-Supervised Fine-Grained Classification with Web Data via Noisy Sample Selection
abstract
For fine-grained classification, it is extremely difficult and costly to acquire the annotated data. Hence, some studies propose to use web data for fine-grained classification. However, the web data contains tremendous noisy labels, which can affect the classification results. Although many previous studies propose to discard noisy data via sample selection, they also discard some valid data. The valid data denotes hard or mislabeled samples that can enhance the robustness of the model. To solve the above problems, we propose a novel method to discard irrelevant noisy data from web data while keeping valid data for fine-grained classification. Specifically, we divide the web data into clean and noisy samples and then distinguish the noisy samples into open-set and close-set noises. Finally, the model is constructed in a semi-supervised manner, where the clean samples are used as the labeled set, and the close-set noises are used as the unlabeled set. Extensive experiments verify that our method can improve the classification performance by an average of 1.89% on three fine-grained benchmark datasets compared with the current methods. The experimental results prove the effectiveness of the combination of sample selection and semi-supervised training strategy.
Meng-Xuan Li, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICPR6
2022 DANet: Dynamic Attention to Spoof Patterns for Face Anti-Spoofing
abstract
Face anti-spoofing is a vital part to protect the security of face recognition systems. Many existing face anti-spoofing methods rely on convolutional neural networks (CNNs) and achieve competitive performance. However, due to the power of CNNs, these methods will extract information that is irrelevant to spoof patterns, such as acquisition equipment and environmental characteristics, which makes the network vulnerable to changes of the illumination or camera. In this work, we propose a plug-and-play module called DyAttention, which can improve the robustness against environmental changes. Moreover, we build a network named DANet with DyAttention, which can accurately capture the spoof patterns from coarse to fine. DANet can dynamically capture the texture differences between live and spoof samples in the facial area. Specifically, we use the spatial attention mechanism to generate a mask of the facial area. Then, we extract the intrinsic texture patterns and piecewise enhance them via dynamic activation for clean representation, where the texture patterns are not affected by the environmental and domain factors. Through experiments on three benchmark datasets, our DANet achieves state-of-the-art intra-dataset accuracy on CASIA-MFSD, Replay-Attack, and OULU-NPU. Meanwhile, DANet can enhance the cross-dataset performance between CASIA-MFSD and Replay-Attack, improving the average HTER by 1.3%.
Chun-Yu Sun, Song-Lu Chen, Xinjie Li 0002, Feng Chen 0040, Xu-Cheng Yin
ICPR5
2022 CAliC: Accurate and Efficient Image-Text Retrieval via Contrastive Alignment and Visual Contexts Modeling
abstract
Image-text retrieval is an essential task of information retrieval, in which the models with the Vision-and-Language Pretraining(VLP) are able to achieve ideal accuracy compared with the ones without VLP. Among different VLP approaches, the single-stream models achieve the overall best retrieval accuracy, but slower inference speed. Recently, researchers have introduced the two-stage retrieval setting commonly used in the information retrieval field to the single-stream VLP model for a better accuracy/efficiency trade-off. However, the retrieval accuracy and efficiency are still unsatisfactory mainly due to the limitations of the patch-based visual unimodal encoder in these VLP models. The unimodal encoders are trained on pure visual data, so the visual features extracted by them are difficult to align with the textual features and it is also difficult for the multi-modal encoder to understand visual information. Under these circumstances, we propose an accurate and efficient two-stage image-text retrieval model via Contrastive Alignment and visual Contexts modeling(CAliC). In the first stage of the proposed model, the visual unimodal encoder is pretrained with cross-modal contrastive learning to extract easily aligned visual features, which improves the retrieval accuracy and the inference speed. In the second stage of the proposed model, we introduce a new visual contexts modeling task during pretraining to help the multi-modal encoder better understand the visual information and get more accurate predictions. Extensive experimental evaluation validates the effectiveness of our proposed approach, which achieves a higher retrieval accuracy while keeping a faster inference speed, and outperforms existing state-of-the-art retrieval methods on image-text retrieval tasks over Flickr30K and COCO benchmarks.
Chao Zhu 0003, Mengyin Liu, Weibo Gu, Hongfa Wang, Wei Liu 0005, Xu-Cheng Yin
ACM Multimedia7
2022 From Token to Word: OCR Token Evolution via Contrastive Learning and Semantic Matching for Text-VQA
abstract
Text-based Visual Question Answering (Text-VQA) is a question-answering task to understand scene text, where the text is usually recognized by Optical Character Recognition (OCR) systems. However, the text from OCR systems often includes spelling errors, such as "pepsi" being recognized as "peosi". These OCR errors are one of the major challenges for Text-VQA systems. To address this, we propose a novel Text-VQA method to alleviate OCR errors via OCR token evolution. First, we artificially create the misspelled OCR tokens in the training time, and make the system more robust to the OCR errors. To be specific, we propose an OCR Token-Word Contrastive (TWC) learning task, which pre-trains word representation by augmenting OCR tokens via the Levenshtein distance between the OCR tokens and words in a dictionary. Second, by assuming that the majority of characters in misspelled OCR tokens are still correct, a multimodal transformer is proposed and fine-tuned to predict the answer using character-based word embedding. Specifically, we introduce a vocabulary predictor with character-level semantic matching, which enables the model to recover the correct word from the vocabulary even with misspelled OCR tokens. A variety of experimental evaluations show that our method outperforms the state-of-the-art methods on both TextVQA and ST-VQA datasets. The code will be released at https://github.com/xiaojino/TWA.
Zanxia Jin, Zheng Shou 0001, Satoshi Tsutsui, Jingyan Qin, Xu-Cheng Yin
ACM Multimedia6
2022 SD-GAN: Semantic Decomposition for Face Image Synthesis with Discrete Attribute
abstract
Manipulating latent code in generative adversarial networks (GANs) for facial image synthesis mainly focuses on continuous attribute synthesis (e.g., age, pose and emotion), while discrete attribute synthesis (like face mask and eyeglasses) receives less attention. Directly applying existing works to facial discrete attributes may cause inaccurate results. In this work, we propose an innovative framework to tackle challenging facial discrete attribute synthesis via semantic decomposing, dubbed SD-GAN. To be concrete, we explicitly decompose the discrete attribute representation into two components, i.e. the semantic prior basis and offset latent representation. The semantic prior basis shows an initializing direction for manipulating face representation in the latent space. The offset latent presentation obtained by 3D-aware semantic fusion network is proposed to adjust prior basis. In addition, the fusion network integrates 3D embedding for better identity preservation and discrete attribute synthesis. The combination of prior basis and offset latent representation enable our method to synthesize photo-realistic face images with discrete attributes. Notably, we construct a large and valuable dataset MEGN (Face Mask and Eyeglasses images crawled from Google and Naver) for completing the lack of discrete attributes in the existing dataset. Extensive qualitative and quantitative experiments demonstrate the state-of-the-art performance of our method. Our code is available at an anonymous website: https://github.com/MontaEllis/SD-GAN.
Kangneng Zhou, Xiaobin Zhu 0001, Daiheng Gao, Kai Lee, Xinjie Li 0002, Xu-Cheng Yin
ACM Multimedia6
2022 Anchor-Free Location Refinement Network for Small License Plate Detection
Zhen-Jia Li, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
PRCV (4)5
2022 In the eye of the beholder: A survey of gaze tracking techniques
Jiahui Liu 0004, Jiannan Chi, Huijie Yang, Xu-Cheng Yin
Pattern Recognit.4
2022 Occluded Pedestrian Detection via Distribution-Based Mutual-Supervised Feature Learning
abstract
Pedestrian detection is a very important task in intelligent transportation system. State-of-the-art detectors work well on non-occluded pedestrians, but they are still far from satisfactory for heavily occluded ones. Recently, to deal with occlusion problems, the popular two-stage approaches are to build a two-branch architecture with the help of additional visible body annotations. However, these methods still have disadvantages. Either the two branches only use score-level fusion, which cannot guarantee the detectors to learn more robust pedestrian features. Or they only focus on the features of visible part via the attention mechanisms. However, the visible body features of heavily occluded pedestrians are only concentrated in a relatively small area, which may easily lead to missed detections. To alleviate the above issues, we propose a novel Distribution-based Mutual-Supervised Feature Learning Network (DMSFLN), to better deal with occluded pedestrian detection. The key DMSFL module in our network is to learn more discriminative feature representations of pedestrians by minimizing the similarity loss between feature distributions of full body and visible body, which has two advantages: enhancing the feature representations of occluded pedestrians and reducing the intra-class variance in pedestrians. To facilitate the DMSFL module, we also propose a novel two-branch network architecture, which is trained in a mutual-supervised way with both full body and visible body annotations respectively. Extensive experiments are conducted on four challenging pedestrian datasets: Caltech, CityPersons, CrowdHuman and CUHK occlusion. Our approach achieves superior performance compared to other state-of-the-art methods, especially on heavy occlusion subsets.
Ye He 0004, Chao Zhu 0003, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.3
2022 Depth-Guided Progressive Network for Object Detection
abstract
Multi-scale object detection in natural scenes is still challenging. To enhance the multi-scale perception capability, some algorithms combine the lower-level and higher-level information via multi-scale feature fusion strategies. However, the inherent spatial properties among instances and relations between foreground and background are ignored. In addition, the human-defined “center-based” regression quality evaluation strategy, predicting a high-to-low score based on a linear relationship with the distance to the center of ground-truth box, is not robust to scale-variant objects. In this work, we propose a Depth-Guided Progressive Network (DGPNet) for multi-scale object detection. Specifically, besides the prediction of classification and localization, the depth is estimated and used to guide the image features in a weighted manner to obtain a better spatial representation. Therefore, depth estimation and 2D object detection are simultaneously learned via a unified network, where the depth features are merged as auxiliary information into the detection branch to enhance the discrimination among multi-scale objects. Moreover, to overcome the difficulty of empirically fitting the localization quality function, high-quality predicted boxes on scale-variant objects are more adaptively obtained by an IoU-aware progressive sampling strategy. We divide the sampling process into two stages, i.e., “statistical-aware” and “IoU-aware”. The former selects thresholds for positive samples based on statistical characteristics of multi-scale instances, and the latter further selects high-quality samples by IoU on the basis of the former. Therefore, the final ranking scores better reflect the quality of localization. Experiments verify that our method outperforms state-of-the-art methods on the KINS and Cityscapes dataset.
Jia-Wei Ma, Song-Lu Chen, Feng Chen 0040, Shu Tian, Jingyan Qin, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.7
2021 Adaptive Pattern-Parameter Matching for Robust Pedestrian Detection
abstract
Pedestrians with challenging patterns, e.g. small scale or heavy occlusion, appear frequently in practical applications like autonomous driving, which remains tremendous obstacle to higher robustness of detectors. Although plenty of previous works have been dedicated to these problems, properly matching patterns of pedestrian and parameters of detector, i.e., constructing a detector with proper parameter sizes for certain pedestrian patterns of different complexity, has been seldom investigated intensively. Pedestrian instances are usually handled equally with the same amount of parameters, which in our opinion is inadequate for those with more difficult patterns and leads to unsatisfactory performance. Thus, we propose in this paper a novel detection approach via adaptive pattern-parameter matching. The input pedestrian patterns, especially the complex ones, are first disentangled into simpler patterns for detection head by Pattern Disentangling Module (PDM) with various receptive fields. Then, Gating Feature Filtering Module (GFFM) dynamically decides the spatial positions where the patterns are still not simple enough and need further disentanglement by the next-level PDM. Cooperating with these two key components, our approach can adaptively select the best matched parameter size for the input patterns according to their complexity. Moreover, to further explore the relationship between parameter sizes and their performance on the corresponding patterns, two parameter selection policies are designed: 1) extending parameter size to maximum, aiming at more difficult patterns for different occlusion types; 2) specializing parameter size by group division, aiming at complex patterns for scale variations. Extensive experiments on two popular benchmarks, Caltech and CityPersons, show that our proposed method achieves superior performance compared with other state-of-the-art methods on subsets of different scales and occlusion types.
Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
AAAI4
2021 Adaptive Boundary Proposal Network for Arbitrary Shape Text Detection
abstract
Arbitrary shape text detection is a challenging task due to the high complexity and variety of scene texts. In this work, we propose a novel adaptive boundary proposal network for arbitrary shape text detection, which can learn to directly produce accurate boundary for arbitrary shape text without any post-processing. Our method mainly consists of a boundary proposal model and an innovative adaptive boundary deformation model. The boundary proposal model constructed by multi-layer dilated convolutions is adopted to produce prior information (including classification map, distance field, and direction field) and coarse boundary proposals. The adaptive boundary deformation model is an encoder-decoder network, in which the encoder mainly consists of a Graph Convolutional Network (GCN) and a Recurrent Neural Network (RNN). It aims to perform boundary deformation in an iterative way for obtaining text instance shape guided by prior information from the boundary proposal model. In this way, our method can directly and efficiently generate accurate text boundaries without complex post-processing. Extensive experiments on publicly available datasets demonstrate the state-of-the-art performance of our method. Code is available at the website: https://github.com/GXYM/TextBPN.
Shi-Xue Zhang, Xiaobin Zhu 0001, Hongfa Wang, Xu-Cheng Yin
ICCV5
2021 Fast Recognition for Multidirectional and Multi-type License Plates with 2D Spatial Attention
Qi Liu 0041, Song-Lu Chen, Zhen-Jia Li, Feng Chen 0040, Xu-Cheng Yin
ICDAR (4)6
2021 Dynamic Receptive Field Adaptation for Attention-Based Text Recognition
Haibo Qin, Xiaobin Zhu 0001, Xu-Cheng Yin
ICDAR (2)4
2021 Robust Chinese License Plate Generation via Foreground Text and Background Separation
Qi Liu 0041, Song-Lu Chen, Xu-Cheng Yin
ICIG (3)5
2021 Online Scene Text Tracking with Spatial-Temporal Relation
Yan Xiu, Shu Tian, Xu-Cheng Yin
ICIG (3)4
2021 Real-World Image Super-Resolution Via Spatio-Temporal Correlation Network
abstract
Super-resolving real-world image is very challenging due to the degradations in real-world low-resolution images are highly complicated. In this paper, we propose a novel Spatio-temporal Correlation Network (STCN) for real-world single image super-resolution. Specifically, we adopt a very deep network which consists of several attention groups. Each attention group (AG) contains a series of residual channel attention blocks (RCABs) and one spatio-temporal correlation block (STCB). Notably, STCB mainly consists of a residual 3D convolution, and aims to fully explore the local spatial and temporal correlations between channels of feature maps generated by RCABs for selectively capturing more informative features. In addition, we propose an innovative dual restriction (DR) through a simple degradation model to reduce the possible space of mapping functions in super-resolution. Experiments conducted on two public available real-world datasets demonstrate the superior performance of our method.
Xiaobin Zhu 0001, Xu-Cheng Yin
ICME4
2021 Hybrid approach for big data localization and semantic annotation
abstract
Summary Most of the data concerning business‐oriented systems are still based on either NoSQL or the relational data model. On the other hand, Semantic Web data model Resource Description Framework (RDF) has become the new standard for data modeling and analysis. Due to this situation integration of NoSQL, Relational Database (RDB) and RDF data models are becoming a required feature of the systems. Many solutions like tools and languages are provided in the shape of the transformation of data from RDB to RDF. This research is aimed to compare and map data models used for transformation between NoSQL, RDB, and Semantic Web. This study will help in achieving much better and enhanced technology‐based systems for retrieval and storage of data among Big‐data annotation using Semantic Web. It is aimed to reduce the response time of queries and offer compatibility with the web and semantically enriched data format. A drugs dataset is being used and transformed to have semantical meaning embedded and linked to support big data localization. At the end of this paper, RDF graph and bar chart are used to represent transformed data after passing through the proposed model. Big data localization helps in gaining fast and accurate results.
Waheed Yousuf Ramay, Xu-Cheng Yin, Shams Rahman, Muhammad Asif Habib
Concurr. Comput. Pract. Exp.2
2021 End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
Song-Lu Chen, Shu Tian, Jia-Wei Ma, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
Neurocomputing7
2021 Multi-orientation scene text detection with scale-guided regression
Jie-Bo Hou, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin
Neurocomputing6
2021 GCCNet: Grouped channel composition network for scene text detection
Chang Liu 0083, Jie-Bo Hou, Long-Huang Wu, Xiaobin Zhu 0001, Lei Xiao 0001, Xu-Cheng Yin
Neurocomputing7
2021 Detecting Text in Scene and Traffic Guide Panels With Attention Anchor Mechanism
abstract
Text detection in complex scene images is a challenging task for intelligent transportation. Recently, anchor mechanisms are widely utilized in scene text detection tasks. However, in existing methods, anchors are generally predefined empirically, degrading robustness to complex scenarios with various sizes and orientation variations. In this paper, we propose a novel Attention Anchor Mechanism (AAM), especially targeting at predicting appropriate anchors for each pixel. To be concrete, we regard a series of predefined anchors as basic anchors and utilize an attention model to predict weights corresponding to basic anchors. Consequently, the weighted sum of basic anchors in each pixel can obtain a predicted anchor. In this way, the gap between the predicted anchors and the corresponding ground truth boxes could be narrowed, making the network easier to regress. For facilitating the design of basic anchors, we adopt a dimension-decomposition mechanism to predict width, height, and angle of anchors, respectively. Extensive experiments on several public datasets demonstrate that our method achieves state-of-the-art performance.
Jie-Bo Hou, Xiaobin Zhu 0001, Chang Liu 0083, Long-Huang Wu, Hongfa Wang, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.7
2020 Deep Relational Reasoning Graph Network for Arbitrary Shape Text Detection
abstract
Arbitrary shape text detection is a challenging task due to the high variety and complexity of scenes texts. In this paper, we propose a novel unified relational reasoning graph network for arbitrary shape text detection. In our method, an innovative local graph bridges a text proposal model via Convolutional Neural Network (CNN) and a deep relational reasoning network via Graph Convolutional Network (GCN), making our network end-to-end trainable. To be concrete, every text instance will be divided into a series of small rectangular components, and the geometry attributes (e.g., height, width, and orientation) of the small components will be estimated by our text proposal model. Given the geometry attributes, the local graph construction model can roughly establish linkages between different text components. For further reasoning and deducing the likelihood of linkages between the component and its neighbors, we adopt a graph-based network to perform deep relational reasoning on local graphs. Experiments on public available datasets demonstrate the state-of-the-art performance of our method. Code is available at https://github.com/GXYM/DRRG.
Shi-Xue Zhang, Xiaobin Zhu 0001, Jie-Bo Hou, Chang Liu 0083, Hongfa Wang, Xu-Cheng Yin
CVPR7
2020 A Hybrid Self-Attention Model for Pedestrians Detection
Chao Zhu 0003, Xu-Cheng Yin
ICONIP (1)3
2020 Mutual-Supervised Feature Modulation Network for Occluded Pedestrian Detection
abstract
State-of-the-art pedestrian detectors have achieved significant progress on non-occluded pedestrians, yet they are still struggling under heavy occlusions. The recent occlusion handling strategy of popular two-stage approaches is to build a two-branch architecture with the help of additional visible body annotations. Nonetheless, these methods still have some weaknesses. Either the two branches are trained independently with only score-level fusion, which cannot guarantee the detectors to learn robust enough pedestrian features. Or the attention mechanisms are exploited to only emphasize on the visible body features. However, the visible body features of heavily occluded pedestrians are concentrated on a relatively small area, which will easily cause missing detections. To address the above issues, we propose in this paper a novel Mutual-Supervised Feature Modulation (MSFM) network, to better handle occluded pedestrian detection. The key MSFM module in our network calculates the similarity loss of full body boxes and visible body boxes corresponding to the same pedestrian so that the full-body detector could learn more complete and robust pedestrian features with the assist of contextual features from the occluding parts. To facilitate the MSFM module, we also propose a novel two-branch architecture, consisting of a standard full body detection branch and an extra visible body classification branch. These two branches are trained in a mutual-supervised way with full body annotations and visible body annotations, respectively. To verify the effectiveness of our proposed method, extensive experiments are conducted on two challenging pedestrian datasets: Caltech and CityPersons, and our approach achieves superior performance compared to other state-of-the-art methods on both datasets, especially in heavy occlusion cases.
Ye He 0004, Chao Zhu 0003, Xu-Cheng Yin
ICPR3
2020 Semantic Bilinear Pooling for Fine-Grained Recognition
abstract
Naturally, fine-grained recognition, e.g., vehicle identification or bird classification, has specific hierarchical labels, where fine categories are always harder to be classified than coarse categories. However, most of the recent deep learning based methods neglect the semantic structure of fine-grained objects and do not take advantage of the traditional fine-grained recognition techniques (e.g. coarse-to-fine classification). In this paper, we propose a novel framework with a two-branch network (coarse branch and fine branch), i.e., semantic bilinear pooling, for fine-grained recognition with a hierarchical label tree. This framework can adaptively learn the semantic information from the hierarchical levels. Specifically, we design a generalized cross-entropy loss for the training of the proposed framework to fully exploit the semantic priors via considering the relevance between adjacent levels and enlarge the distance between samples of different coarse classes. Furthermore, our method leverages only the fine branch when testing so that it adds no overhead to the testing time. Experimental results show that our proposed method achieves state-of-the-art performance on four public datasets.
Xinjie Li 0002, Song-Lu Chen, Chao Zhu 0003, Xu-Cheng Yin
ICPR5
2020 Global Context-Based Network with Transformer for Image2latex
abstract
Image2latex usually means converts mathematical formulas in images into latex markup. It is a very challenging job due to the complex two-dimensional structure, variant scales of input, and very long representation sequence. Many researchers use encoder-decoder based model to solve this task and achieved good results. However, these methods don't make full use of the structure and position information of the formula. To solve this problem, we propose a global context-based network with transformer that can (1) learn a more powerful and robust intermediate representation via aggregating global features and (2) encode position information explicitly and (3) learn latent dependencies between symbols by using self-attention mechanism. The experimental results on the dataset IM2LATEX-100K demonstrate the effectiveness of our method.
Nuo Pang, Xiaobin Zhu 0001, Jixuan Li, Xu-Cheng Yin
ICPR5
2020 SUMAC 2020: The 2nd Workshop on Structuring and Understanding of Multimedia heritAge Contents
abstract
SUMAC 2020 is the second edition of the workshop on Structuring and Understanding of Multimedia heritAge Contents. It is held in Seattle, USA on October 12th, 2020 and is co-located with the 28th ACM International Conference on Multimedia; this year, due to the sanitary crisis, it is organized virtually. Its objective is to present and discuss the latest and most significant trends and challenges in the analysis, structuring and understanding of multimedia contents dedicated to the valorization of heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field.
Valérie Gouet-Brunet, Margarita Khokhlova, Ronak Kosti, Liming Chen 0002, Xu-Cheng Yin
ACM Multimedia5
2020 Ranking via partial ordering for answer selection
Zanxia Jin, Bowen Zhang 0011, Jingyan Qin, Xu-Cheng Yin
Inf. Sci.5
2020 Chinese Short Text Classification with Mutual-Attention Convolutional Neural Networks
abstract
The methods based on the combination of word-level and character-level features can effectively boost performance on Chinese short text classification. A lot of works concatenate two-level features with little processing, which leads to losing feature information. In this work, we propose a novel framework called Mutual-Attention Convolutional Neural Networks, which integrates word and character-level features without losing too much feature information. We first generate two matrices with aligned information of two-level features by multiplying word and character features with a trainable matrix. Then, we stack them as a three-dimensional tensor. Finally, we generate the integrated features using a convolutional neural network. Extensive experiments on six public datasets demonstrate improved performance of our new framework over current methods.
Bo Xu 0002, Jing-Yi Liang, Bowen Zhang 0011, Xu-Cheng Yin
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2020 HAM: Hidden Anchor Mechanism for Scene Text Detection
abstract
Direct regression and anchor are the two mainly effective and prevailing mechanisms in the paradigm of scene text detection. However, the use of direct regression-based methods may be challenging during optimization without the help of anchors as references. Unfortunately, the anchor-based methods always suffer from the careful design of the anchors, degrading the robustness to complex scenes. To address the above-mentioned problems, we propose a novel hidden anchor mechanism (HAM) especially for scene text detection. The predictions of anchors are innovatively regarded as hidden layers, and the weighted sum of the predictions is integrated into a direct regression-based network. Hence, the architecture of our HAM still has the characteristic of simplicity as with direct regression-based methods. Moreover, it is easier to optimize anchors as references with this type of method than with direct regression-based methods. In this way, our network can take advantage of both direct regression and anchor mechanisms. In addition, we decouple three kinds of one-dimensional anchors from three-dimensional anchors, greatly reducing the number of anchors in text bounding box matching without performance degradation. We also propose a post-processing technique for long text detection, named iterative regression box (IRB), which takes a few additional computational costs and can be easily generalized to other methods. Experiments on several public datasets demonstrate that the proposed method achieves state-of-the-art performance. Code is available athttps://github.com/hjbplayer/HAM.
Jie-Bo Hou, Xiaobin Zhu 0001, Chang Liu 0083, Kekai Sheng, Long-Huang Wu, Hongfa Wang, Xu-Cheng Yin
IEEE Trans. Image Process.7
2020 Simultaneous End-to-End Vehicle and License Plate Detection With Multi-Branch Attention Neural Network
abstract
Vehicle and license plate detection plays an important role in intelligent transportation systems and is still a challenging task in real applications, such as on-road scenarios. Recently, Convolutional Neural Network (CNN)-based detectors achieve the state-of-the-art performance. However, it is difficult to efficiently detect the vehicle and license plate simultaneously in most cases. With a single network, the vehicle can affect the detection of the license plate due to the inclusion relation. In this paper, we propose an end-to-end deep neural network for detecting the vehicle and the license plate simultaneously in a given image, where two separate branches with different convolutional layers are designed for vehicle detection and license plate detection, respectively. In consideration of the license plate's small size and fairly obvious features as well as the vehicle's various size and rather complex features, the license plates are detected with low-level features and the vehicles are localized with multi-level features in corresponding convolutional layers. Moreover, a task-specific anchor design strategy is employed to obtain better predictions. Besides, the attention mechanisms and feature-fusion strategies are utilized to improve the detection performance of small-scale objects. A variety of experiments on real datasets and public datasets verify that our proposed method has fairly high accuracy and efficiency.
Song-Lu Chen, Jia-Wei Ma, Feng Chen 0040, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.5
2019 Detecting Text in News Images with Similarity Embedded Proposals
abstract
Text extraction plays an important role in news images analysis tasks. However, the conglutination of subtitles and station logos makes text detection challenging. In this paper, we develop an effective news text detection framework by introducing a novel similarity embedded proposal mechanism. The main idea is to predict similarity for each fine-scale coarse proposal to help construct text bounding boxes. Specifically, a CNN and bi-directional LSTM based network is used to produce vectors embedded in coarse proposals provided by Connectionist Text Proposal Network (CTPN). Notably, similarity embedded proposal mechanism can be generalized to other sub-text level text detection models. Comparing to the state-of-the-art method (CTPN), our framework improves F-measure by 25.2% on our Private News Dataset and 8.9% on ICDAR 2013 benchmarks, respectively.
Miaotong Jiang, Jie-Bo Hou, Xiaobin Zhu 0001, Xu-Cheng Yin
ICDAR5
2019 Combined Correlation Filters with Siamese Region Proposal Network for Visual Tracking
Shugang Cui, Shu Tian, Xu-Cheng Yin
ICONIP (2)3
2019 Image Generation Framework for Unbalanced License Plate Data Set
Xu-Cheng Yin
ICONIP (5)4
2019 Pyramid Memory Block and Timestep Attention for Speech Emotion Recognition
Xu-Cheng Yin
INTERSPEECH4
2019 Joint Rotation-Invariance Face Detection and Alignment with Angle-Sensitivity Cascaded Networks
abstract
Due to the angle variations especially in unconstrained scenarios, face detection and alignment have become challenging tasks. In existing methods, face detection and alignment are always conducted separately, which can greatly increase the computation cost. Moreover, this separation will abandon the inherent correlation underlying the two tasks. In this paper, we propose a simple but effective architecture, named Angle-Sensitivity Cascaded Networks (ASCN), for jointly conducting rotation-invariance face detection and alignment. ASCN mainly consists of three consecutive cascaded networks. Specifically, in the first stage, the rotation angle is predicted and candidate bounding boxes are proposed simultaneously. In the second stage, ASCN further refines the candidates and orientations. In the last stage, ASCN jointly learns the accurate bounding boxes and alignment. Besides, for accurately locating landmarks in hard examples, we introduce a pose-equitable loss to balance the faces with large poses. Extensive experiments conducted on benchmark datasets demonstrate the surprising performance of our method. Notably, our method maintains real-time efficiency for both detection and alignment tasks on the ordinary CPU platform.
Qi Liu 0041, Xu-Cheng Yin
ACM Multimedia4
2019 Health assistant: answering your questions anytime from biomedical literature
abstract
MOTIVATION: With the abundant medical resources, especially literature available online, it is possible for people to understand their own health status and relevant problems autonomously. However, how to obtain the most appropriate answer from the increasingly large-scale database, remains a great challenge. Here, we present a biomedical question answering framework and implement a system, Health Assistant, to enable the search process. METHODS: In Health Assistant, a search engine is firstly designed to rank biomedical documents based on contents. Then various query processing and search techniques are utilized to find the relevant documents. Afterwards, the titles and abstracts of top-N documents are extracted to generate candidate snippets. Finally, our own designed query processing and retrieval approaches for short text are applied to locate the relevant snippets to answer the questions. RESULTS: Our system is evaluated on the BioASQ benchmark datasets, and experimental results demonstrate the effectiveness and robustness of our system, compared to BioASQ participant systems and some state-of-the-art methods on both document retrieval and snippet retrieval tasks. AVAILABILITY AND IMPLEMENTATION: A demo of our system is available at https://github.com/jinzanxia/biomedical-QA.
Zanxia Jin, Bowen Zhang 0011, Fan Fang, Le-Le Zhang, Xu-Cheng Yin
Bioinform.5
2018 TED-KISS: A Known-Item Speech Video Search Benchmark
abstract
Known-item search is an everyday natural scenario that we search for a specific thing (maybe a song) while only remembering some details about it. Existing benchmarks generally focus on brief user requests which specify some metadata like the title, or the time. However, in most cases, the users can hardly recall such information accurately. In order to embrace the research of known-item search, we present a new publicly available known-item speech video search benchmark, namely TED-KISS, which takes TED talks as an example. The video collection is constructed with up-to-date nearly 80,000 TED and TEDx talks on Youtube. These talks cover various topics, and their titles, speakers, descriptions, full-text subtitles, as well as original links are extracted as metadata, which makes the researches on text-based retrieval and multimedia retrieval feasible. Unlike other benchmarks concerning visual contents in segments, the user requests in TED-KISS are generated through a more natural process, partly through original related topics posted on Reddit and Baidu Tieba, and partly through manual imitative requests annotated by volunteers in a scenario simulation. In addition, we analyze the characteristics of our benchmark through evaluations of several existing text-based IR and Neural-IR models, which also can be served as baselines for this task.
Fan Fang, Bowen Zhang 0011, Xu-Cheng Yin, Haixia Man
CIKM3
2018 LSTM2 : Multi-Label Ranking for Document Classification
Yan Yan 0004, Bowen Zhang 0011, Xu-Cheng Yin
Neural Process. Lett.6
2018 A Unified Framework for Tracking Based Text Detection and Recognition from Web Videos
abstract
Video text extraction plays an important role for multimedia understanding and retrieval. Most previous research efforts are conducted within individual frames. A few of recent methods, which pay attention to text tracking using multiple frames, however, do not effectively mine the relations among text detection, tracking and recognition. In this paper, we propose a generic Bayesian-based framework of Tracking based Text Detection And Recognition (T DAR) from web videos for embedded captions, which is composed of three major components, i.e., text tracking, tracking based text detection, and tracking based text recognition. In this unified framework, text tracking is first conducted by tracking-by-detection. Tracking trajectories are then revised and refined with detection or recognition results. Text detection or recognition is finally improved with multi-frame integration. Moreover, a challenging video text (embedded caption text) database (USTB-VidTEXT) is constructed and publicly available. A variety of experiments on this dataset verify that our proposed approach largely improves the performance of text detection and recognition from web videos.
Shu Tian, Xu-Cheng Yin, Ya Su, Hongwei Hao
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 A calculation method for social network user credibility
abstract
Trust plays an important role in helping social network user make appropriate decision from abundant social information about products. A variety of social review information (e.g. ratings, voting and tags) emerge with the huge number of products on the Web. How they are utilized for searching and finding appropriate item is investigated. In this paper we propose a user credibility calculation method, which calculates user's credibility through his previous review information. By evaluating facets of every review, an integrated numerical value which denotes the reviewer's credibility can be calculated. This value can be used to rank products, further to help user making appropriate decision. Experiments on social book search database show that this calculation method improves effectively accuracy of recommended items.
Jian-Lin Jin, Xiaojiang Du, Bowen Zhang 0011, Xu-Cheng Yin
ICC5
2017 ICDAR2017 Robust Reading Challenge on Text Extraction from Biomedical Literature Figures (DeTEXT)
abstract
Hundreds of millions of figures are available in the biomedical literature, representing important biomedical experimental evidence. Since text is a rich source of information in figures, automatically extracting such text may assist in the task of mining figure information and understanding biomedical documents. Unlike images in the open domain, biomedical figures present a variety of unique challenges. For example, biomedical figures typically have complex layouts, small font sizes, short text, specific text, complex symbols and irregular text arrangements. This paper presents the final results of the ICDAR 2017 Competition on Text Extraction from Biomedical Literature Figures (ICDAR2017 DeTEXT Competition), which aims at extracting (detecting and recognizing) text from biomedical literature figures. Similar to text extraction from scene images and web pictures, ICDAR2017 DeTEXT Competition includes three major tasks, i.e., text detection, cropped word recognition and end-to-end text recognition. Here, we describe in detail the data set, tasks, evaluation protocols and participants of this competition, and report the performance of the participating methods.
Xu-Cheng Yin, Dimosthenis Karatzas
ICDAR2
2017 Building Your Own Reading List Anytime via Embedding Relevance, Quality, Timeliness and Diversity
abstract
During every summer holidays, several editions of reading lists are recommended and emerged on mass media, e.g., New York Times, and BBC. However, these reading lists are built for whole people with general topics for some purposes. What if we expect the books of a specific topic at a specific moment? How to generate the requested reading list for our own automatically? In this paper, we propose a searching framework for building a topical reading list anytime, where the Relevance (between topics and books), Quality (of books), Timeliness (of popularities) and Diversity (of results) are embedded into vector representations respectively based on user-generated contents and statistics on social media. We collected 8,197 real-world topics from 198 diverse groups on Librarything.com. The proposed methods are evaluated on the topic collection and the public benchmarks Social Book Search 2012-2016 (SBS). Experimental results demonstrate the robustness and effectiveness of our framework.
Bowen Zhang 0011, Xu-Cheng Yin, Jian-Lin Jin
SIGIR2
2017 Fast alignment for sparse representation based face recognition
Ya Su, Xinbo Gao 0001, Xu-Cheng Yin
Pattern Recognit.3
2017 Tracking Based Multi-Orientation Scene Text Detection: A Unified Framework With Dynamic Programming
abstract
There are a variety of grand challenges for multi-orientation text detection in scene videos, where the typical issues include skew distortion, low contrast, and arbitrary motion. Most conventional video text detection methods using individual frames have limited performance. In this paper, we propose a novel tracking based multi-orientation scene text detection method using multiple frames within a unified framework via dynamic programming. First, a multi-information fusion-based multi-orientation text detection method in each frame is proposed to extensively locate possible character candidates and extract text regions with multiple channels and scales. Second, an optimal tracking trajectory is learned and linked globally over consecutive frames by dynamic programming to finally refine the detection results with all detection, recognition, and prediction information. Moreover, the effectiveness of our proposed system is evaluated with the state-of-the-art performances on several public data sets of multi-orientation scene text images and videos, including MSRA-TD500, USTB-SV1K, and ICDAR 2015 Scene Videos.
Xu-Cheng Yin, Wei-Yi Pei, Shu Tian, Ze-Yu Zuo, Chao Zhu 0003, Junchi Yan
IEEE Trans. Image Process.2
2016 Multi-orientation scene text detection with multi-information fusion
abstract
We construct a robust and precise multi-orientation text detection system in scene images which can extensively locate possible characters with multi-information fusion. In our method, an adaptive multi-channel character grouping algorithm is first proposed to extract all possible character candidates robustly, and an AdaBoost classifier is then to properly identify character candidates as characters or non-characters. A single-link clustering with distance metric learning is thereafter used to adaptively group characters into text regions, and an effective hybrid filter with Convolution Neural Networks (CNN), AdaBoost and Bayesian classifiers is finally designed to precisely verify the extracted text regions. Our proposed technology is extensively evaluated on several public multi-orientation scene text datasets, e.g., MSRA-TD500 and USTB-SV1K, and is much better than state-of-the-art methods.
Wei-Yi Pei, Lih-Jen Kau, Xu-Cheng Yin
ICPR4
2016 Scene Text Detection in Video by Learning Locally and Globally
Shu Tian, Wei-Yi Pei, Ze-Yu Zuo, Xu-Cheng Yin
IJCAI4
2016 A Short Survey of Recent Advances in Graph Matching
abstract
Graph matching, which refers to a class of computational problems of finding an optimal correspondence between the vertices of graphs to minimize (maximize) their node and edge disagreements (affinities), is a fundamental problem in computer science and relates to many areas such as combinatorics, pattern recognition, multimedia and computer vision. Compared with the exact graph (sub)isomorphism often considered in a theoretical setting, inexact weighted graph matching receives more attentions due to its flexibility and practical utility. A short review of the recent research activity concerning (inexact) weighted graph matching is presented, detailing the methodologies, formulations, and algorithms. It highlights the methods under several key bullets, e.g. how many graphs are involved, how the affinity is modeled, how the problem order is explored, and how the matching procedure is conducted etc. Moreover, the research activity at the forefront of graph matching applications especially in computer vision, multimedia and machine learning is reported. The aim is to provide a systematic and compact framework regarding the recent development and the current state-of-the-arts in graph matching.
Junchi Yan, Xu-Cheng Yin, Weiyao Lin, Cheng Deng 0002, Hongyuan Zha, Xiaokang Yang 0001
ICMR2
2016 A generic pseudo relevance feedback framework with heterogeneous social information
Bowen Zhang 0011, Xu-Cheng Yin
Inf. Sci.2
2016 Text Detection, Tracking and Recognition in Video: A Comprehensive Survey
abstract
The intelligent analysis of video data is currently in wide demand because a video is a major source of sensory data in our lives. Text is a prominent and direct source of information in video, while the recent surveys of text detection and recognition in imagery focus mainly on text extraction from scene images. Here, this paper presents a comprehensive survey of text detection, tracking, and recognition in video with three major contributions. First, a generic framework is proposed for video text extraction that uniformly describes detection, tracking, recognition, and their relations and interactions. Second, within this framework, a variety of methods, systems, and evaluation protocols of video text extraction are summarized, compared, and analyzed. Existing text tracking techniques, tracking-based detection and recognition techniques are specifically highlighted. Third, related applications, prominent challenges, and future directions for video text extraction (especially from scene videos and web videos) are also thoroughly discussed.
Xu-Cheng Yin, Ze-Yu Zuo, Shu Tian, Cheng-Lin Liu 0001
IEEE Trans. Image Process.1
2015 Multi-strategy tracking based text detection in scene videos
abstract
Text detection and tracking in scene videos are important prerequisites for content-based video analysis and retrieval, wearable camera systems and mobile devices augmented reality translators. Here, we present a novel multi-strategy tracking based text detection approach in scene videos. In this approach, a state-of-the-art scene text detection module [1] is first used to detect text in each video frame. Then a multi-strategy text tracking technique is proposed, which uses tracking by detection, spatio-temporal context learning, and linear prediction to predict the candidate text location sequentially, and adaptively integrates and selects the best matching text block from the candidate blocks with a rule-based method. This multi-strategy tracking technique can combine the advantages of the three different tracking techniques and afterwards make remedies to the disadvantages of them. Experiments on a variety of scene videos show that our proposed approach is effective and robust to reduce false alarm and improve the accuracy of detection.
Ze-Yu Zuo, Shu Tian, Wei-Yi Pei, Xu-Cheng Yin
ICDAR4
2015 DE2: Dynamic ensemble of ensembles for learning nonstationary data
Xu-Cheng Yin, Kaizhu Huang, Hongwei Hao
Neurocomputing1
2015 Learning Imbalanced Classifiers Locally and Globally with One-Side Probability Machine
Kaizhu Huang, Rui Zhang 0012, Xu-Cheng Yin
Neural Process. Lett.3
2015 Multi-Orientation Scene Text Detection with Adaptive Clustering
abstract
Text detection in natural scene images is an important prerequisite for many content-based image analysis tasks, while most current research efforts only focus on horizontal or near horizontal scene text. In this paper, first we present a unified distance metric learning framework for adaptive hierarchical clustering, which can simultaneously learn similarity weights (to adaptively combine different feature similarities) and the clustering threshold (to automatically determine the number of clusters). Then, we propose an effective multi-orientation scene text detection system, which constructs text candidates by grouping characters based on this adaptive clustering. Our text candidates construction method consists of several sequential coarse-to-fine grouping steps: morphology-based grouping via single-link clustering, orientation-based grouping via divisive hierarchical clustering, and projection-based grouping also via divisive clustering. The effectiveness of our proposed system is evaluated on several public scene text databases, e.g., ICDAR Robust Reading Competition data sets (2011 and 2013), MSRA-TD500 and NEOCR. Specifically, on the multi-orientation text data set MSRA-TD500, the f measure of our system is 71 percent, much better than the state-of-the-art performance. We also construct and release a practical challenging multi-orientation scene text data set (USTB-SV1K), which is available at http://prir.ustb.edu.cn/TexStar/MOMV-text-detection/.
Xu-Cheng Yin, Wei-Yi Pei, Hongwei Hao
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Social Book Search Reranking with Generalized Content-Based Filtering
abstract
Semantically searching and navigating products (e.g., on Taobao.com or Amazon.com) with professional metadata and user-generated content from social media is a hot topic in information retrieval and recommendation systems, while most existing methods are specifically designed as a purely searching system. In this paper, taking Social Book Search as an example, we propose a general search-recommendation hybrid system for this topic. Firstly, we propose a Generalized Content-Based Filtering (GCF) model. In this model, a preference value, which flexibly ranges from 0 to 1, is defined to describe a user's preference for each item to be recommended, unlike conventionally using a set of preferable items. We also design a weighting formulation for the measure of recommendation. Next, assuming that the query in a searching system acts as a user in a recommendation system, a general reranking model is constructed with GCF to rerank the initial resulting list by utilizing a variety of rich social information. Afterwards, we propose a general search-recommendation hybrid framework for Social Book Search, where learning-to-rank is used to adaptively combine all reranking results. Finally, our proposed system is extensively evaluated on the INEX 2012 and 2013 Social Book Search datasets, and has the best performance ([email protected]) on both datasets compared to other state-of-the-art systems. Moreover, our system recently won the INEX 2014 Social Book Search Evaluation.
Bowen Zhang 0011, Xu-Cheng Yin, Xiao-Ping Cui, Jiao Qu, Bin Geng, Hongwei Hao
CIKM2
2014 Social Book Search with Pseudo-Relevance Feedback
Bin Geng, Jiao Qu, Bowen Zhang 0011, Xiao-Ping Cui, Xu-Cheng Yin
ICONIP (2)6
2014 Text Categorization with Diversity Random Forests
Xu-Cheng Yin, Kaizhu Huang
ICONIP (3)2
2014 Diversity-Based Ensemble with Sample Weight Learning
abstract
Given multiple classifiers, one prevalent approach in classifier ensemble is to diversely combine classifier components (diversity-based ensemble), and a lot of previous works show that this approach can improve accuracy in classification. However, how to measure diversity and perform diversity-based learning are still challenges in the literature. Moreover, the learning procedure highly depends upon the distribution of the training data. In this paper, we propose a novel classifier ensemble method which combines classifiers with both diversity and sample weighting. First, by designing a matrix for the (sample) data distribution creatively, we formulate a unified optimization model for diversity-based ensemble with sample weighting, where classifier weights are learned through a convex quadratic programming problem with given sample weights. Second, we propose a new self-training algorithm to iteratively run the convex optimization and automatically learn the sample weights. Moreover, these sample weights are updated with a dynamically damped learning trick, which has a good performance for convergence. This paper also discusses the relationship between our optimization model and the margin theory. Extensive experiments on a variety of 50 UCI classification benchmark data sets show that the proposed approach consistently outperforms conventional ensembles such as Bagging, GASEN, and SDP.
Xu-Cheng Yin, Hongwei Hao
ICPR2
2014 Shallow Classification or Deep Learning: An Experimental Study
abstract
After being kick-started with major breakthrough in 2006 by Hinton, LeCun and Bengio respectively, deep learning has been becoming the mainstream for challenging classification systems, which, however always were with "shallow" discriminative classifiers in the past. In this paper, we argue that in common classification cases with plenty but not enough training examples, mixed-quality examples for dozens of categories, deep learning and shallow classification may have complementary performance. Then, we design a hybrid recognition strategy with classification switching to adaptively fuse deep learning and shallow classification technologies. Finally, we present a variety of experiments with visual recognition tasks, i.e., USPS character recognition, Caltech101 visual object classification, and ICDAR scene text recognition. Specifically, we perform word recognition by dynamically combing the conventional open source OCR engine with the present popular convolutional neural networks, and construct an effective end-to-end scene text recognition system with open-vocabulary. This end-to-end system is evaluated on ICDAR 2011 Robust Reading Competition (Challenge 2) dataset, the f measure of which is 54.5%, much better than 45.2% of the latest state-of-the-art performance.
Xu-Cheng Yin, Wei-Yi Pei, Hongwei Hao
ICPR1
2014 Bayesian network scores based text localization in scene images
abstract
Text localization in scene images is an essential and interesting task to analyze the image contents. In this work, a Bayesian network scores using K2 algorithm in conjunction with the geometric features based effective text localization method with the help of maximally stable extremal regions (MSERs). First, all MSER-based extracted candidate characters are directly compared with an existing text localization method to find text regions. Second, adjacent extracted MSER-based candidate characters are not encompassed into text regions due to strict edges constraint. Therefore, extracted candidate character regions are incorporated into text regions using selection rules. Third, K2 algorithm-based Bayesian networks scores are learned for the complimentary candidate character regions. Bayesian logistic regression classifier is built on the Bayesian network scores by computing the posterior probability of complimentary candidate character region corresponding to non-character candidates. The higher posterior probability of complimentary Candidate character regions are further grouped into words or sentences. Bayesian networks scores based text localization system, named as BayesText, is evaluated on ICDAR 2013 Robust Reading Competition (Challenge 2 Task 2.1: Text Localization) database. Experimental results have established significant competitive performance with the state-of-the-art text detection systems.
Khalid Iqbal, Xu-Cheng Yin, Hongwei Hao, Sohail Asghar, Hazrat Ali
IJCNN2
2014 A central tendency-based privacy preserving model for sensitive XML association rules using Bayesian networks
abstract
The rationale of XML design is to transfer and store data at different levels. A key feature of these levels in an XML document is to identify its components for additional processing. XML components can expose sensitive information after application
Khalid Iqbal, Xu-Cheng Yin, Hongwei Hao, Qazi Mudassar Ilyas, Xuwang Yin
Intell. Data Anal.2
2014 A novel classifier ensemble method with sparsity and diversity
Xu-Cheng Yin, Kaizhu Huang, Hongwei Hao, Khalid Iqbal, Zhi-Bin Wang
Neurocomputing1
2014 Robust Text Detection in Natural Scene Images
abstract
Text detection in natural scene images is an important prerequisite for many content-based image analysis tasks. In this paper, we propose an accurate and robust method for detecting texts in natural scene images. A fast and effective pruning algorithm is designed to extract Maximally Stable Extremal Regions (MSERs) as character candidates using the strategy of minimizing regularized variations. Character candidates are grouped into text candidates by the single-link clustering algorithm, where distance weights and clustering threshold are learned automatically by a novel self-training distance metric learning algorithm. The posterior probabilities of text candidates corresponding to non-text are estimated with a character classifier; text candidates with high non-text probabilities are eliminated and texts are identified with a text classifier. The proposed system is evaluated on the ICDAR 2011 Robust Reading Competition database; the f-measure is over 76%, much better than the state-of-the-art performance of 71%. Experiments on multilingual, street view, multi-orientation and even born-digital databases also demonstrate the effectiveness of the proposed method.
Xu-Cheng Yin, Xuwang Yin, Kaizhu Huang, Hongwei Hao
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Sorting-Based Dynamic Classifier Ensemble Selection
abstract
In ensemble learning, a higher accuracy can be achieved by integrating some classifiers instead of all the classifiers. But, it is very difficult to select the best classifier combination which can be seen as an optimization problem, from a pool of classifiers. To deal with this problem, we propose a new classifier selection method, Sorting-based Dynamic Classifier Ensemble Selection (SDES), which consists of two stages: (1) classifier sorting, and (2) dynamic ensemble selection on sorted classifier sequence. In the first stage, classifiers are sorted based on diversity, to avoid searching for the nearest neighbors in dynamic ensemble selection methods and greatly improve the selection efficiency. In the second stage, the optimal subset of classifiers is selected from the sorted classifier sequence based on confidence of test samples, to guarantee high accuracy of the optimal classifier subset. Experimental results have shown the effectiveness and high efficiency of the proposed method.
Yan Yan 0004, Xu-Cheng Yin, Zhi-Bin Wang, Xuwang Yin, Hongwei Hao
ICDAR2
2013 Dynamic Ensemble of Ensembles in Nonstationary Environments
Xu-Cheng Yin, Kaizhu Huang, Hongwei Hao
ICONIP (2)1
2013 Learning Bayesian Network leveled-structure from support based XML frequent itemsets
abstract
XML (eXtensible Markup Language) is a standard and entirely user-driven language for storage and transfer of information. XML frequent itemsets are usually found for mining XML association rules from XML transactional databases. These XML frequent itemsets lead researchers to find interesting XML patterns in large databases with the use of a threshold value. Apriori algorithm is one of the most leading solutions to discover XML frequent itemsets based on support value. XML frequent itemsets consist of similar items which show evidence of association. This relationship can be found with the use of Bayesian Network by learning structure of XML frequent itemsets. K2 algorithm is used to learn the structure of XML frequent itemsets. In this work, we propose a novel Apriori K2 algorithm. This algorithm is composed novel direction of apriori and K2 algorithms to find XML frequent itemsets and learning a level-wise Bayesian Network structure. For learning each level of this structure, XML frequent itemsets are found from XML candidate itemsets with the use of support measure using apriori algorithm. An updated binary table is prepared based on XML frequent itemset during the execution of apriori algorithm. K2 algorithm is used in conjunction with apriori algorithm to learn Bayesian Network structure of XML large frequent itemsets and find their relationship at each level. We have extensively tested our solution over UCI machine learning datasets and measured its performance. The results have shown that performance of our proposed solution is better than the combined performance of apriori and K2 algorithms.
Khalid Iqbal, Xu-Cheng Yin, Hongwei Hao, Qazi Mudassar Ilyas
IJCNN2
2013 Classifier comparison for MSER-based text classification in scene images
abstract
Text detection in images is an emerging area of interest with a growing motivation to researchers. Various methodologies have been developed to localize text contained in scene images. One main application of localizing scene image text is to produce a real time support to visually impaired persons. To design a real-time support platform for visually impaired persons, classification of textual information, i.e. character and non-character information can provide a baseline for further research. However, the challenge exists in choosing the optimum classifier for this purpose. In this work, first, we used Maximally Stable Extremal Regions (MSERs) to detect character candidates in a scene image; then, we trained several classifiers, i.e., AdaboostM1, Bayesian Logistic Regression, Naïve Bayes, and Bayes Net, to classify MSERs as characters and non-characters; and finally, we compared and analyzed the performances of these classifiers empirically. From experiments, it has been concluded that Bayesian Logistic Regression provides the better accuracy over the other three classifiers. This work argues that MSER based character candidates extraction and Bayesian Logistic Regression based text classification are two prominent and potential techniques in scene text detection.
Khalid Iqbal, Xu-Cheng Yin, Xuwang Yin, Hazrat Ali, Hongwei Hao
IJCNN2
2013 Transfer learning based compressive tracking
abstract
Although existing online tracking algorithms can solve the problems of scene illumination changes, partial or full object occlusions, and pose variation, there are still two weaknesses, inadequacy of training data and drift problem. Considering these, Compressive Tracking algorithm (CT) [1] extracts features from compressed domain, and classified object and background via a naive Bayes classier with online update. To further solve the problems of drift and inadequacy of training data, we introduce transfer learning into CT to take full advantage of prior information and propose a self-traininglike transfer learning algorithm. It selects training samples from samples collection to update classifier by the conduction of the classifier constructed at first frame. Eventually we introduce self-training-like transfer learning algorithm into CT to construct a novel tracking algorithm called Transfer Learning based Compressive Tracking (TLCT). Experimental results on 17 publicly available challenging sequences have shown the effectiveness and robustness of our algorithm.
Shu Tian, Xu-Cheng Yin, Hongwei Hao
IJCNN2
2013 Accurate and robust text detection: a step-in for text retrieval in natural scene images
abstract
We propose and implement a robust text detection system, which is a prominent step-in for text retrieval in natural scene images or videos. Our system includes several key components: (1) A fast and effective pruning algorithm is designed to extract Maximally Stable Extremal Regions as character candidates using the strategy of minimizing regularized variations. (2) Character candidates are grouped into text candidates by the single-link clustering algorithm, where distance weights and threshold of clustering are learned automatically by a novel self-training distance metric learning algorithm. (3) The posterior probabilities of text candidates corresponding to non-text are estimated with an character classifier; text candidates with high probabilities are then eliminated and finally texts are identified with a text classifier. The proposed system is evaluated on the ICDAR 2011 Robust Reading Competition dataset and a publicly available multilingual dataset; the f measures are over 76% and 74% which are significantly better than the state-of-the-art performances of 71% and 65%, respectively.
Xu-Cheng Yin, Xuwang Yin, Kaizhu Huang, Hongwei Hao
SIGIR1
2012 Pedestrian Analysis and Counting System with Videos
Zhi-Bin Wang, Hongwei Hao, Xu-Cheng Yin, Shu Tian
ICONIP (5)4
2012 Classifier Ensemble Using a Heuristic Learning with Sparsity and Diversity
Xu-Cheng Yin, Kaizhu Huang, Hongwei Hao, Khalid Iqbal, Zhi-Bin Wang
ICONIP (2)1
2012 Effective text localization in natural scene images with MSER, geometry-based grouping and AdaBoost
Xuwang Yin, Xu-Cheng Yin, Hongwei Hao, Khalid Iqbal
ICPR2
2012 Automatic segmentation of cervical vertebrae in X-ray images
abstract
Physiological parameters of vertebrae are important for cervical condition assessment. In order to measure the parameters fast and accurately, automatic segmentation instead of manual key point placement has become an imperative for diagnosing. We propose an applicable automatic segmentation system for medical image of cervical spine. The system includes a series of algorithms: a parallel cascade structure based Haar-like features and the AdaBoost learning algorithm used to detect the location of cervical vertebrae as a initial position of Active Appearance Model (AAM), multi-resolution AAM search applied to improve the speed and accuracy of AAM fit, and combination of global AAM and local AAM used to achieve more effective matching of details of vertebrae. Experiments on the cervical spine databases show a significant increase in speed, robustness and quality of fit compared to previous methods.
Hongwei Hao, Xu-Cheng Yin, Shawkat Hasan Shafin
IJCNN3
2011 Robust Vanishing Point Detection for MobileCam-Based Documents
abstract
Document images captured by a mobile phone camera often have perspective distortions. In this paper, fast and robust vanishing point detection methods for such perspective documents are presented. Most of previous methods are either slow or unstable. Based on robust detection of text baselines and character tilt orientations, our proposed technology is fast and robust with the following features: (1) quick detection of vanishing point candidates by clustering and voting on the Gaussian sphere space, and (2) precise and efficient detection of the final vanishing points using a hybrid approach, which combines the results from clustering and projection analysis. The rectified image acceptance rate for Mobile Cam-based documents, signboards and posters is more than 98% with an average speed of about 100ms.
Xu-Cheng Yin, Hongwei Hao, Jun Sun 0004, Satoshi Naoi
ICDAR1
2011 Learning Based Visibility Measuring with Images
Xu-Cheng Yin, Hongwei Hao, Xiao-Zhong Cao, Qing Li 0015
ICONIP (3)1
2011 Handwritten Chinese character identification with Bagged One-Class support vector machines
abstract
Today, more and more foreigners go to China and are emerged into studying Chinese. Thereinto, how to write Chinese characters is a very important and difficult task. As computers and internets develop, many teachers for Chinese Education want to use pattern recognition technologies to automatically evaluate and direct the quality of Chinese characters written by foreign students through document scanning. Actually, this is a handwritten character evaluation and identification problem. In this paper, we investigate and compare several character identification methods for Chinese Education within a classification framework. First, some two-class classification techniques with different features and classifiers (BP neural networks and SVMs) are investigated to identify each handwritten Chinese character. Moreover, in character identification, positive examples are always conjunctive, but negative examples are diffused in most cases. Consequently, we use one-class classification technique (one-class SVMs) to perform this handwritten character identification. In order to overcome the sensitivity to the SVM parameters, we propose a variant one-class SVM system - Bagged One-Class SVMs, which integrate many one-class SVMs with sample bagging. Some experiments of evaluating real handwritten Chinese characters by foreigners are performed, which show that general handwritten character identification is a big challenge and one-class classification technique is a potential researching and developing direction.
Hongwei Hao, Cui-Xia Mu, Xu-Cheng Yin, Zhi-Bin Wang
IJCNN3
2011 An improved topic relevance algorithm for focused crawling
abstract
Topic relevance of pages and hyperlinks is the key issue in focused crawling. In this paper, an improved topic relevance algorithm for focused crawling is proposed. First, we implement a prototype system of the focused crawler - a topic-specific news gathering system which is prepared for comparative experiments on different similarity measures with the anchor text. Second, experiments on Chinese text corpus show that using LSI (Latent Semantic Indexing) outperforms using TF-IDF (term frequency- inverse document frequency) for hyperlink topic relevance prediction and pages topic relevance calculation. Third, in real crawling experiments on the prototype system, the crawler using TF-IDF has high performance with the accumulated topic relevance increasing quickly at the beginning of crawling, however the crawler using LSI can find more related pages and tunnel through. Fourth, combining their advantages of LSI and TF-IDF, we propose TFIDF+LSI algorithm to guide the crawling. Last, the crawler using TFIDF+LSI performs the same crawl task and demonstrates the combination advantage of TF-IDF and LSI. The experiment suggests that the crawler's performance using TFIDF+LSI is greatly superior to that using either TF-IDF or LSI respectively.
Hongwei Hao, Cui-Xia Mu, Xu-Cheng Yin, Zhi-Bin Wang
SMC3
2011 Exchange rate prediction with non-numerical information
Zhi-Bin Wang, Hongwei Hao, Xu-Cheng Yin, Kaizhu Huang
Neural Comput. Appl.3
2011 FMI image based rock structure classification using classifier combination
Xu-Cheng Yin, Hongwei Hao, Zhi-Bin Wang, Kaizhu Huang
Neural Comput. Appl.1
2010 Ellipse Detection with an Improved Randomized Hough Transform for Intellectual Phacoemulsification Surgery Systems
Hongwei Hao, Xu-Cheng Yin, Zhi-Bin Wang, Kaizhu Huang
ICIP3
2009 Rejection Strategies with Multiple Classifiers for Handwritten Character Recognition
abstract
With rejection strategies in a handwriting recognition system, we are able to improve the reliability and accuracy of the recognized characters. In this paper, we propose several rejection strategies with multiple classifiers for handwritten character recognition. First, the rejection strategy for the single classifier is introduced, which is composed of three stages: initial scaling, confidence measure calculation, and rejection performing. Then, we analyze rejection strategies for multiple classifiers. We divided our rejection strategies into two categories: (1) for voting combination; and (2) for linear combination with multiple classifiers. In the voting combination style, three rejection strategies, OR, AND, and VOTING, are proposed. And for the linear combination one, rejection strategies for average and weighted combination are analyzed respectively. We also experiment and compare our rejection strategies with handwritten digit recognition.
Xu-Cheng Yin, Hongwei Hao, Yun-Feng Tang, Jun Sun 0004, Satoshi Naoi
ICDAR1
2009 Exchange Rate Forecasting Using Classifier Ensemble
Zhi-Bin Wang, Hongwei Hao, Xu-Cheng Yin, Kaizhu Huang
ICONIP (1)3
2009 A Rock Structure Recognition System Using FMI Images
Xu-Cheng Yin, Hongwei Hao, Zhi-Bin Wang, Kaizhu Huang
ICONIP (1)1
2007 A Multi-Stage Strategy to Perspective Rectification for Mobile Phone Camera-Based Document Images
abstract
Document images captured by a mobile phone camera often have perspective distortions. Efficiency and accuracy are two important issues in designing a rectification system for such perspective documents. In this paper, we propose a new perspective rectification system based on vanishing point detection. This system achieves both the desired ef- ficiency and accuracy using a multi-stage strategy: at the first stage, document boundaries and straight lines are used to compute vanishing points; at the second stage, text base- lines and block aligns are utilized; and at the last stage, character tilt orientations are voted for the vertical vanish- ing point. A profit function is introduced to evaluate the reliability of detected vanishing points at each stage. If van- ishing points at one stage are reliable, then rectification is ended at that stage. Otherwise, our method continues to seek more reliable vanishing points in the next stage. We have tested this method with more than 400 images includ- ing paper documents, signboards and posters. The image acceptance rate is more than 98.5% with an average speed of only about 60ms.
Xu-Cheng Yin, Jun Sun 0004, Satoshi Naoi, Katsuhito Fujimoto, Yusaku Fujii, Koji Kurokawa, Hiroaki Takebe
ICDAR1
2005 Financial Document Image Coding with Regions of Interest Using JPEG2000
abstract
Document image coding is a very important issue in document analysis and recognition systems provided with vast samples. An image compression algorithm with regions of interest (ROIs) using JPEG2000 is proposed for financial document images which have various categories, complex layouts, and irregular noises. Three types of ROIs: filled information ROIs, seal ROIs, and handwriting ROIs, are detected and extracted through document knowledge analysis and handwriting identification. The first ROIs are detected by document classification, the second are extracted by connected component analysis based on color and shape information, and the third are located by handwriting identification using an incremental Fisher linear discriminant classifier. A ROI mask with a random shape is constructed by thresholding and merging these ROIs. Finally, a financial document image is encoded using JPEG2000 Part I with this ROI mask. Compared to JPEG and DjVu, the method improves visual quality while decreasing storing space.
Xu-Cheng Yin, Chang-Ping Liu, Zhi Han
ICDAR1
2005 Feature combination using boosting
Xu-Cheng Yin, Chang-Ping Liu, Zhi Han
Pattern Recognit. Lett.1