Xu-Cheng Yin

dblp:70/4187 · also Xucheng Yin · DBLP profile ↗
← Back
25ranked-venue papers in the field
5as first author
10since 2021 · last 2026
0000-0003-0023-0220ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 14 (4 first)Information Retrieval & Web Search · 7 (1 first)Database Systems & Data Management · 2Knowledge Engineering, Semantic Web & Information Systems · 2
YearPublicationVenuePosition
2026 Querying Historical $k$-Dense Subgraphs on Temporal Graphs
Yalong Zhang, Xu-Cheng Yin, Guoren Wang
ICDE4
2026 Beyond Optical Flow: Latent Micro-Motion as Visual Evidence in UAV Video
abstract
Optical flow has long dominated motion representation in video by focusing on explicit pixel displacement caused by object or camera movement. In this work, we argue that video also contains a largely overlooked form of motion information, namely latent micro-motion, which arises from subtle, structure-constrained responses of rigid components to physical interaction with the environment. We study this phenomenon in UAV video, a physically grounded setting where onboard structures are continuously exposed to aerodynamic forces. Although such micro-motions are low in amplitude and are often treated as noise or residual vibration, we show that they form a consistent visual signal that becomes observable through structure-aware and temporally aggregated analysis, even when using simple segmentation and coarse motion descriptors. Through an exploratory analysis, we demonstrate that micro-motion patterns exhibit clear structure and respond systematically to changes in wind conditions, with particularly strong sensitivity to wind direction and weaker dependence on wind magnitude in the examined scenarios. These observations suggest that micro-motion constitutes a distinct regime of motion information in video, complementary to explicit displacement, and motivate a broader reconsideration of how motion is represented and exploited in physically grounded multimedia scenarios.
Bowen Zhang 0011, Song-Lu Chen, Xiaobin Zhu 0001, Xu-Cheng Yin
ICMR6
2026 Amplitude-Phase Reconstruction for Non-Stationary Time-Series Forecasting
abstract
Real-world time-series can be decomposed into multiple interacting frequency components whose amplitudes and phases co-evolve over time. Such coupled dynamics can deform spectral trajectories, leading to pronounced non-stationarity in frequency-domain representations and substantial degradation in long-horizon forecasting performance. However, most existing frequency-domain forecasting methods do not explicitly model amplitude–phase interactions and often implicitly assume globally consistent spectral structures, which limits their ability to capture drifting spectra under non-stationary dynamics. To address this challenge, we propose a novel Amplitude–Phase Reconstruction Network (APRNet), a frequency-domain time-series forecasting framework that models amplitude–phase interactions from temporal and channel-wise perspectives. Specifically, we propose an innovative Amplitude-Phase Global Correlation (APGC) module to capture spectrum-wide amplitude-phase dependencies and derive frequency-wise calibration factors under non-stationary dynamics. The calibrated spectral representations are further reconstructed into the time domain, forming a closed time-frequency reconstruction loop that suppresses spectral drift and mitigates non-stationarity in temporal features. In addition, we propose a novel Discrete Kolmogorov–Arnold Network (D-KAN) to further enhance fine-grained amplitude–phase modeling. By combining discretized nonlinear activations with locally supported B-spline basis functions, D-KAN enables frequency-adaptive piecewise nonlinear modeling, improving sensitivity to localized spectral variations and enhancing high-frequency expressiveness. Extensive experiments verify the superior performance of our APRNet. Our codes are available at:https://github.com/LH325/APRNet.
Xiaobin Zhu 0001, Lirui Deng 0001, Xu-Cheng Yin
IEEE Trans. Knowl. Data Eng.5
2025 T-LLaVA: An Effective Saliency-Aware Slicing Strategy for Text Recognition
Mengze Wei, Xiaobin Zhu 0001, Xu-Cheng Yin
ICDAR (1)6
2024 Multi-task Learning for License Plate Recognition in Unconstrained Scenarios
Zhen-Lun Mo, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
ICDAR (1)5
2024 Transformer-based Reasoning for Learning Evolutionary Chain of Events on Temporal Knowledge Graph
abstract
Temporal Knowledge Graph (TKG) reasoning often involves completing missing factual elements along the timeline. Although existing methods can learn good embeddings for each factual element in quadruples by integrating temporal information, they often fail to infer the evolution of temporal facts. This is mainly because of (1) insufficiently exploring the internal structure and semantic relationships within individual quadruples and (2) inadequately learning a unified representation of the contextual and temporal correlations among different quadruples. To overcome these limitations, we propose a novel Transformer-based reasoning model (dubbed ECEformer) for TKG to learn the Evolutionary Chain of Events (ECE). Specifically, we unfold the neighborhood subgraph of an entity node in chronological order, forming an evolutionary chain of events as the input for our model. Subsequently, we utilize a Transformer encoder to learn the embeddings of intra-quadruples for ECE. We then craft a mixed-context reasoning module based on the multi-layer perceptron (MLP) to learn the unified representations of inter-quadruples for ECE while accomplishing temporal knowledge reasoning. In addition, to enhance the timeliness of the events, we devise an additional time prediction task to complete effective temporal information within the learned unified representation. Extensive experiments on six benchmark datasets verify the state-of-the-art performance and the effectiveness of our method.
Zhiyu Fang, Shuai-Long Lei, Xiaobin Zhu 0001, Shi-Xue Zhang, Xu-Cheng Yin, Jingyan Qin
SIGIR6
2023 End-to-End Multi-line License Plate Recognition with Cascaded Perception
Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
ICDAR (5)4
2023 Open-Set Text Recognition via Shape-Awareness Visual Reconstruction
Chang Liu 0083, Xu-Cheng Yin
ICDAR (6)3
2021 Fast Recognition for Multidirectional and Multi-type License Plates with 2D Spatial Attention
Qi Liu 0041, Song-Lu Chen, Zhen-Jia Li, Feng Chen 0040, Xu-Cheng Yin
ICDAR (4)6
2021 Dynamic Receptive Field Adaptation for Attention-Based Text Recognition
Haibo Qin, Xiaobin Zhu 0001, Xu-Cheng Yin
ICDAR (2)4
2020 Ranking via partial ordering for answer selection
Zanxia Jin, Bowen Zhang 0011, Jingyan Qin, Xu-Cheng Yin
Inf. Sci.5
2019 Detecting Text in News Images with Similarity Embedded Proposals
abstract
Text extraction plays an important role in news images analysis tasks. However, the conglutination of subtitles and station logos makes text detection challenging. In this paper, we develop an effective news text detection framework by introducing a novel similarity embedded proposal mechanism. The main idea is to predict similarity for each fine-scale coarse proposal to help construct text bounding boxes. Specifically, a CNN and bi-directional LSTM based network is used to produce vectors embedded in coarse proposals provided by Connectionist Text Proposal Network (CTPN). Notably, similarity embedded proposal mechanism can be generalized to other sub-text level text detection models. Comparing to the state-of-the-art method (CTPN), our framework improves F-measure by 25.2% on our Private News Dataset and 8.9% on ICDAR 2013 benchmarks, respectively.
Miaotong Jiang, Jie-Bo Hou, Xiaobin Zhu 0001, Xu-Cheng Yin
ICDAR5
2018 TED-KISS: A Known-Item Speech Video Search Benchmark
abstract
Known-item search is an everyday natural scenario that we search for a specific thing (maybe a song) while only remembering some details about it. Existing benchmarks generally focus on brief user requests which specify some metadata like the title, or the time. However, in most cases, the users can hardly recall such information accurately. In order to embrace the research of known-item search, we present a new publicly available known-item speech video search benchmark, namely TED-KISS, which takes TED talks as an example. The video collection is constructed with up-to-date nearly 80,000 TED and TEDx talks on Youtube. These talks cover various topics, and their titles, speakers, descriptions, full-text subtitles, as well as original links are extracted as metadata, which makes the researches on text-based retrieval and multimedia retrieval feasible. Unlike other benchmarks concerning visual contents in segments, the user requests in TED-KISS are generated through a more natural process, partly through original related topics posted on Reddit and Baidu Tieba, and partly through manual imitative requests annotated by volunteers in a scenario simulation. In addition, we analyze the characteristics of our benchmark through evaluations of several existing text-based IR and Neural-IR models, which also can be served as baselines for this task.
Fan Fang, Bowen Zhang 0011, Xu-Cheng Yin, Haixia Man
CIKM3
2017 ICDAR2017 Robust Reading Challenge on Text Extraction from Biomedical Literature Figures (DeTEXT)
abstract
Hundreds of millions of figures are available in the biomedical literature, representing important biomedical experimental evidence. Since text is a rich source of information in figures, automatically extracting such text may assist in the task of mining figure information and understanding biomedical documents. Unlike images in the open domain, biomedical figures present a variety of unique challenges. For example, biomedical figures typically have complex layouts, small font sizes, short text, specific text, complex symbols and irregular text arrangements. This paper presents the final results of the ICDAR 2017 Competition on Text Extraction from Biomedical Literature Figures (ICDAR2017 DeTEXT Competition), which aims at extracting (detecting and recognizing) text from biomedical literature figures. Similar to text extraction from scene images and web pictures, ICDAR2017 DeTEXT Competition includes three major tasks, i.e., text detection, cropped word recognition and end-to-end text recognition. Here, we describe in detail the data set, tasks, evaluation protocols and participants of this competition, and report the performance of the participating methods.
Xu-Cheng Yin, Dimosthenis Karatzas
ICDAR2
2017 Building Your Own Reading List Anytime via Embedding Relevance, Quality, Timeliness and Diversity
abstract
During every summer holidays, several editions of reading lists are recommended and emerged on mass media, e.g., New York Times, and BBC. However, these reading lists are built for whole people with general topics for some purposes. What if we expect the books of a specific topic at a specific moment? How to generate the requested reading list for our own automatically? In this paper, we propose a searching framework for building a topical reading list anytime, where the Relevance (between topics and books), Quality (of books), Timeliness (of popularities) and Diversity (of results) are embedded into vector representations respectively based on user-generated contents and statistics on social media. We collected 8,197 real-world topics from 198 diverse groups on Librarything.com. The proposed methods are evaluated on the topic collection and the public benchmarks Social Book Search 2012-2016 (SBS). Experimental results demonstrate the robustness and effectiveness of our framework.
Bowen Zhang 0011, Xu-Cheng Yin, Jian-Lin Jin
SIGIR2
2016 A Short Survey of Recent Advances in Graph Matching
abstract
Graph matching, which refers to a class of computational problems of finding an optimal correspondence between the vertices of graphs to minimize (maximize) their node and edge disagreements (affinities), is a fundamental problem in computer science and relates to many areas such as combinatorics, pattern recognition, multimedia and computer vision. Compared with the exact graph (sub)isomorphism often considered in a theoretical setting, inexact weighted graph matching receives more attentions due to its flexibility and practical utility. A short review of the recent research activity concerning (inexact) weighted graph matching is presented, detailing the methodologies, formulations, and algorithms. It highlights the methods under several key bullets, e.g. how many graphs are involved, how the affinity is modeled, how the problem order is explored, and how the matching procedure is conducted etc. Moreover, the research activity at the forefront of graph matching applications especially in computer vision, multimedia and machine learning is reported. The aim is to provide a systematic and compact framework regarding the recent development and the current state-of-the-arts in graph matching.
Junchi Yan, Xu-Cheng Yin, Weiyao Lin, Cheng Deng 0002, Hongyuan Zha, Xiaokang Yang 0001
ICMR2
2016 A generic pseudo relevance feedback framework with heterogeneous social information
Bowen Zhang 0011, Xu-Cheng Yin
Inf. Sci.2
2015 Multi-strategy tracking based text detection in scene videos
abstract
Text detection and tracking in scene videos are important prerequisites for content-based video analysis and retrieval, wearable camera systems and mobile devices augmented reality translators. Here, we present a novel multi-strategy tracking based text detection approach in scene videos. In this approach, a state-of-the-art scene text detection module [1] is first used to detect text in each video frame. Then a multi-strategy text tracking technique is proposed, which uses tracking by detection, spatio-temporal context learning, and linear prediction to predict the candidate text location sequentially, and adaptively integrates and selects the best matching text block from the candidate blocks with a rule-based method. This multi-strategy tracking technique can combine the advantages of the three different tracking techniques and afterwards make remedies to the disadvantages of them. Experiments on a variety of scene videos show that our proposed approach is effective and robust to reduce false alarm and improve the accuracy of detection.
Ze-Yu Zuo, Shu Tian, Wei-Yi Pei, Xu-Cheng Yin
ICDAR4
2014 Social Book Search Reranking with Generalized Content-Based Filtering
abstract
Semantically searching and navigating products (e.g., on Taobao.com or Amazon.com) with professional metadata and user-generated content from social media is a hot topic in information retrieval and recommendation systems, while most existing methods are specifically designed as a purely searching system. In this paper, taking Social Book Search as an example, we propose a general search-recommendation hybrid system for this topic. Firstly, we propose a Generalized Content-Based Filtering (GCF) model. In this model, a preference value, which flexibly ranges from 0 to 1, is defined to describe a user's preference for each item to be recommended, unlike conventionally using a set of preferable items. We also design a weighting formulation for the measure of recommendation. Next, assuming that the query in a searching system acts as a user in a recommendation system, a general reranking model is constructed with GCF to rerank the initial resulting list by utilizing a variety of rich social information. Afterwards, we propose a general search-recommendation hybrid framework for Social Book Search, where learning-to-rank is used to adaptively combine all reranking results. Finally, our proposed system is extensively evaluated on the INEX 2012 and 2013 Social Book Search datasets, and has the best performance ([email protected]) on both datasets compared to other state-of-the-art systems. Moreover, our system recently won the INEX 2014 Social Book Search Evaluation.
Bowen Zhang 0011, Xu-Cheng Yin, Xiao-Ping Cui, Jiao Qu, Bin Geng, Hongwei Hao
CIKM2
2013 Sorting-Based Dynamic Classifier Ensemble Selection
abstract
In ensemble learning, a higher accuracy can be achieved by integrating some classifiers instead of all the classifiers. But, it is very difficult to select the best classifier combination which can be seen as an optimization problem, from a pool of classifiers. To deal with this problem, we propose a new classifier selection method, Sorting-based Dynamic Classifier Ensemble Selection (SDES), which consists of two stages: (1) classifier sorting, and (2) dynamic ensemble selection on sorted classifier sequence. In the first stage, classifiers are sorted based on diversity, to avoid searching for the nearest neighbors in dynamic ensemble selection methods and greatly improve the selection efficiency. In the second stage, the optimal subset of classifiers is selected from the sorted classifier sequence based on confidence of test samples, to guarantee high accuracy of the optimal classifier subset. Experimental results have shown the effectiveness and high efficiency of the proposed method.
Yan Yan 0004, Xu-Cheng Yin, Zhi-Bin Wang, Xuwang Yin, Hongwei Hao
ICDAR2
2013 Accurate and robust text detection: a step-in for text retrieval in natural scene images
abstract
We propose and implement a robust text detection system, which is a prominent step-in for text retrieval in natural scene images or videos. Our system includes several key components: (1) A fast and effective pruning algorithm is designed to extract Maximally Stable Extremal Regions as character candidates using the strategy of minimizing regularized variations. (2) Character candidates are grouped into text candidates by the single-link clustering algorithm, where distance weights and threshold of clustering are learned automatically by a novel self-training distance metric learning algorithm. (3) The posterior probabilities of text candidates corresponding to non-text are estimated with an character classifier; text candidates with high probabilities are then eliminated and finally texts are identified with a text classifier. The proposed system is evaluated on the ICDAR 2011 Robust Reading Competition dataset and a publicly available multilingual dataset; the f measures are over 76% and 74% which are significantly better than the state-of-the-art performances of 71% and 65%, respectively.
Xu-Cheng Yin, Xuwang Yin, Kaizhu Huang, Hongwei Hao
SIGIR1
2011 Robust Vanishing Point Detection for MobileCam-Based Documents
abstract
Document images captured by a mobile phone camera often have perspective distortions. In this paper, fast and robust vanishing point detection methods for such perspective documents are presented. Most of previous methods are either slow or unstable. Based on robust detection of text baselines and character tilt orientations, our proposed technology is fast and robust with the following features: (1) quick detection of vanishing point candidates by clustering and voting on the Gaussian sphere space, and (2) precise and efficient detection of the final vanishing points using a hybrid approach, which combines the results from clustering and projection analysis. The rectified image acceptance rate for Mobile Cam-based documents, signboards and posters is more than 98% with an average speed of about 100ms.
Xu-Cheng Yin, Hongwei Hao, Jun Sun 0004, Satoshi Naoi
ICDAR1
2009 Rejection Strategies with Multiple Classifiers for Handwritten Character Recognition
abstract
With rejection strategies in a handwriting recognition system, we are able to improve the reliability and accuracy of the recognized characters. In this paper, we propose several rejection strategies with multiple classifiers for handwritten character recognition. First, the rejection strategy for the single classifier is introduced, which is composed of three stages: initial scaling, confidence measure calculation, and rejection performing. Then, we analyze rejection strategies for multiple classifiers. We divided our rejection strategies into two categories: (1) for voting combination; and (2) for linear combination with multiple classifiers. In the voting combination style, three rejection strategies, OR, AND, and VOTING, are proposed. And for the linear combination one, rejection strategies for average and weighted combination are analyzed respectively. We also experiment and compare our rejection strategies with handwritten digit recognition.
Xu-Cheng Yin, Hongwei Hao, Yun-Feng Tang, Jun Sun 0004, Satoshi Naoi
ICDAR1
2007 A Multi-Stage Strategy to Perspective Rectification for Mobile Phone Camera-Based Document Images
abstract
Document images captured by a mobile phone camera often have perspective distortions. Efficiency and accuracy are two important issues in designing a rectification system for such perspective documents. In this paper, we propose a new perspective rectification system based on vanishing point detection. This system achieves both the desired ef- ficiency and accuracy using a multi-stage strategy: at the first stage, document boundaries and straight lines are used to compute vanishing points; at the second stage, text base- lines and block aligns are utilized; and at the last stage, character tilt orientations are voted for the vertical vanish- ing point. A profit function is introduced to evaluate the reliability of detected vanishing points at each stage. If van- ishing points at one stage are reliable, then rectification is ended at that stage. Otherwise, our method continues to seek more reliable vanishing points in the next stage. We have tested this method with more than 400 images includ- ing paper documents, signboards and posters. The image acceptance rate is more than 98.5% with an average speed of only about 60ms.
Xu-Cheng Yin, Jun Sun 0004, Satoshi Naoi, Katsuhito Fujimoto, Yusaku Fujii, Koji Kurokawa, Hiroaki Takebe
ICDAR1
2005 Financial Document Image Coding with Regions of Interest Using JPEG2000
abstract
Document image coding is a very important issue in document analysis and recognition systems provided with vast samples. An image compression algorithm with regions of interest (ROIs) using JPEG2000 is proposed for financial document images which have various categories, complex layouts, and irregular noises. Three types of ROIs: filled information ROIs, seal ROIs, and handwriting ROIs, are detected and extracted through document knowledge analysis and handwriting identification. The first ROIs are detected by document classification, the second are extracted by connected component analysis based on color and shape information, and the third are located by handwriting identification using an incremental Fisher linear discriminant classifier. A ROI mask with a random shape is constructed by thresholding and merging these ROIs. Finally, a financial document image is encoded using JPEG2000 Part I with this ROI mask. Compared to JPEG and DjVu, the method improves visual quality while decreasing storing space.
Xu-Cheng Yin, Chang-Ping Liu, Zhi Han
ICDAR1