Xinguo Yu

dblp:06/3654 · DBLP profile ↗
← Back
67ranked-venue papers
31as first author
15since 2021 · last 2026
0000-0001-8379-6742ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 24 first-author · 1 since 2021Artificial intelligence and machine learning · 26 · 6 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 FLAIR: Steering LLM Mathematical Problem Solving based on A Fuzzy-Logic-AssIsted Reasoner
abstract
Hao Wu, Hongru Sun, Wanqing Li, Xinguo Yu, Hao Ming, Xiao Luo, Wenbin Zhang, Jiahong Zhao, Yi Guo, Jie Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hao Wu 0094, Hongru Sun, Wanqing Li 0001, Xinguo Yu, Hao Ming 0001, Xiao Luo 0001, Wenbin Zhang 0002, Jiahong Zhao, Yi Guo 0001, Jie Yang 0009
ACL (1)4
2026 UDfSolver : Uniform representation and decoupled inference strategy for text-diagram function problems solving
Xinguo Yu, Litian Huang
Expert Syst. Appl.2
2026 A theoretical review on solving algebra problems
abstract
Solving algebra problems (APs) continues to attract significant research interest as evidenced by the large number of algorithms and theories proposed over the past decade. Despite these important research contributions, however, the body of work remains incomplete in terms of theoretical justification and scope. The current contribution intends to fill the gap by developing a review framework that aims to lay a theoretical base, create an evaluation scheme, and extend the scope of the investigation. This paper first develops the State Transform Theory (STT), which emphasizes that the problem-solving algorithms are structured according to states and transforms unlike the understanding that underlies traditional surveys which merely emphasize the progress of transforms. The STT, thus, lays the theoretical basis for a new framework for reviewing algorithms. This new construct accommodates the relation-centric algorithms for solving both word and diagrammatic algebra problems. The latter not only highlights the necessity of introducing new states but also allows revelation of contributions of individual algorithms – obscured in prior reviews without this approach. A review of AP solving algorithms (2014 to date) is subsequently enhanced by applying the STT specifically designed to analyze individual and collective algorithms for states and transforms. Furthermore, the State Transform Analysis (STA) is a core function in the identification of progress in terms of states and transforms. Thirdly, the Perspective Confusion Comparison (PCC) is developed to extend the application of STA to add capabilities for systematic and individual evaluation of transforms, algorithms, and approaches. This is a new evaluation method of being different from other methods that can do more in-depth evaluation for decomposed algorithms for solving problems with multiple types of inputs. Finally, this work identifies several research directions by extracting benefit from theoretical reviews. This work significantly contributes to the advancement of AP-solving by providing the mechanism for identifying and understanding the key contributors to building high-performance problem-solving algorithms at the three levels of transform, algorithm, and approach.
Xinguo Yu, Weina Cheng, Chuanzhi Yang
Expert Syst. Appl.1
2026 Global profits, local decisions: Why global cooperation falters in multi-level games
Jinhua Zhao 0002, Xinguo Yu, Cuiling Gu, Xianjia Wang
Expert Syst. Appl.2
2025 A Compact Model for Mathematics Problem Representations Distilled from BERT
abstract
Large language models (LLMs) have made significant advancements in math problem solving, but their large size and high latency render them impractical for real-world applications in intelligent mathematics solvers. Recently, task-agnostic compact models have been developed to replace LLMs in general natural language processing tasks. However, these models often struggle to acquire sufficient math-related knowledge from LLMs, leading to unsatisfactory performance in solving math word problems (MWPs). To develop a specialized compact model for representing MWPs, we develop the knowledge distillation (KD) technique to extract mathematical semantics knowledge from the large pre-trained model BERT. Effective knowledge types and distillation strategies are explored through extensive experiments. Our KD algorithm employs multi-knowledge distillation to extract fundamental knowledge from hidden states in the middle to lower layers, while also incorporating knowledge of mathematical relations and symbol constraints from higher-layer outputs and math decoder outputs, by leveraging bottleneck networks. Pre-training tasks on MWP datasets, such as masked language modeling and part-of-speech tagging, are also utilized to enhance the generalization of the compact model for MWP understanding. Additionally, a simple parameter mixing strategy is employed to prevent catastrophic forgetting of acquired knowledge. Our findings indicate that our approach can reduce the size of a BERT model by 10% while retaining approximately 95% of its performance on MWP datasets, outperforming the mainstream BERT-based task-agnostic compact models. The efficacy of each component has been validated through ablation studies.
Hao Ming 0001, Xinguo Yu, Xiaotian Cheng, Zhenquan Shen, Xiaopan Lyu
AAAI2
2025 A Scene-Attention Relation-Centric Algorithm for Solving Arithmetic Word Problems
Rao Peng, Xinguo Yu, Chuanzhi Yang, Xiaopan Lyu
Expert Syst. Appl.2
2025 Can question-texts improve the recognition of handwritten mathematical expressions in respondents' solutions?
Xinxin Jin, Xinzi Peng, Jinzheng Liu, Xinguo Yu
Knowl. Based Syst.7
2025 Bayesian math word problem solvers with credibility level: know what they know and what they do not know
Shengbing Tang, Xinguo Yu
Neural Comput. Appl.3
2025 Physics-informed multi-output Gaussian process for dynamical system modeling
Shengbing Tang, Bin He 0007, Xinguo Yu
Neural Networks3
2024 An offer-generating strategy for multiple negotiations with mixed types of issues and issue interdependency
abstract
Agent negotiation in multi-agent systems has been extensively studied, focusing on both theoretical and applied research. However, a limited number of studies have considered proposing an offer-generating strategy for agents to propose offers during the negotiation process in the multiple-negotiation situation where interdependency exist between a mixture of discrete issues and continuous issues across different negotiations. Especially, considering the above common real-life situation, there is little work of proposing such a strategy which is able to generate an approximately Pareto optimal solution . To address such challenges, this paper targets at multiple-negotiation scenarios involving interdependency between mixed types of issues across different negotiations. The contributions of this paper are threefold. Firstly, this paper addresses the research gap in mixed-type of issues in multiple negotiations. Secondly, the paper introduces a formalized negotiation model for multiple-negotiation scenarios, addressing both discrete and continuous issues, enabling automatic agents to obtain goal-aligned offers effectively. Thirdly, this paper introduces a Hybrid of PSO (Particle Swarm Optimization) and GA (Genetic Algorithm) Algorithm (i.e., named as HPGA in this paper) as an offer-generating strategy to assist agents in achieving approximately Pareto optimization in multiple-negotiation scenarios. To support those claims, this paper presents an overall modeling framework, introduces the proposed offer-generation strategy, conducts a series of experiments to demonstrate the superiority of the proposed approach in this paper, and presents two realistic case studies . Overall, this research expands upon existing studies in agent-based negotiation by addressing the overlooked aspects of mixed types of issues and issue interdependency across multiple negotiations. The proposed modeling approach and offer-generation strategy contribute to the advancement of negotiation techniques in multi-agent systems.
Lei Niu, Fenghui Ren, Xinguo Yu
Eng. Appl. Artif. Intell.4
2024 A reputation-aided negotiation mechanism for multi-agent society based on blockchain
Lei Niu, Qihang Cai, Fenghui Ren, Xinguo Yu
Eng. Appl. Artif. Intell.5
2024 Elite GA-based feature selection of LSTM for earthquake prediction
Zhiwei Ye, Wuyang Lan, Wen Zhou 0007, Qiyi He, Xinguo Yu, Yunxuan Gao
J. Supercomput.6
2023 A Numeracy-Enhanced Decoding for Solving Math Word Problem
Rao Peng, Chuanzhi Yang, Litian Huang, Xiaopan Lyu, Xinguo Yu
NLPCC (3)6
2022 Improving Oracle Bone Characters Recognition via A CycleGAN-Based Data Augmentation Method
Ting Zhang 0008, Xinxin Jin, Harold Mouchère, Xinguo Yu
ICONIP (6)6
2022 A memetic algorithm based on edge-state learning for max-cut
abstract
Max-cut is one of the most classic NP-hard combinatorial optimization problems . The symmetry nature of it leads to special difficulty in extracting meaningful configuration information for learning; none of the state-of-the-art algorithms has employed any learning operators. This paper proposes an original learning method for max-cut, namely post-flip edge-state learning (PF-ESL). Different from previous algorithms, PF-ESL regards edge-states (cut or not cut) rather than vertex-positions as the critical information of a configuration, and extracts their statistics over a population for learning. It is based on following observations. 1) Edges are the only factors considered by the objective function. 2) Edge-states keep invariant when rotating a local configuration to its symmetry position, but vertex-positions do not. These suggest that edge-states contain more meaningful information about a configuration than vertex-positions do. It is impossible to set the state of an edge without influencing some other edges’ states due to their dependencies. Therefore, instead of setting edge-states directly, PF-ESL samples the flips on vertices. Flips on vertices are sampled according to their capacities in increasing the similarity on edge-states between the given solution and a population. PF-ESL is employed in an EDA (Estimation of Distribution Algorithm) perturbation operator and a path-relinking operator. Experimental results show that our algorithm is competitive, and show that edge-state learning is value-added for both the two operators. The main contributions of this paper are as follows. Firstly, previous state-of-the-art evolutionary algorithms for max-cut focus on vertex positions in their evolutionary operation, this paper proposes a new and more reasonable perspective suggesting that edge-states are the critical information of divided graphs rather than vertex positions, and introduces a novel method to measure and utilize their similarities based on it. Such a perspective is fundamental to learning based algorithms design for max-cut and other graph partitioning problems, and can shed lights on future researches. Furthermore, since max-cut is one of the most classic and fundamental NP hard problems, many real-world problems involve dividing graph data into different parts to optimize certain functions, this new perspective may inspire related or similar problems. Secondly, besides the original edge-states based perspective, and the post-flip edge-states learning (PFESL) operator based on it, our memetic algorithm also incorporates a novel evolutionary framework which alternates between EDA based Iterated Tabu search (ITS) and path relinking based genetic algorithm . Finally, the proposed algorithm provides competitive results on two mostly used benchmark sets and improves the best-known results of 6 most challenging instances.
Zhizhong Zeng, Zhipeng Lü, Xinguo Yu, Qinghua Wu 0002, Yang Wang 0098
Expert Syst. Appl.3
2020 A relation based algorithm for solving direct current circuit problems
Bin He 0007, Xinguo Yu, Pengpeng Jian
Appl. Intell.2
2019 AZUPT: Adaptive Zero Velocity Update Based on Neural Networks for Pedestrian Tracking
abstract
Zero Velocity Update (ZUPT) has played a key role in Pedestrian Dead Reckoning (PDR) with inertial measurement units (IMU). However, it is both crucial and difficult to determine ZUPT conditions given complex and varying motion types such as walking, fast walking or running, and different walking habits of distinct people, which have direct and significant impact on the tracking accuracy. In this research we proposed a model based on deep neural networks to determine moments when the ZUPT should be conducted. The proposed model ensures nearly identical performance regardless of different motion types. It has been demonstrated by extensive experiments conducted in three different scenarios that our model can work equally well with different pedestrians and walking patterns, enabling the wide use of PDR in real-world applications.
Xinguo Yu, Xinyue Lan, Zhuoling Xiao, Shuisheng Lin, Bo Yan 0007
GLOBECOM1
2019 An End-to-End Trainable System for Offline Handwritten Chemical Formulae Recognition
abstract
In this paper, we propose an end-to-end trainable system for recognizing handwritten chemical formulae. This system recognize once a time a chemical formula, instead of one chemical symbol or a whole chemical equation, which is in line with people's writing habits, at the same time could help to develop methods for the complicated chemical equations recognition. The proposed system adopts the CNN+RNN+CTC framework, which is one of state of the art methods in imagebased sequence labelling tasks. We extend the capability of the CNN+RNN+CTC framework to interpret 2D spatial relationships (such as 'subscript' existing in chemical formula) by introducing additional labels to represent them. The system evaluated on a self-collected data set of 12,224 samples, achieves the recognition rate of 94.98% at the chemical formula level.
Xinguo Yu
ICDAR3
2019 Double Channel 3D Convolutional Neural Network for Exam Scene Classification of Invigilation Videos
Wu Song, Xinguo Yu
PSIVT2
2019 High-Resolution Realistic Image Synthesis from Text Using Iterative Generative Adversarial Network
Anwar Ullah, Xinguo Yu, Hafiz ur Rahman, Muhammad Farhan Mughal
PSIVT2
2019 Automatically Proving Plane Geometry Theorems Stated by Text and Diagram
abstract
This paper presents an algorithm for proving plane geometry theorems stated by text and diagram in a complementary way. The problem of proving plane geometry theorems involves two challenging subtasks, being theorem understanding and theorem proving. This paper proposes to consider theorem understanding as a problem of extracting relations from text and diagram. A syntax–semantics (S2) model method is proposed to extract the geometric relations from theorem text, and a diagram mining method is proposed to extract geometry relations from diagram. Then, a procedure is developed to obtain a set of relations that is consistent with the given theorem with high confidence. Finally, theorem proving is conducted by using the existing proving methods which take the extracted geometric relations as input. The experimental results show that the proposed theorem proving algorithm can prove 86% of plane geometry theorems in the test dataset of 200 theorems, which is all the theorems in the popular textbook. The proposed algorithm outperforms the existing algorithms mainly because it can extract relations not only from text but also from diagram.
Wenbin Gan, Xinguo Yu, Mingshu Wang
Int. J. Pattern Recognit. Artif. Intell.2
2019 An End-to-End Algorithm for Solving Circuit Problems
abstract
This paper presents an end-to-end algorithm for solving circuit problems in secondary physics. A key challenge in solving circuit problems is to automatically understand circuit problems over the modals of both text and schematic. Existing methods have a limited capacity in problem understanding due to the they cannot deal with the numerous expressions of problems in natural language and the various circuit diagrams. In fact that this paper, a batch of methods is proposed to work against the challenge of solving circuit problems. The problem understanding is modeled as a problem of relation extraction and a scheme is proposed to extract relations from both text and schematic. A syntax–semantics model is adopted to extract explicit relations from text, whereas a unit-theorem-based method is proposed to extract implicit relations. And a mesh search method is proposed to extract relations from schematic. Based on the result of problem understanding, an algorithm is proposed to produce the solutions of circuit problems, in which the solutions are presented in a readable way. The experimental results demonstrate the effectiveness of the proposed algorithm in solving circuit problems. To the best of our knowledge, this paper is the first literature which reports the quantitative results in understanding and solving circuit problems.
Pengpeng Jian, Xinguo Yu, Bin He 0007
Int. J. Pattern Recognit. Artif. Intell.3
2019 A Framework for Solving Explicit Arithmetic Word Problems and Proving Plane Geometry Theorems
abstract
This paper presents a framework for solving math problems stated in a natural language (NL) and applies the framework to develop algorithms for solving explicit arithmetic word problems and proving plane geometry theorems. We focus on problem understanding, that is, the transformation of a NL description of a math problem to a formal representation. We view this as a relation extraction problem, and adopt a greedy algorithm to extract the mathematical relations using a syntax-semantics model, which is a set of patterns describing how a syntactic pattern is mapped to its formal semantics. Our method yields a human readable solution that shows how the mathematical relations are extracted one at a time. We apply our framework to solve arithmetic word problems and prove plane geometry theorems. For arithmetic word problems, the extracted relations are transformed into a system of equations, and the equations are then solved to produce the solution. For plane geometry theorems, these extracted relations are input to an inference system to generate the proof. We evaluate our approach on a set of arithmetic word problems stated in Chinese, and two sets of plane geometry theorems stated in Chinese and English. Our algorithms achieve high accuracies on these datasets and they also show some desirable properties such as brevity of algorithm description and legibility of algorithm actions.
Xinguo Yu, Mingshu Wang, Wenbin Gan, Bin He 0007
Int. J. Pattern Recognit. Artif. Intell.1
2017 Multimodal Prediction of Affective Dimensions via Fusing Multiple Regression Techniques
Dong-Yan Huang, Wan Ding, Huaiping Ming, Minghui Dong, Xinguo Yu, Haizhou Li 0001
INTERSPEECH6
2017 Understanding Plane Geometry Problems by Integrating Relations Extracted from Text and Diagram
Wenbin Gan, Xinguo Yu, Bin He 0007, Mingshu Wang
PSIVT2
2017 A New Scheme for QoE Management of Live Video Streaming in Cloud Environment
Dheyaa Jasim Kadhim, Xinguo Yu, Saba Qasim Jabbar, Wenxing Luo
PSIVT2
2017 Automatic Problem Understanding from Circuit Schematics
Xinguo Yu, Pengpeng Jian, Bin He 0007
PSIVT1
2016 Audio and face video emotion recognition in the wild using deep neural networks and small datasets
abstract
This paper presents the techniques used in our contribution to Emotion Recognition in the Wild 2016’s video based sub-challenge. The purpose of the sub-challenge is to classify the six basic emotions (angry, sad, happy, surprise, fear & disgust) and neutral. Compared to earlier years’ movie based datasets, this year’s test dataset introduced reality TV videos containing more spontaneous emotion. Our proposed solution is the fusion of facial expression recognition and audio emotion recognition subsystems at score level. For facial emotion recognition, starting from a network pre-trained on ImageNet training data, a deep Convolutional Neural Network is fine-tuned on FER2013 training data for feature extraction. The classifiers, i.e., kernel SVM, logistic regression and partial least squares are studied for comparison. An optimal fusion of classifiers learned from different kernels is carried out at the score level to improve system performance. For audio emotion recognition, a deep Long Short-Term Memory Recurrent Neural Network (LSTM-RNN) is trained directly using the challenge dataset. Experimental results show that both subsystems individually and as a whole can achieve state-of-the art performance. The overall accuracy of the proposed approach on the challenge test dataset is 53.9%, which is better than the challenge baseline of 40.47% .
Wan Ding, Dong-Yan Huang, Weisi Lin, Minghui Dong, Xinguo Yu, Haizhou Li 0001
ICMI6
2016 A framework of timestamp replantation for panorama video surveillance
Xinguo Yu, Shiqian Wu, Wu Song
Multim. Tools Appl.1
2015 Reading Digital Video Clocks
abstract
This paper presents an algorithm for reading digital video clocks reliably and quickly. Reading digital clocks from videos is difficult due to the challenges such as color variety, font diversity, noise, and low resolution. The proposed algorithm overcomes these challenges by using the novel methods derived from the domain knowledge. This algorithm first localizes the digits of a digital video clock and then recognizes the digits representing the time of digital video clock. It is a robust three-step algorithm. The first step is an efficient procedure that directly identifies the region of the second digit at a very low computational cost, which replaces the traditional tedious image processing procedure of identifying the second digit region. The success of the first step mainly leverages on the novel second-pixel periodicity method. Using the acquired second digit region as input, the second step is a clock digit localization procedure. It first acquires the colors of the digits of the digital video clock and performs the color conversion. Then it localizes the remaining clock digits. Finally, the last step is a clock digit recognition procedure. It first employs an enhanced digit-sequence recognition method to robustly recognize the digits on the second; it then adopts a deep learning procedure to recognize the remaining digits. The proposed algorithm is tested on a prepared benchmark of 1000 videos that is publicly available and the experimental results show that it can read digital video clocks with a 100% accuracy at a low computational cost.
Xinguo Yu, Wan Ding, Zhizhong Zeng, Hon Wai Leong
Int. J. Pattern Recognit. Artif. Intell.1
2015 Guest editorial: selected papers from ICIMCS 2012
Zhengjun Zha, Yan Liu 0004, Shin'ichi Satoh 0001, Xinguo Yu, Rainer Lienhart
Multim. Syst.4
2013 A Novel and Robust System for Time Recognition of the Digital Video Clock Using the Domain Knowledge
Xinguo Yu, Tie Rong, Hon Wai Leong
MMM (1)1
2013 An Automatic Timestamp Replanting Algorithm for Panorama Video Surveillance
Xinguo Yu, Wu Song, Bin He 0007
PSIVT1
2012 Vision-based attention estimation and selection for social robot to perform natural interaction in the open world
abstract
In this paper, a novel vision system is proposed to estimate attention of people from rich visual clues for social robot to perform natural interactions with multiple participants in public environments. The vision detection and recognition modules include multi-person detection and tracking, upper-body pose recognition, face and gaze detection, lip motion analysis for speaking recognition, and facial expression recognition. A computational approach is proposed to generate a quantitative estimation of human attention. The vision system is implemented on a robotic receptionist "EVE" and encouraging results have been obtained.
Liyuan Li, Xinguo Yu, Jun Li 0005, Gang S. Wang, Ji Yu Shi, Yeow Kee Tan, Haizhou Li 0001
HRI2
2012 Localization and extraction of the four clock-digits using the knowledge of the digital video clock
Xinguo Yu
ICPR1
2012 Robust Multiperson Detection and Tracking for Mobile Service and Social Robots
abstract
This paper proposes an efficient system which integrates multiple vision models for robust multiperson detection and tracking for mobile service and social robots in public environments. The core technique is a novel maximum likelihood (ML)-based algorithm which combines the multimodel detections in mean-shift tracking. First, a likelihood probability which integrates detections and similarity to local appearance is defined. Then, an expectation-maximization (EM)-like mean-shift algorithm is derived under the ML framework. In each iteration, the E-step estimates the associations to the detections, and the M-step locates the new position according to the ML criterion. To be robust to the complex crowded scenarios for multiperson tracking, an improved sequential strategy to perform the mean-shift tracking is proposed. Under this strategy, human objects are tracked sequentially according to their priority order. To balance the efficiency and robustness for real-time performance, at each stage, the first two objects from the list of the priority order are tested, and the one with the higher score is selected. The proposed method has been successfully implemented on real-world service and social robots. The vision system integrates stereo-based and histograms-of-oriented-gradients-based human detections, occlusion reasoning, and sequential mean-shift tracking. Various examples to show the advantages and robustness of the proposed system for multiperson tracking from mobile robots are presented. Quantitative evaluations on the performance of multiperson tracking are also performed. Experimental results indicate that significant improvements have been achieved by using the proposed method.
Liyuan Li, Shuicheng Yan, Xinguo Yu, Yeow Kee Tan, Haizhou Li 0001
IEEE Trans. Syst. Man Cybern. Part B3
2010 HOG based multi-stage object detection and pose recognition for service robot
abstract
This paper develops a HOG-based multistage approach for object detection and object pose recognition for service robots. This approach makes use of the merits of both multi-class and bi-class HOG-based detectors to form a three-stage algorithm at low computing cost. In the first stage, the multi-class classifier with coarse features is employed to estimate the orientation of a potential target object in the image; in the second stage, a bi-class detector corresponding to the detected orientation with intermediate level features is used to filter out most of false positives; and in the third stage, a bi-class detector corresponding to the detected orientation using fine features is used to achieve accurate detection with low rate of false positives. The training of multi-class and bi-class SVMs with their respective features in different levels is described. Experiments in real-world environments have shown that the proposed method is much more accurate than the detection method as it uses only multi-class detector. The proposed method is also much more efficient than the detection method as it uses a bi-class detector for each possible orientation. The approach works well on the scenarios where the SIFT-based detector may fail. The method can achieve real-time object detection, localization, and pose recognition on a P4 2.4GHz PC.
Xinguo Yu, Liyuan Li, Kah Eng Hoe
ICARCV2
2009 Lift-button detection and recognition for service robot in buildings
abstract
Lift operation is one of critical functions for mobile service robot to move across levels in buildings. Lift operation poses the problem of lift-button detection and recognition for computer vision. This paper presents a framework for vision-based lift-button detection and recognition. This problem is challenging due to reflection and complex background. To achieve the robustness in lift operation, we adopt multiple techniques to combat the challenges. First, we propose a multiple partial models method to increase the robustness of button panel detection. Second, we do button recognition by combining structural inference, Hough transform, and multi-symbol recognition techniques. Lastly, we use ultrasonic distance measure devices to aid the vision of the robot. The experimental results show that our framework can achieve the promising results in recognizing both the internal and external buttons of lift.
Xinguo Yu, Liyuan Li, Kah Eng Hoe
ICIP1
2009 Fall Detection and Alert for Ageing-at-Home of Elderly
Xinguo Yu, Panachit Kittipanya-ngam, How-Lung Eng, Loong Fah Cheong
ICOST1
2009 ML-fusion based multi-model human detection and tracking for robust human-robot interfaces
abstract
A novel stereo vision system for real-time human detection and tracking on a mobile service robot is presented in this paper. The system integrates the individually enhanced stereo-based human detection, HOG-based human detection, color-based tracking, and motion estimation for the robust detection and tracking of humans with large appearance and scale variations in real-world environments. A new framework of maximum likelihood based multi-model fusion is proposed to fuse these four human detection and tracking models according to the detection-track associations in 3D space, which is robust to the possible missed detections, false detections, and duplicated responses from the individual models. Multi-person tracking is implemented in a sequential near-to-far way, which well alleviates the difficulties caused by human-over-human occlusions. Extensive experimental results demonstrate the robustness of the proposed system under real-world scenarios with large variations in lighting conditions, cluttered backgrounds, human clothes and postures, and complex occlusion situations. Significant improvements in human detection and tracking have been achieved. The system has been deployed on six robot butlers to serve drinks, and showed encouraging performance in open ceremony events.
Liyuan Li, Kah Eng Hoe, Shuicheng Yan, Xinguo Yu
WACV4
2009 Automatic camera calibration of broadcast tennis video with applications to 3D virtual content insertion and ball detection and tracking
Xinguo Yu, Nianjuan Jiang, Loong Fah Cheong, Hon Wai Leong, Xin Yan 0001
Comput. Vis. Image Underst.1
2009 Interactive broadcast services for live soccer video based on instant semantics acquisition
Xinguo Yu, Liyuan Li, Hon Wai Leong
J. Vis. Commun. Image Represent.1
2008 Tool-aided semantics acquisition for live soccer video
abstract
The semantics of live sports video has big value in catering the services demanded by users (e.g. live event alert and on-the-fly language selection). This paper develops a tool-aided approach to acquire semantics for live soccer video. First, we develop a gamelog acquisition tool, which not only acquires accurate gamelogs but also gets rid of the heavy manual work of the text-way input. Then we develop algorithms to detect multiple boundaries of each event. The amount of manual work of our approach is very limited thanks to the good attributes of our gamelog acquisition tool: (a). it visualizes and symbolizes names of players, actions, etc; (b). it also requires fewer key strokes and mouse clicks for acquiring event description compared with text-way input. Experimental results show that we only need about 5% of mouse-clicks or key-strokes, 50% of input time, and 10% of file size, compared with the text-way input in English, and that our system can quickly obtain the event multiple boundaries of live soccer video.
Xinguo Yu, Xin Yan 0001
ICME1
2008 Robust time recognition of video clock based on digit transition detection and digit-sequence recognition
abstract
This paper presents an algorithm for robust time recognition of video clock. The existing OCR algorithms cannot recognize time properly due to digits of time are in very low resolution and blur. To confront the challenges of time recognition, our algorithm employs three techniques. The first one is a digit transition detection, which identifies SECOND transit frames. The second is a digit-sequence recognition, which uses the property that digits in clock appear in cycle of 0 to 9 to form digit sequence. The third is an on-the-fly template creation. Informally, the robustness of our algorithm benefits from the facts that both digit transition detection and digit-sequence recognition are more reliable than direct character recognition. Experimental results show that our algorithm can achieve a high accuracy in recognizing time.
Xinguo Yu, Wei San Lee
ICPR1
2007 Accurate and Stable Camera Calibration of Broadcast Tennis Video
abstract
This paper presents an original algorithm for accurate and stable camera calibration of broadcast tennis video (BTV). That frame-data of BTV is often erroneous results in wildly fluctuating camera parameters. To meet this challenge, we propose aframegroupingtechnique, which groups frames together according to camera viewpoint. We then use a group-wise data analysis to obtain more stable parameters. Recognizing the fact that some of these parameters do vary somewhat even if they have a similar camera viewpoint, we further employ aHough-likesearch to tune them, maximizing the reprojection similarity. This two-tiered process gains stability of the camera parameters, and yet ensures large reprojection similarity via the tuning step. The experimental results show that our algorithm is able to acquire accurate camera matrix.
Xinguo Yu, Nianjuan Jiang, Loong Fah Cheong
ICIP (3)1
2007 Trajectory-Based Ball Detection and Tracking in Broadcast Soccer Video with the Aid of Camera Motion Recovery
abstract
This paper presents an enhanced trajectory-based ball detection and tracking algorithm, which acquires 2.5D ball position with the aid of camera motion recovery (CMR). Informally, CMR enhances the algorithm by obtaining better ball candidates and forming longer ball trajectories via computing 2.5D position of the ball. The algorithm in this paper comprises two phases. In the first phase, we achieve CMR by a procedure combining homography and global motion estimation. In the second phase, we employ the trajectory-based procedure to the sequence of transformed frames. The experimental results show that the algorithm presented in this paper achieves higher accuracy in identifying the ball than our previous trajectory-based algorithm. More importantly, the obtained ball position is 2.5D, i.e. the ball projection position.
Xinguo Yu, Xiaoying Tu, Ee-Luang Ang
ICME1
2007 A Framework of Context-Aware Object Recognition for Smart Home
Xinguo Yu, Weimin Huang 0002, Boon Fong Chew, Junfeng Dai
ICOST1
2007 Trajectory-based ball detection and tracking with aid of homography in broadcast tennis video
abstract
Ball-detection-and-tracking in broadcast tennis video (BTV) is a crucial but challenging task in tennis video semantics analysis. Informally, the challenges are due to camera motion and the other causes such as the presence of many ball-like objects and the small size of the tennis ball. The trajectory-based approach proposed by us in our previous papers mainly counteracted the challenges imposed by causes other than camera motion and achieves a good performance. This paper proposes an improved trajectory-based ball detection and tracking algorithm in BTV with the aid of homography, which counteracts the challenges caused by camera motion and bring us multiple new merits. Firstly, it acquires an accurate homography, which transforms each frame into the "standard" frame. Secondly, it achieved higher accuracy of ball identification. Thirdly, it obtains the ball projection position in the real world, instead of ball location in the image. Lastly, it also identifies landing frames and positions of the ball. The experimental results show that the improved algorithm can obtain not only higher accuracy in ball identification and in ball position alike, but also ball landing frames and positions. With the intent of using homography to improve the video-based event detection for smart home we also do some experiments on acquiring the homography for home surveillance video.
Xinguo Yu, Nianjuan Jiang, Ee-Luang Ang
VCIP1
2006 Video Clock Time Reconition Based on Temporal Periodic Pattern Change of the Digit Characters
abstract
A novel Temporal Neighboring Pattern Similarity (TNPS) measure is proposed to recognize the time of a digital clock overlay. TNPS detects the presence of a clock overlay by monitoring the periodic changes of the clock digit, and infers the clock time by its natural transition cycle. Compared to traditional methods such as OCR, this method is faster and more reliable because it converts a pattern recognition problem to a pattern change detection problem. Experiments show the recognition result is promising and accurate. One of the applications for this method is to detect the start time of soccer game for our real time live soccer video highlights and event alerts.
Kong-Wah Wan, Xin Yan 0001, Xinguo Yu, Changsheng Xu
ICASSP (2)4
2006 A system for 3D projected virtual content insertion into broadcast tennis video
abstract
This demonstration presents our system for inserting projected virtual content into broadcast tennis video based on camera matrix acquired. This system can automatically acquire the accurate camera matrix for each frame with the tennis court and insert projected virtual content consistent with camera motion. We achieve the accuracy of camera matrices via the proposed techniques of clip-wise data analysis and Hough-like search.
Xin Yan 0001, Xinguo Yu, Tran Thi Phuong Chi
ACM Multimedia2
2006 Inserting 3D projected virtual content into broadcast tennis video
abstract
The ability to acquire the accurate camera matrix of each frame of a video clip is essential if virtual content is inserted into the images in a believable way. This paper presents our system for inserting projected virtual content into broadcast tennis video based on camera matrix acquired. To achieve the accurate camera matrix, we develop a new algorithm for the 3D camera calibration of broadcast tennis video, which improves the accuracy of camera matrices via the proposed techniques of clip-wise data analysis and Hough-like search. We divide all the camera parameters determining a camera matrix into two categories: clip-varying and frame-varying. For the clip-varying ones, we use a clip-wise data analysis procedure to achieve their accuracy. For the framevarying ones, we use a Hough-like search to tune them for each frame. Preliminary experiments results show that we can seamlessly insert projected virtual content into each frame.
Xinguo Yu, Xin Yan 0001, Tran Thi Phuong Chi, Loong Fah Cheong
ACM Multimedia1
2006 Trajectory-Based Ball Detection and Tracking in Broadcast Soccer Video
abstract
This paper presents a novel trajectory-based detection and tracking algorithm for locating the ball in broadcast soccer video (BSV). The problem of ball detection and tracking in BSV is well known to be very challenging because of the wide variation in the appearance of the ball over frames. Direct detection algorithms do not work well because the image of the ball may be distorted due to the high speed of the ball, occlusion, or merging with other objects in the frame. To overcome these challenges, we propose a two-phase trajectory-based algorithm in which we first generate a set of ball-candidates for each frame, and then use them to compute the set of ball trajectories. Informally, the two key ideas behind our strategy are 1) while it is very challenging to achieve high accuracy in locating the precise location of the ball, it is relatively easy to achieve very high accuracy in locating the ball among a set of ball-like candidates and 2) it is much better to study the trajectory information of the ball since the ball is the "most active" object in the BSV. Once the ball trajectories are computed, the ball locations can be reliably recovered from them. One important advantage of our algorithm is that it is able to reliably detect partially occluded or merged balls in the sequence. Two videos from the 2002 FIFA World Cup were used to evaluate our algorithm. It achieves a high accuracy of about 81% for ball location
Xinguo Yu, Hon Wai Leong, Changsheng Xu, Qi Tian 0002
IEEE Trans. Multim.1
2005 Current and Emerging Topics in Sports Video Processing
abstract
Sports video processing is an interesting topic for research, since the clearly defined game rules in sports provide the rich domain knowledge for analysis. Moreover, it is interesting because many specialized applications for sports video processing are emerging. This paper gives an overview of sports video research, where we describe both basic algorithmic techniques and applications.
Xinguo Yu, Dirk Farin
ICME1
2005 A Player-Possession Acquisition System for Broadcast Soccer Video
abstract
A semi-auto system is developed to acquire player possession for broadcast soccer video, whose objective is to minimize the manual work. This research is important because acquiring player-possession by pure manual work is very time-consuming. For completeness, this system integrates the ball detection-and-tracking algorithm, view classification algorithm, and play/break analysis algorithm. First, it produces the ball locations, play/break structure, and the view classes of frames. Then it finds the touching points based on ball locations and player detection. Next it estimates the touching-place in the field for each touching point based on the view-class of the touching frame. Last, for each touching-point it acquires the touching-player candidates based on the touching-place and the roles of players. The system provides the graphical user interfaces to verify touching-points and finalize the touching-player for each touching-point. Experimental results show that the proposed system can obtain good results in touching-point detection and touching-player candidate inference, which save a lot of time compared with the pure manual way.
Xinguo Yu, Tze Sen Hay, Xin Yan 0001, Chng Eng Siong
ICME1
2005 A gridding Hough transform for detecting the straight lines in sports video
abstract
A gridding Hough transform (GHT) is proposed to detect the straight lines in sports video, which is much faster and requires much less memory than the previous Hough transforms. The GHT uses the active gridding to replace the random point selection in the random Hough transforms because forming the linelets from the actively selected points is easier than from the randomly selected points. Existing straight-line Hough transforms require a lot of resources because they were designed for all kinds of straight lines. Considering the fact that the straight lines interested in sports video are long and sparse, this paper proposes two techniques: the active gridding and linelets process. On account of these two techniques, the proposed GHT is fast and uses little memory. The experimental results show that the proposed GHT is faster than the random Hough transform (RHT) and the standard Hough transform (SHT) by 30% and 700% respectively and achieves a 97.5% recall, higher than those achieved by either the SHT or the RHT.
Xinguo Yu, Hoe Chee Lai, Sophie X. Liu, Hon Wai Leong
ICME1
2004 Event detection based on non-broadcast sports video
Jinjun Wang, Changsheng Xu, Chng Eng Siong, Xinguo Yu, Qi Tian 0002
ICIP4
2004 A trajectory-based ball detection and tracking algorithm in broadcast tennis video
abstract
Ball locations over frames facilitate tennis video analysis to a great extent. But so far no algorithm is able to obtain satisfactory result in locating the ball in broadcast tennis video (BTV). This paper presents a trajectory-based algorithm to detect and track the ball in BTV. Unlike the object-based algorithm, it does not decide whether an object is the ball. Instead it decides whether a candidate trajectory is a ball trajectory. This algorithm is able to obtain ball locations for most frames in a BTV, making use of four cues, namely, (1) an antimodel method to produce ball candidates from each frame, (2) a trajectory-based scheme to generate, identify and extend the ball trajectories from a set of candidates, (3) a method to infer the ball locations according to players' locations and the points of hitting, (4) a method to estimate missing ball locations from known ball locations. The experimental results show that our algorithm obtains the ball locations for above 96% frames in a sufficient accuracy for summarization.
Xinguo Yu, Chern-Horng Sim, Jenny R. Wang, Loong Fah Cheong
ICIP1
2004 A robust Hough-based algorithm for partial ellipse detection in broadcast soccer video
abstract
This work presents a robust Hough-based algorithm for partial slightly oblique ellipse detection in broadcast soccer video. The successful identification of the ellipses significantly facilitate soccer video analysis. The existing standard and various modified ellipse Hough transforms measure a cell in the Hough space as though the ellipse defined by the cell were a complete ellipse. Hence, they are not robust when they are applied to detect the partial ellipses appearing in broadcast soccer video. This paper proposes a new measure function that is able to fairly measure whole and partial ellipses. With the improved measure function, we propose an algorithm to detect the partial ellipses in broadcast soccer video. The proposed algorithm first estimates the target ellipse by using the symmetry of the ellipse and the domain knowledge of soccer video. Then for each estimated ellipse, the algorithm searches around the estimated ellipses to find the ellipse with the highest measured value. Our algorithm is efficient and memory-small, i.e. it overcomes two main problems of the standard ellipse Hough transform. More importantly, the proposed algorithm is much more robust than the existing ellipse Hough transforms. Experimental results show that the proposed algorithm achieves above 96% recall and 100% precision. Our algorithm may be the first one that is able to detect ellipses from commercial video in pseudo real-time.
Xinguo Yu, Hon Wai Leong, Changsheng Xu, Qi Tian 0002
ICME1
2004 A 3D reconstruction and enrichment system for broadcast soccer video
abstract
In this demonstration, we present a 3D reconstruction and enrichment system for broadcast soccer video to enhance the consumers' viewing experience. The system can reconstruct not only the goalmouth scene but also the midfield scene as well. Furthermore, the reconstructed video is enriched by music and illustrations of the video contents. In the reconstructed videos, we have eliminated the ball deformation and unnecessary camera changes though smoothing the camera parameters.
Xin Yan 0001, Xinguo Yu, Tze Sen Hay
ACM Multimedia2
2004 A robust and accumulator-free ellipse hough transform
abstract
The ellipse Hough transform (EHT) is a widely-used technique. Most of the previous modifications to the standard EHT improved either the voting procedure that computes the absolute measure function (AMF) or the peak detection of the AMF. However, existing EHTs are not robust for detecting partial slightly-oblique ellipses. This paper presents a Robust and Accumulator-Free Ellipse Hough Transform (RAF-EHT), an improved EHT that is robust even for partial slightly-oblique ellipses. Our RAF-EHT is based on two main ideas, namely, (1) an improved measure function (IMF) for handling the partiality and the obliqueness of ellipses, (2) a new accumulator-free computation scheme for finding the top k peaks of the IMF, without complex peak detection. Experimental results show that the RAF-EHT is more robust than the existing EHTs in detecting the partial slightly-oblique ellipses. In addition, the RAF-EHT needs only a little memory because it is accumulator-free.
Xinguo Yu, Hon Wai Leong, Changsheng Xu, Qi Tian 0002
ACM Multimedia1
2004 3D reconstruction and enrichment of broadcast soccer video
abstract
Recently, it has become a new trend to reconstruct sports video for various purposes. This paper presents a 3D reconstruction and enrichment system that not only reconstructs broadcast soccer video but also enriches reconstructed video with music and illustrations of the video contents. The system can reconstruct not only the goalmouth scene but also the midfield scene, which cannot be reconstructed by the existing systems. To quickly find the feature points for calibrating the camera, we propose a fast algorithm to detect the lines in the goalmouth scene and use the algorithm proposed in our previous papers to detect the partial ellipses in the midfield scene. The reconstruction is conducted on several video sequences of two scenes. The reconstructed videos eliminate the ball deformation and unnecessary camera changes through smoothing the camera parameters. This system also serves as an experimental system for our project that reconstructs the on-going soccer game in real time.
Xinguo Yu, Xin Yan 0001, Tze Sen Hay, Hon Wai Leong
ACM Multimedia1
2003 Real-time camera field-view tracking in soccer video
abstract
Soccer video content-based analysis remains a challenging problem due to the lack of structure in a soccer game. To automate game and tactic analysis, we need to detect and track important activities such as ball possession in a soccer video that is highly correlated to the camera's field-view. In this paper, we present a system that tracks the camera's field-view in a soccer video in real-time. It utilizes a host of content-based visual cues that are obtained by independent threads running in parallel. The result is visualized as an active rectangular bounding box that approximates the camera's field of view superimposed on a virtual soccer field. Experimental results show that the system can reliably track the camera field-view as the game progresses.
Kong-Wah Wan, Joo-Hwee Lim, Changsheng Xu, Xinguo Yu
ICASSP (3)4
2003 A novel ball detection framework for real soccer video
abstract
Despite a lot of research efforts in sports video analysis, soccer video indexing remains a challenging task due to the lack of structure in a soccer game that could help in structure analysis. In particular, little work was done in detecting and tracking the ball whose trajectory could play a crucial role for detecting key events. We propose a novel framework for accurately detecting the ball for broadcast soccer video. Our framework combines both direct and indirect insights to identify the ball rather than conventional simple template matching methods. It has three key components. First we infer the ball size range from the player size. Next non-ball objects are removed to reduce the possible ball candidates. Last but not least, a Kalman filer-based procedure mines candidate trajectories in candidate feature images. Then, a procedure selects the reliable ball trajectories from them. The experimental results on two 1000-frame sequences confirm that the proposed framework is very effective and obtain a better result than existing methods.
Xinguo Yu, Qi Tian 0002, Kong-Wah Wan
ICME1
2003 A ball tracking framework for broadcast soccer video
abstract
It is challenging to detect and track the ball from the broadcast soccer video. The feature-based tracking method to judge if a sole object is a target are inadequate because the features of the balls change fast over frames and we cannot differ the ball from other objects by them. This paper proposes a new framework to find the ball position by creating and analyzing the trajectory. The ball trajectory is obtained from the candidate collection by use of the heuristic false candidate reduction, the Kalman filter-based trajectory mining, and the trajectory evaluation. The ball trajectory is extended via a localized Kalman filter-based model matching procedure. The experimental results on two consecutive 1000-frame sequences illustrate that the proposed framework is very effective and can obtain a very high accuracy that is much better than existing methods.
Xinguo Yu, Changsheng Xu, Qi Tian 0002, Hon Wai Leong
ICME1
2003 Real-time goal-mouth detection in MPEG soccer video
abstract
We report our work in real-time detection of goal-mouth appearances in MPEG soccer video. Processing on sub-optimal quality images after MPEG-decoding, the system constrains the Hough Transform-based line-mark detection to only the dominant green regions typically seen in soccer video. The vertical goal-posts and horizontal goal-bar are then isolated by color-based region (pole)-growing. We demonstrate its application for quick video browsing and virtual content insertion. Extensive test over a large data set of about 15 hours of MPEG-1 soccer video @1.15Mbps, CIF-resolution, shows the robustness of our method.
Kong-Wah Wan, Xin Yan 0001, Xinguo Yu, Changsheng Xu
ACM Multimedia3
2003 Robust goal-mouth detection for virtual content insertion
abstract
In this paper, we describe a working system that detects and segments goal-mouth appearances of soccer video in real-time. Processing on sub-optimal quality images after MPEG-decoding, the system constrains the Hough Transform-based line-mark detection to only the dominant green regions. The vertical goal-posts and horizontal goal-bar are then isolated by color-based region (pole)-growing. We demonstrate its application for quick video browsing and virtual content insertion.
Kong-Wah Wan, Xin Yan 0001, Xinguo Yu, Changsheng Xu
ACM Multimedia3
2003 Trajectory-based ball detection and tracking with applications to semantic analysis of broadcast soccer video
abstract
This paper first presents an improved trajectory-based algorithm for automatically detecting and tracking the ball in broadcast soccer video. Unlike the object-based algorithms, our algorithm does not evaluate whether a sole object is a ball. Instead, it evaluates whether a candidate trajectory, which is generated from the candidate feature image by a candidate verification procedure based on Kalman filter,, which is generated from the candidate feature image by a candidate verification procedure based on Kalman filter, is a ball trajectory. Secondly, a new approach for automatically analyzing broadcast soccer video is proposed, which is based on the ball trajectory. The algorithms in this approach not only improve play-break analysis and high-level semantic event detection, but also detect the basic actions and analyze team ball possession, which may not be analyzed based only on the low-level feature. Moreover, experimental results show that our ball detection and tracking algorithm can achieve above 96% accuracy for the video segments with the soccer field. Compared with the existing methods, a higher accuracy is achieved on goal detection and play-break segmentation. To the best of our knowledge, we present the first solution in detecting the basic actions such as touching and passing, and analyzing the team ball possession in broadcast soccer video.
Xinguo Yu, Changsheng Xu, Hon Wai Leong, Qi Tian 0002, Kong-Wah Wan
ACM Multimedia1