Xiaomeng Wu

dblp:26/2086 · DBLP profile ↗
← Back
53ranked-venue papers
34as first author
13since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 28 first-author · 7 since 2021Artificial intelligence and machine learning · 14 · 9 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2027 Multi-view hypergraph networks with aggregation and estrangement dependencies for multi-source spatiotemporal forecasting
Xiaomeng Wu, Dongmei Fu, Tao Yang 0010, Lizhen Shao
Expert Syst. Appl.1
2026 Graph-enhanced & monotonic embeddings: A novel approach to tabular data representation
Xiaomeng Wu, Dongmei Fu
Neurocomputing1
2025 Interpersonal Communication Skills Training Platform for Teachers: Using Generative AI to Simulate Educational Conflict Resolution
abstract
Traditional teacher interpersonal communication training is often limited to face-to-face formats, offering few opportunities for practice and relying on fixed role-play scenarios. Artificial intelligence (AI) presents a solution to these limitations. In this study, we developed a ChatGPT-4.0-based platform for pre-service teachers to practice and enhance their interpersonal communication skills. A total of 29 pre-service teachers par-ticipated in a one-month training program on the platform, where they could repeatedly simulate real-world communication dilemmas and receive iterative feedback from generative AI. The results demonstrated that the training significantly improved the teachers' interpersonal communication skills and teaching efficacy. This study highlights the potential of AI to empower teacher professional development.
Chengze Zeng, Junjie Shang 0001, Xiaomeng Wu
ICALT5
2025 Semi-Supervised Change Detection With Boundary Refinement Teacher
abstract
High-quality pseudo-labels are critical for guiding learning in semi-supervised change detection (SSCD). Recently, many SSCD methods based on consistency regularization (CR) have achieved advanced performance. These methods typically generate pseudo-labels by setting a fixed threshold. However, this strategy struggles to generate pseudo-labels with high-quality boundary details. To this end, we propose a novel SSCD method boundary refinement teacher (BRT) to enhance the boundary quality of pseudo-labels. A bi-temporal image boundary refinement (BIBR) module is designed to uncover boundary details in unlabeled images at first. BIBR explores the boundary characteristics of pseudo-label by extracting the boundary blocks from its change map and re-delineates the boundary decision points with their magnified views. Then, a stable teacher parameter update (STPU) module is devised to sustain the semi-supervised learning state steady, avoiding frequent updates to the teacher model parameters. These stable updates to teacher model parameters provide continuous and high-quality guidance, preventing pseudo-label fluctuations from disrupting the student model’s acquisition of new knowledge. Extensive experiments are conducted on three commonly used change detection datasets, encompassing buildings and multiple categories, covering SSCD settings, boundary metrics, and detailed ablation studies. Our results show that simply enhancing the boundary quality of the pseudo-labels allows BRT to consistently deliver state-of-the-art (SOTA) performance in SSCD. The code is available at: https://github.com/yoghurts-sy/BRT.
You Su, Yonghong Song, Xiaomeng Wu, Jingqi Chen, Zehan Wen
IEEE Trans. Geosci. Remote. Sens.3
2024 Video Anomaly Detection Via Self-Supervised Learning With Frame Interval and Rotation Prediction
abstract
Video Anomaly Detection (VAD) presents a substantial challenge in the field of computer vision. The emergence of self-supervised learning has played a crucial role in tackling this challenge, with the design of self-supervised pretext tasks proving to be an exceptionally effective method. However, how to imbue the devised pretext task with a more comprehensive understanding of video content is still a problem worth exploring. In this paper, we contribute by introducing two object-level self-supervised tasks tailored for mining temporal and spatial information in videos, subsequently applied to VAD. The self-supervised pretext tasks we have formulated are as follows: (i) Frame Rotation Prediction and (ii) Frame Interval Prediction. These two pretext tasks Focus on handling abnormality of appearance and action respectively. Our approach follows an end-to-end methodology that does not depend on pre-trained models, skeleton data, or optical flow information. In our experimental evaluations, our model has demonstrated superior performance, outperforming established competitors on two publicly available benchmarks. In particular, we achieved a micro-AUROC of 86.9 on the ShanghaiTech dataset.
Ke Jia, Yonghong Song, Xiaomeng Wu, You Su
ICME3
2023 Deep Quantigraphic Image Enhancement via Comparametric Equations
abstract
Most recent methods of deep image enhancement can be generally classified into two types: decompose-and-enhance and illumination estimation-centric. The former is usually less efficient, and the latter is constrained by a strong assumption regarding image reflectance as the desired enhancement result. To alleviate this constraint while retaining high efficiency, we propose a novel trainable module that diversifies the conversion from the low-light image and illumination map to the enhanced image. It formulates image enhancement as a comparametric equation parameterized by a camera response function and an exposure compensation ratio. By incorporating this module in an illumination estimation-centric DNN, our method improves the flexibility of deep image enhancement, limits the computational burden to illumination estimation, and allows for fully unsupervised learning adaptable to the diverse demands of different tasks.
Xiaomeng Wu, Yongqing Sun, Akisato Kimura
ICASSP1
2023 Deep attentive time warping
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
Pattern Recognit.2
2022 Contrast enhancement based on reflectance-oriented probabilistic equalization
Xiaomeng Wu, Yongqing Sun, Akisato Kimura, Kunio Kashino
Signal Process.1
2021 Reflectance-Oriented Probabilistic Equalization for Image Enhancement
abstract
Despite recent advances in image enhancement, it remains difficult for existing approaches to adaptively improve the brightness and contrast for both low-light and normal-light images. To solve this problem, we propose a novel 2D histogram equalization approach. It assumes intensity occurrence and co-occurrence to be dependent on each other and derives the distribution of intensity occurrence (1D histogram) by marginalizing over the distribution of intensity co-occurrence (2D histogram). This scheme improves global contrast more effectively and reduces noise amplification. The 2D histogram is defined by incorporating the local pixel value differences in image reflectance into the density estimation to alleviate the adverse effects of dark lighting conditions. Over 500 images were used for evaluation, demonstrating the superiority of our approach over existing studies. It can sufficiently improve the brightness of low-light images while avoiding over-enhancement in normal-light images.
Xiaomeng Wu, Yongqing Sun, Akisato Kimura, Kunio Kashino
ICASSP1
2021 Attention to Warp: Deep Metric Learning for Multivariate Time Series
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
ICDAR (3)2
2021 Deep Reinforcement Image Matching with Self-Termination
abstract
Deep reinforcement learning-based image matching sequentially searches only the promising regions in the reference image that match the query, leading to a significantly small number of steps compared to traditional methods. Since existing methods do not have any function to judge whether the target region has been successfully identified or not, they continue to search until the preset maximum number of search steps is reached. In this paper, we propose a deep image matching network that can terminate the matching process by itself. Our network is designed to have a halting module that identifies whether the current reference region matches the query based on the image features and the search history. The entire network is effectively trained end-to-end in a framework of deep reinforcement learning that incorporates a new loss function to evaluate the accuracy of the termination decision. Experimental results demonstrate that our method can achieve highly competitive or better matching accuracy with fewer search steps than the existing methods.
Onkar Krishna, Go Irie, Xiaomeng Wu, Akisato Kimura, Kunio Kashino
ICIP3
2021 Contrast enhancement based on discriminative co-occurrence statistics
Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
Multim. Tools Appl.1
2021 Reflectance-Guided Histogram Equalization and Comparametric Approximation
abstract
Existing image enhancement methods fall short of expectations because with them it is difficult to improve global and local image contrast simultaneously. To address this issue, we propose a histogram equalization-based method called RG-CACHE. It adapts to the data-dependent requirements of brightness enhancement and improves the visibility of details without losing the global contrast. RG-CACHE incorporates the spatial information provided by image context into density estimation for discriminative histogram equalization. To minimize the adverse effect of nonuniform illumination, we propose defining spatial information on the basis of image reflectance estimated with edge-preserving smoothing. RG-CACHE works particularly well for determining how the background brightness should be adaptively adjusted and for revealing useful image details hidden in the dark. To handle the loss of details due to the monotonicity of the intensity mapping function, we further propose a post-processing method to approximate RG-CACHE with a brightness transformation function corresponding to a parameterized camera response function. This method is called comparametric approximation. It takes into account a regression problem, in which the parameters of the camera response function are chosen so that the converted intensities are optimally matched to the image enhanced by RG-CACHE. Comparametric approximation is especially suitable for recovering useful image details that tend to be suppressed due to insufficient reflectance contrast.
Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
IEEE Trans. Circuits Syst. Video Technol.1
2020 Adaptive Spotting: Deep Reinforcement Object Search in 3D Point Clouds
Onkar Krishna, Go Irie, Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
ACCV (3)3
2020 Reflectance-Guided, Contrast-Accumulated Histogram Equalization
abstract
Existing image enhancement methods fall short of expectations because with them it is difficult to improve global and local image contrast simultaneously. To address this problem, we propose a histogram equalization-based method that adapts to the data-dependent requirements of brightness enhancement and improves the visibility of details without losing the global contrast. This method incorporates the spatial information provided by image context in density estimation for discriminative histogram equalization. To minimize the adverse effect of non-uniform illumination, we propose defining spatial information on the basis of image reflectance estimated with edge preserving smoothing. Our method works particularly well for determining how the background brightness should be adaptively adjusted and for revealing useful image details hidden in the dark.
Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
ICASSP1
2020 Total Whitening for Online Signature Verification Based on Deep Representation
abstract
In deep metric learning targeted at time series, the correlation between feature activations may be easily enlarged through highly nonlinear neural networks, leading to suboptimal embedding effectiveness. An effective solution to this problem is whitening. For example, in online signature verification, whitening can be derived for three individual Gaussian distributions, namely the distributions of local features at all temporal positions 1) for all signatures of all subjects, 2) for all signatures of each particular subject, and 3) for each particular signature of each particular subject. This study proposes a unified method called total whitening that integrates these individual Gaussians. Total whitening rectifies the layout of multiple individual Gaussians to resemble a standard normal distribution, improving the balance between intraclass invariance and interclass discriminative power. Experimental results demonstrate that total whitening achieves state-of-the-art accuracy when tested on online signature verification benchmarks.
Xiaomeng Wu, Akisato Kimura, Kunio Kashino, Seiichi Uchida
ICPR1
2019 Learning Search Path for Region-level Image Matching
abstract
Finding a region of an image which matches to a query from a large number of candidates is a fundamental problem in image processing. The exhaustive nature of the sliding window approach has encouraged works that can reduce the run time by skipping unnecessary windows or pixels that do not play a substantial role in search results. However, such a pruning-based approach still needs to evaluate the non-ignorable number of candidates, which leads to a limited efficiency improvement. We propose an approach to learn efficient search paths from data. Our model is based on a CNN-LSTM architecture which is designed to sequentially determine a prospective location to be searched next based on the history of the locations attended. We propose a reinforcement learning algorithm to train the model in an end-to-end manner, which allows to jointly learn the search paths and deep image features for matching. These properties together significantly reduce the number of windows to be evaluated and makes it robust to background clutters. Our model gives remarkable matching accuracy with the reduced number of windows and run time on MNIST and FlickrLogos-32 datasets.
Onkar Krishna, Go Irie, Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
ICASSP3
2019 Prewarping Siamese Network: Learning Local Representations for Online Signature Verification
abstract
We propose a neural network-based framework for learning local representations of multivariate time series, and demonstrate its effectiveness for online signature verification. In contrast to related works that optimize a global distance objective, we incorporate a Siamese network into dynamic time warping (DTW), leading to a novel prewarping Siamese network (PSN) optimized with a local embedding loss. PSN learns a feature space that preserves the temporal location-wise distances of local structures. Local embedding, along with the alignment conditions of DTW, imposes a temporal consistency constraint on the sequence-level distance measure while achieving invariance as regards non-linear distortions. Validation on online signature verification datasets demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Seiichi Uchida, Kunio Kashino
ICASSP1
2019 Deep Dynamic Time Warping: End-to-End Local Representation Learning for Online Signature Verification
abstract
Siamese networks have been shown to be successful in learning deep representations for multivariate time series verification. However, most related studies optimize a global distance objective and suffer from a low discriminative power due to the loss of temporal information. To address this issue, we propose an end-to-end, neural network-based framework for learning local representations of time series, and demonstrate its effectiveness for online signature verification. This framework optimizes a Siamese network with a local embedding loss, and learns a feature space that preserves the temporal location-wise distances between time series. To achieve invariance to non-linear temporal distortion, we propose building a dynamic time warping block on top of the Siamese network, which will greatly improve the accuracy for local correspondences across intra-personal variability. Validation with respect to online signature verification demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Brian Kenji Iwana, Seiichi Uchida, Kunio Kashino
ICDAR1
2018 Query Expansion with Diffusion On Mutual Rank Graphs
abstract
In query expansion for object retrieval, there is substantial danger of query drift, where irrelevant information is inferred from pseudo-relevant images to enrich the query. To address this issue, we propose a query expansion method from the viewpoint of diffusion. It explores the structure of highly ranked images in a topological space, assuming that false positives reside on different manifolds from the query. For this purpose, a mutual rank graph is defined on pseudo-relevant images, and their distribution is learned by diffusing their query similarities through the graph. The relevance of a database image can thus be obtained by marginalizing over the learned distribution. The mutual rank graph accounts for varying local density in the image space, leading to great robustness as regards query drift and high generalization ability. The proposed method experimentally shows a consistent boost in the performance of object retrieval with handcrafted features on standard benchmarks.
Xiaomeng Wu, Go Irie, Kaoru Hiramatsu, Kunio Kashino
ICASSP1
2018 Weighted Generalized Mean Pooling for Deep Image Retrieval
abstract
Spatial pooling over convolutional activations (e.g., max pooling or sum pooling) has been shown to be successful in learning deep representations for image retrieval. However, most pooling techniques assume that every activation is equally important, and as a result they suffer from the presence of uninformative image regions that play a negative role as regards matching or lead to the confusion of particular visual instances. To address this issue, we propose a trainable building block that steers pooling to local information important to the task at hand. The method formulates pooling as a weighted generalized mean (wGeM), in which weights are learned on activations, reflecting the discriminative power of each activation in image matching. Embedding wGeM in a deep network improves image representation and boosts retrieval performance on standard benchmarks. wGeM does not require any bounding box annotations, but instead learns the latent probabilities of activations from scratch. It even goes beyond objectness, and learns to look at important visual details rather than the whole region of the object of interest.
Xiaomeng Wu, Go Irie, Kaoru Hiramatsu, Kunio Kashino
ICIP1
2018 Label Propagation with Ensemble of Pairwise Geometric Relations: Towards Robust Large-Scale Retrieval of Object Instances
Xiaomeng Wu, Kaoru Hiramatsu, Kunio Kashino
Int. J. Comput. Vis.1
2017 Deep salience map guided arbitrary direction scene text recognition
abstract
Irregular scene text such as curved, rotated or perspective texts commonly appear in natural scene images due to different camera view points, special design purposes etc. In this work, we propose a text salience map guided model to recognize these arbitrary direction scene texts. We train a deep Fully Convolutional Network (FCN) to calculate the precise salience map for texts. Then we estimate the positions and rotations of the text and utilize this information to guide the generation of CNN sequence features. Finally the sequence is recognized with a Recurrent Neural Network (RNN) model. Experiments on various public datasets show that the proposed approach is robust to different distortions and performs superior or comparable to the state-of-the-art techniques.
Xinhao Liu 0001, Takahito Kawanishi, Xiaomeng Wu, Kaoru Hiramatsu, Kunio Kashino
ICASSP3
2017 Edited film alignment via selective Hough transform and accurate template matching
abstract
Edited film alignment is the post-production process of finding small parts of unedited footage that temporally and spatially match an edited film. The huge amount of data to be processed makes significant downsampling of the videos essential in real-life applications. Simultaneously, professional users demand that the task be achieved with frame and pixel-level accuracy. We propose a novel selective Hough transform (SHT) and an accurate template matching method to address the difficult trade-off between accuracy and scalability. For robust temporal alignment, SHT investigates the selectivity of frame-level similarities and advantageously reduces the weights of mismatches. The template matching method encompasses spatial Hough transform and sum of squared differences (SSD) minimization. SSD is efficiently approximated by exploiting the second-order derivative of image intensity. Experiments conducted on real-world data show the superiority of our methods.
Xiaomeng Wu, Takahito Kawanishi, Minoru Mori, Kaoru Hiramatsu, Kunio Kashino
ICASSP1
2017 Contrast-accumulated histogram equalization for image enhancement
abstract
Among image enhancement methods, histogram equalization (HE) has received the most attention because of its intuitive implementation quality, high efficiency, and the monotonicity of its intensity mapping function. However, HE is indiscriminate and overemphasizes the contrast around intensities with large pixel populations but little visual importance. To address this issue, we propose an HE-based method that adaptively controls the contrast gain according to the potential visual importance of intensities and pixels. Observing that in natural scenes image details are usually hidden in darker regions that have noticeable local differences, we formulate the potential visual importance on the basis of the multi-resolution, dark-pass filtered gradients in the image. Experiments show that our method is highly discriminating in terms of noises and trivial image gradients, and it guarantees great global contrast preservation.
Xiaomeng Wu, Xinhao Liu 0001, Kaoru Hiramatsu, Kunio Kashino
ICIP1
2016 Scene text recognition with high performance CNN classifier and efficient word inference
abstract
The recognition of text in natural scene images is a practical yet challenging task due to the large variations in backgrounds, textures, fonts, and illumination conditions. In this paper, we propose a highly accurate character recognition model by utilizing the representational power of a specially designed Convolutional Neural Network (CNN). Based on the recognition model, we also develop an efficient post processing approach for error correction and hypothesis re-verification. Character and word image recognition experiments on two public datasets, namely the ICDAR 2003 Robust Reading dataset and the Street View Text (SVT) dataset both show that the proposed approach provides superior or comparable results to the state-of-the-art techniques.
Xinhao Liu 0001, Takahito Kawanishi, Xiaomeng Wu, Kunio Kashino
ICASSP3
2016 Scene text recognition with CNN classifier and WFST-based word labeling
abstract
Natural scene text recognition has proved to be challenging due to the unconstrained wild conditions. In this paper, to solve this problem we propose a method which first detects and recognizes characters by utilizing the high performance Convolutional Neural Network (CNN). Then for post-processing, inspired by its success in speech recognition, we employ the efficient and flexible Weight Finite State Transducer (WFST) based word labeling model for incorporation with a lexicon or high order language model. In the experiments we show that the proposed approach can correctly and robustly recognize the text in the scene images and the results for serveral public datasets (ICDAR 2003, SVT and IIIT5K) show comparable or superior performance to the state-of-the-art algorithms.
Xinhao Liu 0001, Takahito Kawanishi, Xiaomeng Wu, Kunio Kashino
ICPR3
2015 Robust Spatial Matching as Ensemble of Weak Geometric Relations
Xiaomeng Wu, Kunio Kashino
BMVC1
2015 Trademark Image Retrieval Using Inverse Total Feature Frequency and Multiple Detectors
Minoru Mori, Xiaomeng Wu, Kunio Kashino
CAIP (1)2
2015 Assessing Students' Learning Experience and Achievements in a Medium-Sized Massively Open Online Course
abstract
A medium-sized Massively Open Online Course (MOOC) was hosted by Peking University in 2013. Altogether 192 Chinese students came to Peking University and learned face-to-face, another 311 Chinese students participated online. This study targeted one research question: Were there any significant differences between the learning experiences and outcomes of onsite and online students? Although onsite students had lower attrition and higher completion rates than their online peers, no significant difference was detected between the average assignment scores of the onsite and online participants who had completed all the assignments. Learners also responded to a survey asking for their learning experiences. There were no significant differences between the online and onsite students' ratings of technology quality and usability of course management system, instructional content, and the design of learning assessment. Findings from this first empirical study on a Chinese MOOC will inform researchers and practitioners interested in introducing MOOCs to Chinese students.
Jiyou Jia, Jingmin Miao, Xiaomeng Wu, Aihua Wang, Baijie Yang
ICALT4
2015 Adaptive Dither Voting for Robust Spatial Verification
abstract
Hough voting in a geometric transformation space allows us to realize spatial verification, but remains sensitive to feature detection errors because of the inflexible quantization of single feature correspondences. To handle this problem, we propose a new method, called adaptive dither voting, for robust spatial verification. For each correspondence, instead of hard-mapping it to a single transformation, the method augments its description by using multiple dithered transformations that are deterministically generated by the other correspondences. The method reduces the probability of losing correspondences during transformation quantization, and provides high robustness as regards mismatches by imposing three geometric constraints on the dithering process. We also propose exploiting the non-uniformity of a Hough histogram as the spatial similarity to handle multiple matching surfaces. Extensive experiments conducted on four datasets show the superiority of our method. The method outperforms its state-of-the-art counterparts in both accuracy and scalability, especially when it comes to the retrieval of small, rotated objects.
Xiaomeng Wu, Kunio Kashino
ICCV1
2015 Data-driven taxonomy forest for fine-grained image categorization
abstract
Fine-grained image categorization must handle huge cross-class ambiguities and a large number of classes. Inspired by the success of rigid hierarchical classification, we propose a new flexible hierarchical classification method, called a data-driven taxonomy forest. It constructs a multitude of taxonomies, each of which converts a complex multi-class problem to a more easily tractable path-finding problem. We demonstrate how a stochastic representation of local classification hypotheses incorporated in multiple taxonomies deals skillfully with error propagation and over-fitting. Various strategies for instance space decomposition are investigated from the viewpoint of taxonomy complexity. We comprehensively evaluate our data-driven taxonomy forest using Oxford Flower 102 and Oxford Pet benchmarks and show its superiority in effectiveness and generality to rigid hierarchical classification in fine-grained image categorization tasks.
Xiaomeng Wu, Minoru Mori, Kunio Kashino
ICME1
2015 Interest point selection by topology coherence for multi-query image retrieval
Xiaomeng Wu, Kunio Kashino
Multim. Tools Appl.1
2015 Second-Order Configuration of Local Features for Geometrically Stable Image Matching and Retrieval
abstract
Local features offer high repeatability, which supports efficient matching between images, but they do not provide sufficient discriminative power. Imposing a geometric coherence constraint on local features improves the discriminative power but makes the matching sensitive to anisotropic transformations. We propose a novel feature representation approach to solve the latter problem. Each image is abstracted by a set of tuples of local features. We revisit affine shape adaptation and extend its conclusion to characterize the geometrically stable feature of each tuple. The representation thus provides higher repeatability with anisotropic scaling and shearing than found in previous research. We develop a simple matching model by voting in the geometrically stable feature space, where votes arise from tuple correspondences. To make the required index space linear as regards the number of features, we propose a second approach called a centrality-sensitive pyramid to select potentially meaningful tuples of local features on the basis of their spatial neighborhood information. It achieves faster neighborhood association and has a greater robustness to errors in interest point detection and description. We comprehensively evaluated our approach using Flickr Logos 32, Holiday, Oxford Buildings, and Flickr 100 K benchmarks. Extensive experiments and comparisons with advanced approaches demonstrate the superiority of our approach in image retrieval tasks.
Xiaomeng Wu, Kunio Kashino
IEEE Trans. Circuits Syst. Video Technol.1
2014 Tri-Map Self-Validation Based on Least Gibbs Energy for Foreground Segmentation
Xiaomeng Wu, Kunio Kashino
BMVC1
2014 Image retrieval based on spatial context with Relaxed Gabriel Graph pyramid
abstract
Imposing the coherence of the spatial context on local features is becoming a necessity for object retrieval and recognition. Motivated by the success of proximity graphs in topological decomposition, clustering, and gradient estimation, we introduce a variation on and a generalization of Delaunay Triangulation, called a Relaxed Gabriel Graph (RGG), as the apex of spatial neighborhood association and design a Centrality-Sensitive Pyramid (CSP) model for hierarchical spatial context modeling. RGG is parameterized, and so allows the tuning of various applications and datasets. CSP achieves better neighborhood association and is more robust as regards feature description error than other related work. Our method is evaluated on Flickr Logos 32, Holiday, and Oxford Buildings benchmarks. Experimental results and comparisons demonstrate the superiority of our method in an image retrieval scenario.
Xiaomeng Wu, Kunio Kashino
ICASSP1
2014 Image Retrieval Based on Anisotropic Scaling and Shearing Invariant Geometric Coherence
abstract
Imposing a spatial coherence constraint on image matching is becoming a necessity for local feature based object retrieval. We tackle the affine invariance problem of the prior spatial coherence model and propose a novel approach for geometrically stable image retrieval. Compared with related studies focusing simply on translation, rotation, and isotropic scaling, our approach can deal with more significant transformations including anisotropic scaling and shearing. Our contribution consists of revisiting the first-order affine adaptation approach and extending its application to represent the geometric coherence of a second-order local feature structure. We comprehensively evaluated our approach using Flickr Logos 32, Holiday, and Oxford Buildings benchmarks. Extensive experimentation and comparisons with state-of-the-art spatial coherence models demonstrate the superiority of our approach in image retrieval tasks.
Xiaomeng Wu, Kunio Kashino
ICPR1
2014 Tell Me about TV Commercials of This Product
Cai-Zhi Zhu, Siriwat Kasamwattanarote, Xiaomeng Wu, Shin'ichi Satoh 0001
MMM (1)3
2013 Connect commercial films with realities
abstract
Broadcast TV program is a quite informative media resource which records our daily life over the time. While for emphasizing real-time reporting, those out-of-date video archives once were elaborately created with high quality are always left without being fully used. In this paper, many known state-of-the-art retrieval technologies are integrated into a commercial film retrieval system, which manages to index a huge commercial dataset archived from five TV channels within recent three years. The final purpose is to connect images queried by users with our archived broadcast video dataset via searching relevant commercials and accessing their broadcast information, such as air time and replay frequency. This system also serves as one part of our ongoing broadcast TV program reusing project.
Cai-Zhi Zhu, Siriwat Kasamwattanarote, Xiaomeng Wu, Shin'ichi Satoh 0001
ICMR3
2013 Ultrahigh-Speed TV Commercial Detection, Extraction, and Matching
abstract
We describe a system based on exact-duplicate matching for detecting and localizing TV commercials in a video stream, clustering the exact duplicates, and detecting duplicate exact-duplicate clusters across video streams. A two-stage temporal recurrence hashing algorithm is used for the detection, localization, and clustering. The algorithm is fully unsupervised, generic, and ultrahigh speed. Another algorithm is used to integrate the video and audio streams to achieve higher performance extraction. Its sequence- and frame-level accuracies in testing were respectively 98.1% and 97.4%. A third algorithm uses a new bag-of-fingerprints model to detect duplicate exact-duplicate clusters across multiple streams. It is robust against decoding errors. Its contributions include: 1) fully unsupervised detection, extraction, and matching of exact duplicates; 2) more generic commercial detection than with the knowledge-based techniques; 3) ultrahigh-speed processing, which detected the TV commercials from a one-month video stream in less than 42 minutes, which is more than ten times faster than with state-of-the-art algorithms; and 4) more generic operation in terms of signal input, the performance of which is consistent between video and audio streams. Testing using a video database containing a ten-hour, a one-month, and a five-year video stream comprehensively demonstrates the effectiveness and efficiency of this system.
Xiaomeng Wu, Shin'ichi Satoh 0001
IEEE Trans. Circuits Syst. Video Technol.1
2011 Temporal recurrence hashing algorithm for mining commercials from multimedia streams
abstract
We propose a dual-stage algorithm for fully-unsupervised and super fast TV commercial mining in this paper. The two stages involved in process include: 1) searching for recurring short segments, and 2) assembling these short segments into sets of long and complete commercial sequences. The first stage is achieved by frame hashing. Different from the related studies that depend on brute-force pairwise matching, we propose applying a second-stage hashing algorithm for the recurring segment assemblage, which is the key idea in this pa per. A large-scale archive containing a 10-hour and a 1-month stream was used for the experimentation. The algorithm mined commercials from the 1-month stream in less than 50 minutes, which was ten times faster than that of related studies, with a 98.05% sequence level and 97.39% frame-level accuracy. We demonstrate the performance consistency of the algorithm on both audio and video streams, and investigate the computational cost from both the theoretical and experimental viewpoints.
Xiaomeng Wu, Shin'ichi Satoh 0001
ICASSP1
2011 Commercial mining basedon temporal recurrence hashing algorithm and bag-of-fingerprints model
abstract
We propose two novel algorithms for fully-unsupervised, super-fast, and cross-channel TV commercial mining in this paper. The tasks involved in the process include: 1) mining commercial clusters from streams of individual channels, and 2) grouping identical commercial clusters across multiple channels. The first process is achieved with a dual-stage hashing algorithm, which searches for recurring short segments by hashing frames, and it assembles these short segments into sets of commercial sequences by hashing temporal recurrences. The algorithm mined commercials from a one-month stream in less than 42 minutes, which was ten times faster than that in related studies. A new bag-of-fingerprints model is proposed for the second process to encode the temporal clues of local fingerprints. The model is abundantly robust against framing and fingerprinting errors in recurring sequences, and discovers false matches of local fingerprints. A five-month database was used for comprehensively demonstrating the effectiveness and efficiency of the model.
Xiaomeng Wu, Shin'ichi Satoh 0001
ICIP1
2010 PageRank with Text Similarity and Video Near-Duplicate Constraints for News Story Re-ranking
Xiaomeng Wu, Ichiro Ide, Shin'ichi Satoh 0001
MMM1
2008 Scene duplicate detection based on the pattern of discontinuities in feature point trajectories
abstract
The paper is aiming to detect and retrieve videos of the same scene (scene duplicates) from broadcast video archives. Scene duplicate is composed of different pieces of footage of the same scene, the same event, at the same time, but from the different viewpoints. Scene duplicate detection would be particularly useful to identify the same event reported in different programs from different broadcast stations. The approach should be invariant to viewpoint changes. We focused on object motion in videos and devised a video matching approach based on the temporal pattern of discontinuities obtained from feature point trajectories. We developed an acceleration method based on the discontinuity pattern, which is more robust to variations in camerawork and editing than conventional features, to dramatically reduce the computation burden. We compared our approach with an existing video matching method based on the local feature of keyframe. The spatial registration strategy of this method was also used with the proposed approach to cope with visually different unrelated video pairs. The performance and effectiveness of our approach was demonstrated on actual broadcasted videos.
Xiaomeng Wu, Masao Takimoto, Shin'ichi Satoh 0001, Jun Adachi
ACM Multimedia1
2007 Nucleotide composition string selection in HIV-1 subtyping using whole genomes
abstract
MOTIVATION: The availability of the whole genomic sequences of HIV-1 viruses provides an excellent resource for studying the HIV-1 phylogenies using all the genetic materials. However, such huge volumes of data create computational challenges in both memory consumption and CPU usage. RESULTS: We propose the complete composition vector representation for an HIV-1 strain, and a string scoring method to extract the nucleotide composition strings that contain the richest evolutionary information for phylogenetic analysis. In this way, a large-scale whole genome phylogenetic analysis for thousands of strains can be done both efficiently and effectively. By using 42 carefully curated strains as references, we apply our method to subtype 1156 HIV-1 strains (10.5 million nucleotides in total), which include 825 pure subtype strains and 331 recombinants. Our results show that our nucleotide composition string selection scheme is computationally efficient, and is able to define both pure subtypes and recombinant forms for HIV-1 strains using the 5000 top ranked nucleotide strings. AVAILABILITY: The Java executable and the HIV-1 datasets are accessible through 'http://www.cs.ualberta.ca/~ghlin/src/WebTools/hiv.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaomeng Wu, Zhipeng Cai 0001, Xiu-Feng Wan, Tin Hoang, Randy Goebel, Guohui Lin
Bioinform.1
2006 Interactive Object Annotation for Construction of Video Information System
abstract
In this paper, an interactive object annotation approach for video database will be proposed. The same semantic objects such as characters, backgrounds, and the main subjects in keyframes of each video shots can be queried and annotated based on the similarity of low-level features such as the color, area, and position of each region. An adaptive image enhancement algorithm is used for handling the lighting changes, and a region-based fuzzy feature matching approach is used for addressing the typical feature representation impreciseness. The content provider can then select relevant keyframes interactively from the results to annotate matched objects in them according to the descriptions that are added into the model. Based on this approach, a video information system is proposed for supporting video content generation. Furthermore, a novel practical application is constructed by using this support system and implemented to show the practicability of it
Xiaomeng Wu, Shunsuke Kamijo, Masao Sakauchi
ISM1
2006 Selection Measure of Illumination Instability for Multimedia Data Indexing
abstract
In this paper, a novel approach, which automatically measures the illumination instability of the video, is proposed to provide selection measure of target video for color-based multimedia system. An information-theoretic measure is proposed as a quantitative measure of the information distribution within an image. This measure is further extended to the video case and used to quantitatively represent the lighting condition of each scene. The illumination instability of the video is thus measured by calculating the instability of the features extracted from the extended measure. Experiments are generated to demonstrate how the proposed approach using the information-theoretic measure to take information distribution within an image into account can reflect the instability more effectively than other simple and straight-forward measures, and how to use it to provide selection measure of target video for color-based multimedia data indexing system
Xiaomeng Wu, Shunsuke Kamijo, Masao Sakauchi
ISM1
2006 Semantic video database system with semi-automatic secondary-content generation capability
Xiaomeng Wu, Shunsuke Kamijo, Masao Sakauchi
Multim. Tools Appl.2
2005 Faster solution to the maximum quartet consistency problem with constraint programming
Gang Wu 0020, Guohui Lin, Jia-Huai You, Xiaomeng Wu
APBC4
2005 Selected String Representation for Whole Genomes
Xiaomeng Wu, Guohui Lin
CIBCB1
2003 A proposal for a video content generation support system and its application
abstract
A video content generation support system, based on an interactive approach that maps low-level features to high-level concepts, is proposed. By consulting an ontological semantic object model database, the same semantic objects such as characters, backgrounds, and the main subjects in key frames of each video shots can be queried and automatically annotated based on the similarity of low-level features such as the color, area, and position of each region. Since image recognition techniques are limited in their ability to fully identify and compare images, an additional function is proposed, which uses a coarse model to recover a higher number of similar key frames to provide more relevant results. The content provider can then select relevant key frames interactively from the results to annotate matched objects in them according to the descriptions that are added into the model. Therefore, more complex content can be generated with a higher accuracy by using a combination of the application-oriented operations. The system has high potential for use in object-based interactive multimedia applications. One prototype application is also presented.
Xiaomeng Wu, Shunsuke Kamijo, Masao Sakauchi
ICME2
2003 Multiple Agents Moving Target Search
Meir Goldenberg, Alexander Kovarsky, Xiaomeng Wu, Jonathan Schaeffer 0001
IJCAI3
2003 Construction of interactive video information system by applying results of object recognition
abstract
Although numerous attempts have been made to determine algorithms and approaches for building up a video information system, not many practical applications have been proposed. In this paper, a novel interactive video information system called the Drama Characters' Popularity Voting System (DCPVS) is constructed by applying the results of off-line object recognition. The system's purpose is to provide description annotation, retrieval, and statistics in the video associated with an object, such as a character as a basic unit, over the Internet. By using the proposed system, multiple users in a network can enjoy the same video and can vote for the characters they like in it. The voting information is collected and stored in the server, which then provides the statistics regarding the popularity of different characters or the voting rates within different periods of the video.
Xiaomeng Wu, Shunsuke Kamijo, Masao Sakauchi
ACM Multimedia1