Baoxin Li

dblp:36/395 · DBLP profile ↗
← Back
145ranked-venue papers
11as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 115 · 10 first-author · 18 since 2021Artificial intelligence and machine learning · 51 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 1 since 2021Human-computer interaction and ubiquitous computing · 9Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Structuring the Unstructured: A Zero-Shot Approach to Video Chaptering and Title Generation
Nupur Thakur, Riti Paul, Baoxin Li
ICPR (15)3
2026 AUCp: Pseudo-AUC for Inference Model Selection With Unlabeled Validation Data in Abnormality Detection
Md Mahfuzur Rahman Siddiquee, Fazle Rafsani, Jay Shah, Teresa Wu, Catherine D. Chong, Todd J. Schwedt, Baoxin Li
IEEE Trans. Medical Imaging7
2025 TFM2: Training-Free Mask Matching for Open-Vocabulary Semantic Segmentation
abstract
The potential of Open-Vocabulary Semantic Segmentation (OVSS) in few-shot scenarios is not fully explored due to the complexity of extending few-shot concepts to seman-tic segmentation tasks. To address this challenge, we propose Training-Free Mask Matching (TFM2), an efficient, mask-based adapter method that enhances OVSS models for the few-shot open vocabulary semantic segmentation task. TFM2 is a key-value cache that explicitly designed for image masks. We introduce three modules to construct and refine the mask cache, subsequently enhancing the OVSS mask classification performance. Comprehensive experiments demonstrate that TFM2 improves the performance of state-of-the-art OVSS methods by a margin of 1% to 5% across different settings. Moreover, TFM2 is not limited to any specific methods or backbones. This work underscores the importance and potential of few-shot data in OVSS and presents a significant step toward leveraging this potential
Yaoxin Zhuo, Zachary Bessinger, Lichen Wang, Naji Khosravan, Baoxin Li, Sing Bing Kang
WACV5
2024 Transformer-Based Selective Super-resolution for Efficient Image Refinement
abstract
Conventional super-resolution methods suffer from two drawbacks: substantial computational cost in upscaling an entire large image, and the introduction of extraneous or potentially detrimental information for downstream computer vision tasks during the refinement of the background. To solve these issues, we propose a novel transformer-based algorithm, Selective Super-Resolution (SSR), which partitions images into non-overlapping tiles, selects tiles of interest at various scales with a pyramid architecture, and exclusively reconstructs these selected tiles with deep features. Experimental results on three datasets demonstrate the efficiency and robust performance of our approach for super-resolution. Compared to the state-of-the-art methods, the FID score is reduced from 26.78 to 10.41 with 40% reduction in computation cost for the BDD100K dataset.
Kishore Kasichainula, Yaoxin Zhuo, Baoxin Li, Jae-sun Seo, Yu Cao 0001
AAAI4
2024 PatchRot: Self-Supervised Training of Vision Transformers by Rotation Prediction
Sachin Chhabra, Hemanth Venkateswara, Baoxin Li
BMVC3
2024 Label Smoothing++: Enhanced Label Regularization for Training Neural Networks
Sachin Chhabra, Hemanth Venkateswara, Baoxin Li
BMVC3
2024 A-FSL: Adaptive Few-Shot Learning via Task-Driven Context Aggregation and Attentive Feature Refinement
Riti Paul, Sahil Vora, Nupur Thakur, Baoxin Li
ICPR (26)4
2024 Ordinal Classification with Distance Regularization for Robust Brain Age Prediction
abstract
Age is one of the major known risk factors for Alzheimer's Disease (AD). Detecting AD early is crucial for effective treatment and preventing irreversible brain damage. Brain age, a measure derived from brain imaging reflecting structural changes due to aging, may have the potential to identify AD onset, assess disease risk, and plan targeted interventions. Deep learning-based regression techniques to predict brain age from magnetic resonance imaging (MRI) scans have shown great accuracy recently. However, these methods are subject to an inherent regression to the mean effect, which causes a systematic bias resulting in an overestimation of brain age in young subjects and underestimation in old subjects. This weakens the reliability of predicted brain age as a valid biomarker for downstream clinical applications. Here, we reformulate the brain age prediction task from regression to classification to address the issue of systematic bias. Recognizing the importance of preserving ordinal information from ages to understand aging trajectory and monitor aging longitudinally, we propose a novel ORdinal Distance Encoded Regularization (ORDER) loss that incorporates the order of age labels, enhancing the model's ability to capture age-related patterns. Extensive experiments and ablation studies demonstrate that this framework reduces systematic bias, outperforms state-of-art methods by statistically significant margins, and can better capture subtle differences between clinical groups in an independent AD dataset. Our implementation is publicly available at https://github.com/jaygshah/Robust-Brain-Age-Prediction.
Jay Shah, Md Mahfuzur Rahman Siddiquee, Teresa Wu, Baoxin Li
WACV5
2024 Brainomaly: Unsupervised Neurologic Disease Detection Utilizing Unannotated T1-weighted Brain MR Images
abstract
Harnessing the power of deep neural networks in the medical imaging domain is challenging due to the difficulties in acquiring large annotated datasets, especially for rare diseases, which involve high costs, time, and effort for annotation. Unsupervised disease detection methods, such as anomaly detection, can significantly reduce human effort in these scenarios. While anomaly detection typically focuses on learning from images of healthy subjects only, real-world situations often present unannotated datasets with a mixture of healthy and diseased subjects. Recent studies have demonstrated that utilizing such unannotated images can improve unsupervised disease and anomaly detection. However, these methods do not utilize knowledge specific to registered neuroimages, resulting in a subpar performance in neurologic disease detection. To address this limitation, we propose Brainomaly, a GAN-based image-to-image translation method specifically designed for neurologic disease detection. Brainomaly not only offers tailored image-to-image translation suitable for neuroimages but also leverages unannotated mixed images to achieve superior neurologic disease detection. Additionally, we address the issue of model selection for inference without annotated samples by proposing a pseudo-AUC metric, further enhancing Brainomaly's detection performance. Extensive experiments and ablation studies demonstrate that Brainomaly outperforms existing state-of-the-art unsupervised disease and anomaly detection methods by significant margins in Alzheimer's disease detection using a publicly available dataset and headache detection using an institutional dataset. The code is available from https://github.com/mahfuzmohammad/Brainomaly.
Md Mahfuzur Rahman Siddiquee, Jay Shah, Teresa Wu, Catherine D. Chong, Todd J. Schwedt, Gina Dumkrieger, Simona Nikolova, Baoxin Li
WACV8
2024 Graph(Graph): A Nested Graph-Based Framework for Early Accident Anticipation
abstract
Anticipating traffic accidents early using dashcam videos is an important task for ensuring road safety and building reliable intelligent autonomous vehicles. However, factors like high traffic on the roads, different types of accidents, limited angles of vision, etc. make this task very challenging. Using the early frames, a lot of existing methods predict a large number of false positives which poses a huge risk for all vehicles on the road. In this paper, we propose a novel end-to-end learning, nested graph-based framework named Graph(Graph) for early accident anticipation. It uses interactions between the objects in the same as well as the neighboring frames along with the global features to make precise predictions as early as possible. This way it is able to embed the local as well as global temporal information into the extracted features. Graph(Graph) outperforms state-of-the-art methods on different datasets by a large margin demonstrating its effectiveness. With empirical evidence, we highlight the importance of each component in Graph(Graph) and show their effect on the final performance. Our code is available at https://github.com/thakurnupur/Graph-Graph.
Nupur Thakur, PrasanthSai Gouripeddi, Baoxin Li
WACV3
2024 Patch-based Selection and Refinement for Early Object Detection
abstract
Early object detection (OD) is a crucial task for the safety of many dynamic systems. Current OD algorithms have limited success for small objects at a long distance. To improve the accuracy and efficiency of such a task, we propose a novel set of algorithms that divide the image into patches, select patches with objects at various scales, elaborate the details of a small object, and detect it as early as possible. Our approach is built upon a transformer-based network and integrates the diffusion model to improve the detection accuracy. As demonstrated on BDD100K, our algorithms enhance the mAP for small objects from 1.03 to 8.93, and reduce the data volume in computation by more than 77%.
Kishore Kasichainula, Yaoxin Zhuo, Baoxin Li, Jae-sun Seo, Yu Cao 0001
WACV4
2024 FELGA: Unsupervised Fragment Embedding for Fine-Grained Cross-Modal Association
abstract
Vision-and-Language Pre-trained (VLP) models have demonstrated their powerful zero-shot ability in multiple downstream tasks. Most of these models are designed to learn joint embeddings of images and their paired sentences, with both modalities considered globally. This does not lead to optimal solutions for applications where what matters more is the local-level cross-modal association, such as the situation where a user may want to retrieve images with query words that link to only small parts of the images. While a VLP model could in principle be retrained to learn a new embedding capturing such fine-grained association, expensive annotation would be needed, making it impractical for big data applications. This paper proposes a novel method named Fragment Embedding by Local and Global Alignment (FELGA), which learns fragment-level embeddings that capture fine-grained cross-modal association through utilizing visual entity proposals and semantic concept proposals in an unsupervised manner. Comprehensive experiments conducted on three VLP models and two datasets demonstrate that FELGA is not limited to specific VLP models and outperforms the original VLP features. In particular, the learned embeddings support cross-modal fragment association tasks including query-driven object discovery and description assignment.
Yaoxin Zhuo, Baoxin Li
WACV2
2023 Improving the Efficiency of CMOS Image Sensors through In-Sensor Selective Attention
abstract
Inspired by the selective attention mechanism in human vision, we propose to introduce a saliency-based processing step in the CMOS image sensor, to continuously select pixels corresponding to salient objects and feedback such information to the sensor, instead of blindly passing all pixels to the sensor output. To minimize the overhead of saliency detection in this feedback loop, we propose two techniques: (1) saliency detection with low-precision, down-sampled grayscale images, and (2) Optimization of the loss function and model structure. Finally, we pad the minimum number of pixels around the selected pixels to maintain the accuracy of object detection (OD). Our method is experimented with two types of OD algorithms on three representative datasets. At the similar OD accuracy with the full image, our proposed selective feedback method successfully achieves 70.5% reduction in the volume of output pixels for BDD100K, which translates to 4.3× and 3.4× reduction in power consumption and latency, respectively.
Kishore Kasichainula, Dong-Woo Jee, Injune Yeo, Yaoxin Zhuo, Baoxin Li, Jae-sun Seo, Yu Cao 0001
ISCAS6
2023 Generative Alignment of Posterior Probabilities for Source-free Domain Adaptation
abstract
Existing domain adaptation literature comprises multiple techniques that align the labeled source and unlabeled target domains at different stages, and predict the target labels. In a source-free domain adaptation setting, the source data is not available for alignment. We present a source-free generative paradigm that captures the relations between the source categories and enforces them onto the unlabeled target data, thereby circumventing the need for source data without introducing any new hyper-parameters. The adaptation is performed through the adversarial alignment of the posterior probabilities of the source and target categories. The proposed approach demonstrates competitive performance against other source-free domain adaptation techniques and can also be used for source-present settings.
Sachin Chhabra, Hemanth Venkateswara, Baoxin Li
WACV3
2023 BMISP: Bidirectional mapping of image signal processing pipeline
Yahui Tang, Kan Chang, Mengyuan Huang, Baoxin Li
Signal Process.4
2022 PatchSwap: A Regularization Technique for Vision Transformers
Sachin Chhabra, Hemanth Venkateswara, Baoxin Li
BMVC3
2022 CLIP4Hashing: Unsupervised Deep Hashing for Cross-Modal Video-Text Retrieval
abstract
With the ever-increasing multimedia data on the Web, cross-modal video-text retrieval has received a lot of attention in recent years. Deep cross-modal hashing approaches utilize the Hamming space for achieving fast retrieval. However, most existing algorithms have difficulties in seeking or constructing a well-defined joint semantic space. In this paper, an unsupervised deep cross-modal video-text hashing approach (CLIP4Hashing) is proposed, which mitigates the difficulties in bridging between different modalities in the Hamming space through building a single hashing net by employing the pre-trained CLIP model. The approach is enhanced by two novel techniques, the dynamic weighting strategy and the design of the min-max hashing layer, which are found to be the main sources of the performance gain. Compared with conventional deep cross-modal hashing algorithms, CLIP4Hashing does not require data-specific hyper-parameters. With evaluation using three challenging video-text benchmark datasets, we demonstrate that CLIP4Hashing is able to significantly outperform existing state-of-the-art hashing algorithms. Additionally, with larger bit sizes (e.g., 2048 bits), CLIP4Hashing can even deliver competitive performance compared with the results based on non-hashing features.
Yaoxin Zhuo, Yikang Li 0001, Jenhao Hsiao, Chiuman Ho, Baoxin Li
ICMR5
2022 A Two-Stage Convolutional Neural Network for Joint Demosaicking and Super-Resolution
abstract
As two practical and important image processing tasks, color demosaicking (CDM) and super-resolution (SR) have been studied for decades. However, most literature studies these two tasks independently, ignoring the potential benefits of a joint solution. In this paper, aiming at efficient and effective joint demosaicking and super-resolution (JDSR), a well-designed two-stage convolutional neural network (CNN) architecture is proposed. For the first stage, by making use of the sampling-pattern information, a pattern-aware feature extraction (PFE) module extracts features directly from the Bayer-sampled low-resolution (LR) image, while keeping the resolution of the extracted features the same as the input. For the second stage, a dual-branch feature refinement (DFR) module effectively decomposes the features into two components with different spatial frequencies, on which different learning strategies are applied. On each branch of the DFR module, the feature refinement unit, namely, densely-connected dual-path enhancement blocks (DDEB), establishes a sophisticated nonlinear mapping from the LR space to the high-resolution (HR) space. To achieve strong representational power, two paths of transformations and the channel attention mechanism are adopted in DDEB. Extensive experiments demonstrate that the proposed method is superior to the sequential combination of state-of-the-art (SOTA) CDM and SR methods. Moreover, with much smaller model size, our approach also surpasses other SOTA JDSR methods.
Kan Chang, Hengxin Li, Yufei Tan, Pak Lun Kevin Ding, Baoxin Li
IEEE Trans. Circuits Syst. Video Technol.5
2021 Does a GAN leave distinct model-specific fingerprints?
Yuzhen Ding, Nupur Thakur, Baoxin Li
BMVC3
2021 Fedns: Improving Federated Learning for Collaborative Image Classification on Mobile Clients
abstract
Federated Learning (FL) is a paradigm that aims to support loosely connected clients in learning a global model collaboratively with the help of a centralized server. The most popular FL algorithm is Federated Averaging (FedAvg), which is based on taking weighted average of the client models, with the weights determined largely based on dataset sizes at the clients. In this paper, we propose a new approach, termed Federated Node Selection (FedNS), for the server’s global model aggregation in the FL setting. FedNS filters and re-weights the clients’ models at the node/kernel level, hence leading to a potentially better global model by fusing the best components of the clients. Using collaborative image classification as an example, we show with experiments from multiple datasets and networks that FedNS can consistently achieve improved performance over FedAvg.
Yaoxin Zhuo, Baoxin Li
ICME2
2021 Deep Latent Graph Matching
abstract
Deep learning for graph matching (GM) has emerged as an important research topic due to its superior performance over traditional methods and insights it provides for solving other combinatorial problems on graph. While recent deep methods for GM extensively investigated effective node/edge feature learning or downstream GM solvers given such learned features, there is little existing work questioning if the fixed connectivity/topology typically constructed using heuristics (e.g., Delaunay or k-nearest) is indeed suitable for GM. From a learning perspective, we argue that the fixed topology may restrict the model capacity and thus potentially hinder the performance. To address this, we propose to learn the (distribution of) latent topology, which can better support the downstream GM task. We devise two latent graph generation procedures, one deterministic and one generative. Particularly, the generative procedure emphasizes the across-graph consistency and thus can be viewed as a matching-guided co-generative model. Our methods deliver superior performance over previous state-of-the-arts on public benchmarks, hence supporting our hypothesis.
Tianshu Yu 0001, Runzhong Wang, Junchi Yan, Baoxin Li
ICML4
2021 Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
abstract
Machine learning (ML) models are widely used in many important domains. For efficiently processing these computational- and memory-intensive applications, tensors of these overparameterized models are compressed by leveraging sparsity, size reduction, and quantization of tensors. Unstructured sparsity and tensors with varying dimensions yield irregular computation, communication, and memory access patterns; processing them on hardware accelerators in a conventional manner does not inherently leverage acceleration opportunities. This article provides a comprehensive survey on the efficient execution of sparse and irregular tensor computations of ML models on hardware accelerators. In particular, it discusses enhancement modules in the architecture design and the software support, categorizes different hardware designs and acceleration techniques, analyzes them in terms of hardware and execution costs, analyzes achievable accelerations for recent DNNs, and highlights further opportunities in terms of hardware/software/model codesign optimizations (inter/intramodule). The takeaways from this article include the following: understanding the key challenges in accelerating sparse, irregular shaped, and quantized tensors; understanding enhancements in accelerator systems for supporting their efficient computations; analyzing tradeoffs in opting for a specific design choice for encoding, storing, extracting, communicating, computing, and load-balancing the nonzeros; understanding how structured sparsity can improve storage efficiency and balance computations; understanding how to compile and map models with sparse tensors on the accelerators; and understanding recent design trends for efficient accelerations and further opportunities.
Shail Dave, Riyadh Baghdadi, Tony Nowatzki, Sasikanth Avancha, Aviral Shrivastava, Baoxin Li
Proc. IEEE6
2020 Determinant Regularization for Gradient-Efficient Graph Matching
abstract
Graph matching refers to finding vertex correspondence for a pair of graphs, which plays a fundamental role in many vision and learning related tasks. Directly applying gradient-based continuous optimization on graph matching can be attractive for its simplicity but calls for effective ways of converting the continuous solution to the discrete one under the matching constraint. In this paper, we show a novel regularization technique with the tool of determinant analysis on the matching matrix which is relaxed into continuous domain with gradient based optimization. Meanwhile we present a theoretical study on the property of our relaxation technique. Our paper strikes an attempt to understand the geometric properties of different regularization techniques and the gradient behavior during the optimization. We show that the proposed regularization is more gradient-efficient than traditional ones during early update stages. The analysis will also bring about insights for other problems under bijection constraints. The algorithm procedure is simple and empirical results on public benchmark show its effectiveness on both synthetic and real-world data.
Tianshu Yu 0001, Junchi Yan, Baoxin Li
CVPR3
2020 RhyRNN: Rhythmic RNN for Recognizing Events in Long and Complex Videos
Tianshu Yu 0001, Yikang Li 0001, Baoxin Li
ECCV (10)3
2020 Deep Learning of Determinantal Point Processes via Proper Spectral Sub-gradient
Tianshu Yu 0001, Yikang Li 0001, Baoxin Li
ICLR3
2020 Learning deep graph matching with channel-independent embedding and Hungarian attention
Tianshu Yu 0001, Runzhong Wang, Junchi Yan, Baoxin Li
ICLR4
2020 Improving Batch Normalization with Skewness Reduction for Deep Neural Networks
abstract
Batch Normalization (BN) is a well-known technique used in training deep neural networks. The main idea behind batch normalization is to normalize the features of the layers (i.e., transforming them to have a mean equal to zero and a variance equal to one). Such a procedure encourages the optimization landscape of the loss function to be smoother, and improves the learning of the networks for both speed and performance. In this paper, we demonstrate that the performance of the network can be improved, if the distributions of the features of the output in the same layer are similar. As normalizing based on mean and variance does not necessarily make the features to have the same distribution, we propose a new normalization scheme: Batch Normalization with Skewness Reduction (BNSR). Comparing with other normalization approaches, BNSR transforms not just only the mean and variance, but also the skewness of the data. By tackling this property of a distribution, we are able to make the output distributions of the layers to be further similar. The nonlinearity of BNSR may further improve the expressiveness of the underlying network. Comparisons with other normalization schemes are tested on the CIFAR-100 and ImageNet datasets. Experimental results show that the proposed approach can outperform other state-of-the-arts that are not equipped with BNSR.
Pak Lun Kevin Ding, Sarah Martin, Baoxin Li
ICPR3
2020 Accurate single image super-resolution using multi-path wide-activated residual network
Kan Chang, Minghong Li, Pak Lun Kevin Ding, Baoxin Li
Signal Process.4
2020 Unsupervised Learning of Optical Flow With CNN-Based Non-Local Filtering
abstract
Estimating optical flow from successive video frames is one of the fundamental problems in computer vision and image processing. In the era of deep learning, many methods have been proposed to use convolutional neural networks (CNNs) for optical flow estimation in an unsupervised manner. However, the performance of unsupervised optical flow approaches is still unsatisfactory and often lagging far behind their supervised counterparts, primarily due to over-smoothing across motion boundaries and occlusion. To address these issues, in this paper, we propose a novel method with a new post-processing term and an effective loss function to estimate optical flow in an unsupervised, end-to-end learning manner. Specifically, we first exploit a CNN-based non-local term to refine the estimated optical flow by removing noise and decreasing blur around motion boundaries. This is implemented via automatically learning weights of dependencies over a large spatial neighborhood. Because of its learning ability, the method is effective for various complicated image sequences. Secondly, to reduce the influence of occlusion, a symmetrical energy formulation is introduced to detect the occlusion map from refined bi-directional optical flows. Then the occlusion map is integrated to the loss function. Extensive experiments are conducted on challenging datasets, i.e. FlyingChairs, MPI-Sintel and KITTI to evaluate the performance of the proposed method. The state-of-the-art results demonstrate the effectiveness of our proposed method.
Zhigang Tu 0001, Dejun Zhang, Jun Liu 0036, Baoxin Li, Junsong Yuan 0001
IEEE Trans. Image Process.5
2019 Weakly Supervised Deep Image Hashing Through Tag Embeddings
abstract
Many approaches to semantic image hashing have been formulated as supervised learning problems that utilize images and label information to learn the binary hash codes. However, large-scale labelled image data is expensive to obtain, thus imposing a restriction on the usage of such algorithms. On the other hand, unlabelled image data is abundant due to the existence of many Web image repositories. Such Web images may often come with images tags that contains useful information, although raw tags in general do not readily lead to semantic labels. Motivated by this scenario, we formulate the problem of semantic image hashing as a weakly-supervised learning problem. We utilize the information contained in the user-generated tags associated with the images to learn the hash codes. More specifically, we extract the word2vec semantic embeddings of the tags and use the information contained in them for constraining the learning. Accordingly, we name our model Weakly Supervised Deep Hashing using Tag Embeddings (WDHT). WDHT is tested for the task of semantic image retrieval and is compared against several state-of-art models. Results show that our approach sets a new state-of-art in the area of weekly supervised image hashing.
Vijetha Gattupalli, Yaoxin Zhuo, Baoxin Li
CVPR3
2019 Improving Robustness of Random Forest Under Label Noise
abstract
Random forest is a well-known and widely-used machine learning model. In many applications where the training data arise from real-world sources, there may be labeling errors in the data. In spite of its superior performance, the basic model of random forest dose not consider potential label noise in learning, and thus its performance can suffer significantly in the presence of label noise. In order to solve this problem, we present a new variation of random forest - a novel learning approach that leads to an improved noise robust random forest (NRRF) model. We incorporate the noise information by introducing a global multi-class noise tolerant loss function into the training of the classic random forest model. This new loss function was found to significantly boost the performance of random forest. We evaluated the proposed NRRF by extensive experiments of classification tasks on standard machine learning/computer vision datasets like MNIST, letter and Cifar10. The proposed NRRF produced very promising results under a wide range of noise settings.
Pak Lun Kevin Ding, Baoxin Li
WACV3
2019 Data-adaptive low-rank modeling and external gradient prior for single image super-resolution
Kan Chang, Xueyu Zhang, Pak Lun Kevin Ding, Baoxin Li
Signal Process.4
2019 A survey of variational and CNN-based optical flow techniques
Zhigang Tu 0001, Wei Xie 0008, Dejun Zhang, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Signal Process. Image Commun.6
2019 Semantic Cues Enhanced Multimodality Multistream CNN for Action Recognition
abstract
This paper addresses the issue of video-based action recognition by exploiting an advanced multistream convolutional neural network (CNN) to fully use semantics-derived multiple modalities in both spatial (appearance) and temporal (motion) domains, since the performance of the CNN-based action recognition methods heavily relates to two factors: semantic visual cues and the network architecture. Our work consists of two major parts. First, to extract useful human-related semantics accurately, we propose a novel spatiotemporal saliency-based video object segmentation (STS) model. By fusing different distinctive saliency maps, which are computed according to object signatures of complementary object detection approaches, a refined STS maps can be obtained. In this way, various challenges in the realistic video can be handled jointly. Based on the estimated saliency maps, an energy function is constructed to segment two semantic cues: the actor and one distinctive acting part of the actor. Second, we modify the architecture of the two-stream network (TS-Net) to design a multistream network that consists of three TS-Nets with respect to the extracted semantics, which is able to use deeper abstract visual features of multimodalities in multi-scale spatiotemporally. Importantly, the performance of action recognition is significantly boosted when integrating the captured human-related semantics into our framework. Experiments on four public benchmarks-JHMDB, HMDB51, UCF-Sports, and UCF101-demonstrate that the proposed method outperforms the state-of-the-art algorithms.
Zhigang Tu 0001, Wei Xie 0008, Justin Dauwels, Baoxin Li, Junsong Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 Action-Stage Emphasized Spatiotemporal VLAD for Video Action Recognition
abstract
Despite outstanding performance in image recognition, convolutional neural networks (CNNs) do not yet achieve the same impressive results on action recognition in videos. This is partially due to the inability of CNN for modeling long-range temporal structures especially those involving individual action stages that are critical to human action recognition. In this paper, we propose a novel action-stage (ActionS) emphasized spatiotemporal Vector of Locally Aggregated Descriptors (ActionS-STVLAD) method to aggregate informative deep features across the entire video according to adaptive video feature segmentation and adaptive segment feature sampling (AVFS-ASFS). In our ActionSST- VLAD encoding approach, by using AVFS-ASFS, the key frame features are chosen and the corresponding deep features are automatically split into segments with the features in each segment belonging to a temporally coherent ActionS. Then, based on the extracted key frame feature in each segment, a flow-guided warping technique is introduced to detect and discard redundant feature maps, while the informative ones are aggregated by using our exploited similarity weight. Furthermore, we exploit an RGBF modality to capture motion salient regions in the RGB images corresponding to action activity. Extensive experiments are conducted on four public benchmarks - HMDB51, UCF101, Kinetics and ActivityNet for evaluation. Results show that our method is able to effectively pool useful deep features spatiotemporally, leading to state-of-the-art performance for videobased action recognition.
Zhigang Tu 0001, Hongyan Li 0003, Dejun Zhang, Justin Dauwels, Baoxin Li, Junsong Yuan 0001
IEEE Trans. Image Process.5
2018 Simultaneous Event Localization and Recognition in Surveillance Video
abstract
The ubiquity of video-based surveillance demands automated approaches to analysis of ever-increasing video footages. Action/Event localization and recognition are two critical capabilities in surveillance video analysis, which have been largely addressed separately in the literature. In this paper, we propose an approach to simultaneously localize and recognize visual events from raw surveillance videos, employing an end-to-end learning strategy. Our approach formulates the task as weakly-supervised sequential semantic segmentation, in which we utilize a specific convolutional RNN to capture not only the appearance and the motion information but also their temporal evolution patterns. We tested our approach on the VIRAT 2.0 dataset. The experimental results, in comparison with relevant existing state-of-the-art, suggest that the proposed approach is promising in delivering a practical solution.
Yikang Li 0001, Tianshu Yu 0001, Baoxin Li
AVSS3
2018 Joint Cuts and Matching of Partitions in One Graph
abstract
As two fundamental problems, graph cuts and graph matching have been intensively investigated over the decades, resulting in vast literature in these two topics respectively. However the way of jointly applying and solving graph cuts and matching receives few attention. In this paper, we first formalize the problem of simultaneously cutting a graph into two partitions i.e. graph cuts and establishing their correspondence i.e. graph matching. Then we develop an optimization algorithm by updating matching and cutting alternatively, provided with theoretical analysis. The efficacy of our algorithm is verified on both synthetic dataset and real-world images containing similar regions or structures.
Tianshu Yu 0001, Junchi Yan, Jieyi Zhao, Baoxin Li
CVPR4
2018 Incremental Multi-graph Matching via Diversity and Randomness Based Graph Clustering
Tianshu Yu 0001, Junchi Yan, Wei Liu 0005, Baoxin Li
ECCV (13)4
2018 Generalizing Graph Matching beyond Quadratic Assignment Model
abstract
Graph matching has received persistent attention over decades, which can be formulated as a quadratic assignment problem (QAP). We show that a large family of functions, which we define as Separable Functions, can approximate discrete graph matching in the continuous domain asymptotically by varying the approximation controlling parameters. We also study the properties of global optimality and devise convex/concave-preserving extensions to the widely used Lawler's QAP form. Our theoretical findings show the potential for deriving new algorithms and techniques for graph matching. We deliver solvers based on two specific instances of Separable Functions, and the state-of-the-art performance of our method is verified on popular benchmarks.
Tianshu Yu 0001, Junchi Yan, Yilin Wang 0002, Wei Liu 0005, Baoxin Li
NeurIPS5
2018 Weakly Supervised Facial Attribute Manipulation via Deep Adversarial Network
abstract
Automatically manipulating facial attributes is challenging because it needs to modify the facial appearances, while keeping not only the person's identity but also the realism of the resultant images. Unlike the prior works on the facial attribute parsing, we aim at an inverse and more challenging problem called attribute manipulation by modifying a facial image in line with a reference facial attribute. Given a source input image and reference images with a target attribute, our goal is to generate a new image (i.e., target image) that not only possesses the new attribute but also keeps the same or similar content with the source image. In order to generate new facial attributes, we train a deep neural network with a combination of a perceptual content loss and two adversarial losses, which ensure the global consistency of the visual content while implementing the desired attributes often impacting on local pixels. The model automatically adjusts the visual attributes on facial appearances and keeps the edited images as realistic as possible. The evaluation shows that the proposed model can provide a unified solution to both local and global facial attribute manipulation such as expression change and hair style transfer. Moreover, we further demonstrate that the learned attribute discriminator can be used for attribute localization.
Yilin Wang 0002, Suhang Wang, Guo-Jun Qi, Jiliang Tang, Baoxin Li
WACV5
2018 Multi-stream CNN: Learning representations based on human-related regions for action recognition
Zhigang Tu 0001, Wei Xie 0008, Qianqing Qin, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Pattern Recognit.6
2018 Single image super-resolution using collaborative representation and non-local self-similarity
Kan Chang, Pak Lun Kevin Ding, Baoxin Li
Signal Process.3
2018 Single Image Super Resolution Using Joint Regularization
abstract
This letter proposes a reconstruction-based single image super resolution method by using joint regularization, where a group-residual-based regularization (GRR) and a ridge-regression-based regularization (3R) are combined. In GRR, nonlocal similar patches are grouped together, and the group weights are calculated so as to adaptively constrain the residual values in the gradient domain. In 3R, we adopt the ridge-regression-based method to establish the projection matrices from an external high-resolution (HR) training set, so that the external HR information can be utilized. To obtain an estimation of the targeted HR image, an efficient algorithm is designed for solving the joint formulation. Experimental results on different image datasets indicate that the proposed method is able to achieve the state-of-the-art performance.
Kan Chang, Pak Lun Kevin Ding, Baoxin Li
IEEE Signal Process. Lett.3
2017 CLARE: A Joint Approach to Label Classification and Tag Recommendation
abstract
Data classification and tag recommendation are both important and challenging tasks in social media. These two tasks are often considered independently and most efforts have been made to tackle them separately. However, labels in data classification and tags in tag recommendation are inherently related. For example, a Youtube video annotated with NCAA, stadium, pac12 is likely to be labeled as football, while a video/image with the class label of coast is likely to be tagged with beach, sea, water and sand. The existence of relations between labels and tags motivates us to jointly perform classification and tag recommendation for social media data in this paper. In particular, we provide a principled way to capture the relations between labels and tags, and propose a novel framework CLARE, which fuses data CLAssification and tag REcommendation into a coherent model. With experiments on three social media datasets, we demonstrate that the proposed framework CLARE achieves superior performance on both tasks compared to the state-of-the-art methods.
Yilin Wang 0002, Suhang Wang, Jiliang Tang, Guo-Jun Qi, Huan Liu 0001, Baoxin Li
AAAI6
2017 Convex dictionary learning for single image super-resolution
abstract
In recent years, dictionary learning approaches have been used in image super-resolution, achieving promising results. Such approaches train a dictionary from image patches and reconstruct a new patch by sparse combination of the atoms of the dictionary. Typical training methods do not constrain the dictionary atoms. In this paper, we propose a convex dictionary learning (CDL) algorithm by constraining the dictionary atoms to be formed by non-negative linear combination of the training data, which is a natural, desired property. We evaluate our approach by demonstrating its performance gain over typical approaches.
Pak Lun Kevin Ding, Baoxin Li, Kan Chang
ICIP2
2017 Non-negative dictionary learning with pairwise partial similarity constraint
abstract
Discriminative dictionary learning has been widely used in many applications such as face retrieval / recognition and image classification, where the labels of the training data are utilized to improve the discriminative power of the learned dictionary. This paper deals with a new problem of learning a dictionary for associating pairs of images in applications such as face image retrieval. Compared with a typical supervised learning task, in this case the labeling information is very limited (e.g. only some training pairs are known to be associated). Further, associated pairs may be considered similar only after excluding certain regions (e.g. sunglasses in a face image). We formulate a dictionary learning problem under these considerations and design an algorithm to solve the problem. We also provide a proof for the convergence of the algorithm. Experiments and results suggest that the proposed method is advantageous over common baselines.
Pak Lun Kevin Ding, Baoxin Li
ICME3
2017 Embedded Supervised Feature Selection for Multi-class Data
abstract
Supervised multi-class learning arises in many application domains such as biology, computer vision, social network analysis, and information retrieval. These applications often involve high-dimensional data, which not only significantly increase the time and space requirement of the underlying algorithms but also degrade their performance due to the curse of dimensionality. Feature selection has been proven effective and efficient for preparing high-dimensional data for many learning tasks. Traditional feature selection algorithms for multi-class data assume the independence of label categories and select features with the capability to distinguish samples from different classes. However, class labels in multi-class data may be correlated and little work exists for exploiting label correlation in multi-class feature selection. In this paper, we investigate label correlation in feature selection for multi-class data. In particular, we provide a principled approach for capturing label correlation and propose an Embedded Supervised Feature Selection (ESFS) framework, which embeds label correlation modeling in supervised feature selection for multi-class data. Experiments on both synthetic data and various types of public benchmark datasets show that the proposed framework effectively captures the multi-class label correlation and significantly outperforms existing state-of-the-art baseline methods.
Lin Chen 0010, Jiliang Tang, Baoxin Li
SDM3
2017 Joint Regression and Ranking for Image Enhancement
abstract
Research on automated image enhancement has gained momentum in recent years, partially due to the need for easy-to-use tools for enhancing pictures captured by ubiquitous cameras on mobile devices. Many of the existing leading methods employ machine-learning-based techniques, by which some enhancement parameters for a given image are found by relating the image to the training images with known enhancement parameters. While knowing the structure of the parameter space can facilitate search for the optimal solution, none of the existing methods has explicitly modeled and learned that structure. This paper presents an end-to-end, novel joint regression and ranking approach to model the interaction between desired enhancement parameters and images to be processed, employing a Gaussian process (GP). GP allows searching for ideal parameters using only the image features. The model naturally leads to a ranking technique for comparing images in the induced feature space. Comparative evaluation using the ground-truth based on the MIT-Adobe FiveK dataset plus subjective tests on an additional data-set were used to demonstrate the effectiveness of the proposed approach.
Parag S. Chandakkar, Baoxin Li
WACV2
2017 Understanding and Discovering Deliberate Self-harm Content in Social Media
abstract
Studies suggest that self-harm users found it easier to discuss self-harm-related thoughts and behaviors using social media than in the physical world. Given the enormous and increasing volume of social media data, on-line self-harm content is likely to be buried rapidly by other normal content. To enable voices of self-harm users to be heard, it is important to distinguish self-harm content from other types of content. In this paper, we aim to understand self-harm content and provide automatic approaches to its detection. We first perform a comprehensive analysis on self-harm social media using different input cues. Our analysis, the first of its kind in large scale, reveals a number of important findings. Then we propose frameworks that incorporate the findings to discover self-harm content under both supervised and unsupervised settings. Our experimental results on a large social media dataset from Flickr demonstrate the effectiveness of the proposed frameworks and the importance of our findings in discovering self-harm content.
Yilin Wang 0002, Jiliang Tang, Jundong Li, Baoxin Li, Yali Wan, Clayton Mellina, Neil O'Hare, Yi Chang 0001
WWW4
2017 Fusing disparate object signatures for salient object detection in video
Zhigang Tu 0001, Zuwei Guo, Wei Xie 0008, Mengjia Yan 0003, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Pattern Recognit.6
2016 Weakly hierarchical lasso based learning to rank in best answer prediction
abstract
In community question and answering sites, pairs of questions and their high-quality answers (like best answers selected by askers) can be valuable knowledge available to others. However lots of questions receive multiple answers but askers do not label either one as the accepted or best one even when some replies answer their questions. To solve this problem, high-quality answer prediction or best answer prediction has been one of important topics in social media. These user-generated answers often consist of multiple “views”, each capturing different (albeit related) information (e.g., expertise of the asker, length of the answer, etc.). Such views interact with each other in complex manners that should carry a lot of information for distinguishing a potential best answer from others. Little existing work has exploited such interactions for better prediction. To explicitly model these information, we propose a new learning-to-rank method, ranking support vector machine (RankSVM) with weakly hierarchical lasso in this paper. The evaluation of the approach was done using data from Stack Overflow. Experimental results demonstrate that the proposed approach has superior performance compared with approaches in state-of-the-art.
Qiongjie Tian, Baoxin Li
ASONAM2
2016 Finding needles of interested tweets in the haystack of Twitter network
abstract
Drug use and abuse is a serious societal problem. The fast development and adoption of social media and smart mobile devices in recent years bring about new opportunities for advancing computer-based strategies for understanding and intervention of drug-related behaviors. However, the existing literature still lacks principled ways of building computational models for supporting effective analysis of large-scale, often unstructured social media data. Part of the challenge stems from the difficulty of obtaining so-called ground-truth data that are typically required for training computational models. This paper presents a progressive semi-supervised learning approach to identifying Twitter tweets that are related to personal and recreational use of marijuana. Based on a small, labeled dataset, the proposed approach first learns optimal mapping of raw features from the tweets for classification, using a method of weakly hierarchical lasso. The learned feature model is then used to support unsupervised clustering of Web-scale data. Experiments with realistic data crawled from Twitter are used to validate the proposed approach, demonstrating its effectiveness.
Qiongjie Tian, Jashmi Lagisetty, Baoxin Li
ASONAM3
2016 PPP: Joint Pointwise and Pairwise Image Label Prediction
abstract
Pointwise label and pairwise label are both widely used in computer vision tasks. For example, supervised image classification and annotation approaches use pointwise label, while attribute-based image relative learning often adopts pairwise labels. These two types of labels are often considered independently and most existing efforts utilize them separately. However, pointwise labels in image classification and tag annotation are inherently related to the pairwise labels. For example, an image labeled with "coast" and annotated with "beach, sea, sand, sky" is more likely to have a higher ranking score in terms of the attribute "open", while "men shoes" ranked highly on the attribute "formal" are likely to be annotated with "leather, lace up" than "buckle, fabric". The existence of potential relations between pointwise labels and pairwise labels motivates us to fuse them together for jointly addressing related vision tasks. In particular, we provide a principled way to capture the relations between class labels, tags and attributes, and propose a novel framework PPP(Pointwise and Pairwise image label Prediction), which is based on overlapped group structure extracted from the pointwise-pairwise-label bipartite graph. With experiments on benchmark datasets, we demonstrate that the proposed framework achieves superior performance on three vision tasks compared to the state-of-the-art methods.
Yilin Wang 0002, Suhang Wang, Jiliang Tang, Huan Liu 0001, Baoxin Li
CVPR5
2016 Recognizing unseen actions in a domain-adapted embedding space
abstract
With the sustaining bloom of multimedia data, Zero-shot Learning (ZSL) techniques have attracted much attention in recent years for its ability to train learning models that can handle “unseen” categories. Existing ZSL algorithms mainly take advantages of attribute-based semantic space and only focus on static image data. Besides, most ZSL studies merely consider the semantic embedded labels and fail to address domain shift problem. In this paper, we purpose a deep two-output model for video ZSL and action recognition tasks by computing both spatial and temporal features from video contents through distinct Convolutional Neural Networks (CNNs) and training a Multi-layer Perceptron (MLP) upon extracted features to map videos to semantic embedding word vectors. Moreover, we introduce a domain adaptation strategy named “ConSSEV” - by combining outputs from two distinct output layers of our MLP to improve the results of zero-shot learning. Our experiments on UCF101 dataset demonstrate the purposed model has more advantages associated with more complex video embedding schemes, and outperforms the state-of-the-art zero-shot learning techniques.
Yikang Li 0001, Sheng-hung Hu, Baoxin Li
ICIP3
2016 On the generality of neural image features
abstract
Often the filters learned by Convolutional Neural Networks (CNNs) from different image datasets appear similar. This similarity of filters is often exploited for the purposes of transfer learning. This is also being used as an initialization technique for different tasks in the same dataset or for the same task in similar datasets. Off-the-shelf CNN features have capitalized on this idea to promote their networks as best transferable and most general and are used in a cavalier manner in day-to-day computer vision tasks. While the filters learned by these CNNs are related to the atomic structures of the images from which they are learnt, all datasets learn similar looking low-level filters. With the understanding that a dataset that contains many such atomic structures learn general filters and are therefore useful to initialize other networks with, we propose a way to analyse and quantify generality. We applied this metric on several popular character recognition, natural image and a medical image dataset, and arrive at some interesting conclusions. On further experimentation we also discovered that particular classes in a dataset themselves are more general than others.
Ragav Venkatesan, Vijetha Gattupalli, Baoxin Li
ICIP3
2016 Scale-adaptive EigenEye for fast eye detection in wild web images
abstract
Detecting eyes in images is fundamental for many computer vision applications including face detection, face recognition, and human-computer interaction. Most existing methods are designed and tested on datasets acquired under controlled lab settings (e.g., fixed scale, known poses, clean background, etc.), leaving their performance to be further examined on real-world, uncontrolled images, such as on-line images. This paper presents an effort on developing a fast and accurate eye detector for on-line images for which the acquisition condition is unknown and varies from one image to another, resulting in unpredictable background and variable scales for the eyes/faces. The key idea is to develop a scale-adaptive EigenEye approach, which employs an approximate scale estimated from face detection to modulate the pre-trained EigenEye basis in searching for the best match in a test image. The effort also includes building a 2845-image dataset with accurately-annotated eye locations and size, which will be made public to the community for future comparative study. Evaluation using this dataset, with comparison with a few leading state-of-the-art approaches, demonstrates the advantages of the proposed method.
Yilin Wang 0002, Peng Zhang 0017, Baoxin Li
ICIP4
2016 MSR-CNN: Applying motion salient region based descriptors for action recognition
abstract
In recent years the most popular video-based human action recognition methods rely on extracting feature representations using Convolutional Neural Networks (CNN) and then using these representations to classify actions. In this work, we propose a fast and accurate video representation that is derived from the motion-salient region (MSR), which represents features most useful for action labeling. By improving a well-performed foreground detection technique, the region of interest (ROI) corresponding to actors in the foreground in both the appearance and the motion field can be detected under various realistic challenges. Furthermore, we propose a complementary motion salient measure to select a secondary ROI - the major moving part of the human. Accordingly, a MSR-based CNN descriptor (MSR-CNN) is formulated to recognize human action, where the descriptor incorporates appearance and motion features along with tracks of MSR. The computation can be efficiently implemented due to two characteristics: 1) only part of the RGB image and the motion field need to be processed; 2) less data is used as input for the CNN feature extraction. Comparative evaluation on JHMDB and UCF Sports datasets shows that our method outperforms the state-of-the-art in both efficiency and accuracy.
Zhigang Tu 0001, Yikang Li 0001, Baoxin Li
ICPR4
2016 A computational approach to relative aesthetics
abstract
Computational visual aesthetics has recently become an active research area. Existing state-of-art methods formulate this as a binary classification task where a given image is predicted to be beautiful or not. In many applications such as image retrieval and enhancement, it is more important to rank images based on their aesthetic quality instead of binary-categorizing them. Furthermore, in such applications, it may be possible that all images belong to the same category. Hence determining the aesthetic ranking of the images is more appropriate. To this end, we formulate a novel problem of ranking images with respect to their aesthetic quality. We construct a new dataset of image pairs with relative labels by carefully selecting images from the popular AVA dataset. Unlike in aesthetics classification, there is no single threshold which would determine the ranking order of the images across our entire dataset. We propose a deep neural network based approach that is trained on image pairs by incorporating principles from relative learning. Results show that such relative training procedure allows our network to rank the images with a higher accuracy than a state-of-art network trained on the same set of images using binary labels.
Vijetha Gattupalli, Parag S. Chandakkar, Baoxin Li
ICPR3
2016 Video2vec: Learning semantic spatio-temporal embeddings for video representation
abstract
We propose to learn semantic spatio-temporal embeddings for videos to support high-level video analysis. The first step of the proposed embedding employs a deep architecture consisting of two channels of convolutional neural networks (capturing appearance and local motion) followed by their corresponding Gated Recurrent Unit encoders for capturing longer-term temporal structure of the CNN features. The resultant spatio-temporal representation (a vector) is used to learn a mapping via a multilayer perceptron to the word2vec semantic embedding space, leading to a semantic interpretation of the video vector that supports high-level analysis. We demonstrate the usefulness and effectiveness of this new video representation by experiments on action recognition, zero-shot video classification, and “word-to-video” retrieval, using the UCF-101 dataset.
Sheng-hung Hu, Yikang Li 0001, Baoxin Li
ICPR3
2016 Clustering-Based Joint Feature Selection for Semantic Attribute Prediction
Lin Chen 0010, Baoxin Li
IJCAI2
2016 A structured approach to predicting image enhancement parameters
abstract
Social networking on mobile devices has become a commonplace of everyday life. In addition, photo capturing process has become trivial due to the advances in mobile imaging. Hence people capture a lot of photos everyday and they want them to be visually-attractive. This has given rise to automated, one-touch enhancement tools. However, the inability of those tools to provide personalized and content-adaptive enhancement has paved way for machine-learned methods to do the same. The existing typical machine-learned methods heuristically (e.g. kNN-search) predict the enhancement parameters for a new image by relating the image to a set of similar training images. These heuristic methods need constant interaction with the training images which makes the parameter prediction sub-optimal and computationally expensive at test time which is undesired. This paper presents a novel approach to predicting the enhancement parameters given a new image using only its features, without using any training images. We propose to model the interaction between the image features and its corresponding enhancement parameters using the matrix factorization (MF) principles. We also propose a way to integrate the image features in the MF formulation. We show that our approach outperforms heuristic approaches as well as recent approaches in MF and structured prediction on synthetic as well as real-world data of image enhancement.
Parag S. Chandakkar, Baoxin Li
WACV2
2016 Simultaneous semantic segmentation of a set of partially labeled images
abstract
Semantic segmentation, by which an image is decomposed into regions with their respective semantic labels, is often the first step towards image understanding. Existing research on this regard is mainly performed under two conditions: the fully-supervised setting that relies on a set of images with pixel-level labels and the weakly-supervised one that uses only image-level labels. In both cases, the labeling task is time-consuming and laborious, and thus training data are always limited. In practice, there are voluminous on-line images, which unfortunately often have only incomplete image-level labels (tags) but would otherwise be potentially useful for a learning-based algorithm. Only limited efforts have been attempted on using such coarsely and incompletely labelled data for semantic segmentation. This paper proposes a new approach to semantic segmentation of a set of partially-labelled images, using a formulation considering information from multiple visual similar images. Experiments on several popular datasets, with comparison with existing methods, demonstrate evident performance improvement of the proposed approach.
Qiongjie Tian, Baoxin Li
WACV2
2016 Efficient unsupervised abnormal crowd activity detection based on a spatiotemporal saliency detector
abstract
Approaches to abnormality detection in crowded scene largely rely on supervised methods using discriminative models. In this paper, we presents a novel and efficient unsupervised learning method for video analysis. We start from visual saliency, which has been used in several vision tasks, e.g., image classification, object detection, and foreground segmentation. To detect saliency regions in video sequences, we propose a new approach for detecting spatiotemporal visual saliency based on the phase spectrum of the videos, which is easy to implement and computationally efficient. With the proposed algorithm, we also study how the spatiotemporal saliency can be used in two important vision tasks, saliency prediction and abnormality detection. The proposed algorithm is evaluated on several benchmark datasets with comparison to the state-of-the-art methods from the literature. The experiments demonstrate the effectiveness of the proposed approach to spatiotemporal visual saliency detection and its application to the above vision tasks.
Yilin Wang 0002, Qiang Zhang 0021, Baoxin Li
WACV3
2016 Affordable, web-based surgical skill training and evaluation tool
Gazi Islam, Kanav Kahol, Baoxin Li, Marshall L. Smith, Vimla L. Patel
J. Biomed. Informatics3
2016 Compressive Sensing Reconstruction of Correlated Images Using Joint Regularization
abstract
This letter proposes a novel compressive sensing reconstruction method for correlated images by using joint regularization, where a compensation-based adaptive total variation (CATV) regularization and a multi-image nonlocal low-rank (MNLR) regularization are included. In CATV, local weights are assigned to the residual values in the gradient domain so as to constrain the regularization strength at each pixel. In MNLR, the search of similar patches goes across different images so that both self-similarity and inter-image similarity are explored. Afterward, an efficient algorithm is proposed to solve the joint formulation, using a Split-Bregman-based technique. The effectiveness of the proposed approach is demonstrated with experiments on both multiview images and video sequences.
Kan Chang, Pak Lun Kevin Ding, Baoxin Li
IEEE Signal Process. Lett.3
2015 Simpler Non-Parametric Methods Provide as Good or Better Results to Multiple-Instance Learning
abstract
Multiple-instance learning (MIL) is a unique learning problem in which training data labels are available only for collections of objects (called bags) instead of individual objects (called instances). A plethora of approaches have been developed to solve this problem in the past years. Popular methods include the diverse density, MILIS and DD-SVM. While having been widely used, these methods, particularly those in computer vision have attempted fairly sophisticated solutions to solve certain unique and particular configurations of the MIL space. In this paper, we analyze the MIL feature space using modified versions of traditional non-parametric techniques like the Parzen window and k-nearest-neighbour, and develop a learning approach employing distances to k-nearest neighbours of a point in the feature space. We show that these methods work as well, if not better than most recently published methods on benchmark datasets. We compare and contrast our analysis with the well-established diverse-density approach and its variants in recent literature, using benchmark datasets including the Musk, Andrews' and Corel datasets, along with a diabetic retinopathy pathology diagnosis dataset. Experimental results demonstrate that, while enjoying an intuitive interpretation and supporting fast learning, these method have the potential of delivering improved performance even for complex data arising from real-world applications.
Ragav Venkatesan, Parag S. Chandakkar, Baoxin Li
ICCV3
2015 Real-time vehicle back-up warning system with a single camera
abstract
In this paper, we propose a real-time system using vehicle back-up camera to alert for potential back-up collisions. We developed a highly efficient algorithm, combining segmenting pedestrians and vehicles from moving background using local optical flow value, and a scale adaptive method using Deformable Part Model to detect objects at different distances. To test out algorithm, we created our own vehicle back-up dataset that contains rich scenes recorded from a back-up camera on moving/stationary vehicles with unique and challenging scenarios such as frequent occlusion with cluttered and moving background, and we made this dataset available to public for other researchers. Experiments on the dataset shows that our algorithm achieves high accuracy in near real-time, and it is about 10 times faster than the comparable state-of-the-art algorithm.
Yilin Wang 0002, Baoxin Li
ICIP3
2015 Relative learning from web images for content-adaptive enhancement
abstract
Personalized and content-adaptive image enhancement can find many applications in the age of social media and mobile computing. This paper presents a relative-learning-based approach, which, unlike previous methods, does not require matching original and enhanced images for training. This allows the use of massive online photo collections to train a ranking model for improved enhancement. We first propose a multi-level ranking model, which is learned from only relatively-labeled inputs that are automatically crawled. Then we design a novel parameter sampling scheme under this model to generate the desired enhancement parameters for a new image. For evaluation, we first verify the effectiveness and the generalization abilities of our approach, using images that have been enhanced/labeled by experts. Then we carry out subjective tests, which show that users prefer images enhanced by our approach over other existing methods.
Parag S. Chandakkar, Qiongjie Tian, Baoxin Li
ICME3
2015 Instructive video retrieval for surgical skill coaching using attribute learning
abstract
Video-based coaching systems have seen increasing adoption in various applications including dance, sports, and surgery training. Most existing systems are either passive (for data capture only) or barely active (with limited automated feedback to a trainee). In this paper, we present a video-based skill coaching system for simulation-based surgical training by exploring a newly proposed problem of instructive video retrieval. By introducing attribute learning into video for high-level skill understanding, we aim at providing automated feedback and providing an instructive video, to which the trainees can refer for performance improvement. This is achieved by ensuring the feedback is weakness-specific, skill-superior and content-similar. A suite of techniques was integrated to build the coaching system with these features. In particular, algorithms were developed for action segmentation, video attribute learning, and attribute-based video retrieval. Experiments with realistic surgical videos demonstrate the feasibility of the proposed method and suggest areas for further improvement.
Lin Chen 0010, Qiang Zhang 0021, Peng Zhang 0017, Baoxin Li
ICME4
2015 Structure-preserving Image Quality Assessment
abstract
Perceptual Image Quality Assessment (IQA) has many applications. Existing IQA approaches typically work only for one of three scenarios: full-reference, non-reference, or reduced-reference. Techniques that attempt to incorporate image structure information often rely on hand-crafted features, making them difficult to be extended to handle different scenarios. On the other hand, objective metrics like Mean Square Error (MSE), while being easy to compute, are often deemed ineffective for measuring perceptual quality. This paper presents a novel approach to perceptual quality assessment by developing an MSE-like metric, which enjoys the benefit of MSE in terms of inexpensive computation and universal applicability while allowing structural information of an image being taken into consideration. The latter was achieved through introducing structure-preserving kernelization into a MSE-like formulation. We show that the method can lead to competitive FR-IQA results. Further, by developing a feature coding scheme based on this formulation, we extend the model to improve the performance of NR-IQA methods. We report extensive experiments illustrating the results from both our FR-IQA and NR-IQA algorithms with comparison to existing state-of-the-art methods.
Yilin Wang 0002, Qiang Zhang 0021, Baoxin Li
ICME3
2015 Inferring Sentiment from Web Images with Joint Inference on Visual and Social Cues: A Regulated Matrix Factorization Approach
Yilin Wang 0002, Yuheng Hu, Subbarao Kambhampati, Baoxin Li
ICWSM4
2015 Unsupervised Sentiment Analysis for Social Media Images
Yilin Wang 0002, Suhang Wang, Jiliang Tang, Huan Liu 0001, Baoxin Li
IJCAI5
2015 Fusing Pointwise and Pairwise Labels for Supporting User-adaptive Image Retrieval
abstract
User-adaptive image retrieval/recommendation has drawn a lot of research interests in recent years, owing to fast development of various Web applications where retrieving images is a key enabling task. Existing challenges include the lack of user-adaptive training data, the ambiguity of user query and the real-time interactivity of a system. This paper proposes a hybrid learning strategy that fuses knowledge from both pointwise and pairwise training data into one framework for attribute-based, user-adaptive image retrieval. Under this framework, we develop an online learning algorithm for updating the ranking performance based on user feedback. Furthermore, we derive the framework into a kernel form, allowing easy application of kernel techniques. The proposed approach is evaluated on two image datasets and experimental results show that it achieves obvious performance gains over ranking and zero-shot learning from either type of training data independently. In addition, the online learning algorithm is able to deliver much better performance than batch learning, given the same elapsed running time, or can achieve better performance in much less time.
Lin Chen 0010, Peng Zhang 0017, Baoxin Li
ICMR3
2015 Retrieving Unfamiliar Faces: Towards Understanding Human Performance
abstract
Face image retrieval is to find from a dataset all images containing the same person in the query image. Automatic face retrieval has seen fast development in recent years, although humans still appear to be the better performer on this task. This paper reports a study towards understanding human performance on retrieving unfamiliar faces. Wild Web face images are utilized in the study, and two experiments are designed to assess human performance and behavior on the retrieval task. The experiments help to identify a set of important features and also to understand how human behaved when facing the task of retrieving unfamiliar faces. Such observations/conclusions may provide guidelines for improving existing automated algorithms.
Baoxin Li
ACM Multimedia2
2015 Improving Vision-Based Self-Positioning in Intelligent Transportation Systems via Integrated Lane and Vehicle Detection
abstract
Traffic congestion is a widespread problem. Dynamic traffic routing systems and congestion pricing are getting importance in recent research. Lane prediction and vehicle density estimation is an important component of such systems. We introduce a novel problem of vehicle self positioning which involves predicting the number of lanes on the road and vehicle's position in those lanes using videos captured by a dashboard camera. We propose an integrated closed-loop approach where we use the presence of vehicles to aid the task of self-positioning and vice versa. To incorporate multiple factors and high-level semantic knowledge into the solution, we formulate this problem as a Bayesian framework. In the framework, the number of lanes, the vehicle's position in those lanes and the presence of other vehicles are considered as parameters. We also propose a bounding box selection scheme to reduce the number of false detections and increase the computational efficiency. We show that the number of box proposals decreases by a factor of 6 using the selection approach. It also results in large reduction in the number of false detections. The entire approach is tested on real-world videos and is found to give acceptable results.
Parag S. Chandakkar, Yilin Wang 0002, Baoxin Li
WACV3
2015 Joint modeling and reconstruction of a compressively-sensed set of correlated images
Kan Chang, Baoxin Li
J. Vis. Commun. Image Represent.2
2015 Relative Hidden Markov Models for Video-Based Evaluation of Motion Skills in Surgical Training
abstract
A proper temporal model is essential to analysis tasks involving sequential data. In computer-assisted surgical training, which is the focus of this study, obtaining accurate temporal models is a key step towards automated skill-rating. Conventional learning approaches can have only limited success in this domain due to insufficient amount of data with accurate labels. We propose a novel formulation termed Relative Hidden Markov Model and develop algorithms for obtaining a solution under this formulation. The method requires only relative ranking between input pairs, which are readily available from training sessions in the target application, hence alleviating the requirement on data labeling. The proposed algorithm learns a model from the training data so that the attribute under consideration is linked to the likelihood of the input, hence supporting comparing new sequences. For evaluation, synthetic data are first used to assess the performance of the approach, and then we experiment with real videos from a widely-adopted surgical training platform. Experimental results suggest that the proposed approach provides a promising solution to video-based motion skill evaluation. To further illustrate the potential of generalizing the method to other applications of temporal analysis, we also report experiments on using our model on speech-based emotion recognition.
Qiang Zhang 0021, Baoxin Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Color image demosaicking using inter-channel correlation and nonlocal self-similarity
Kan Chang, Pak Lun Kevin Ding, Baoxin Li
Signal Process. Image Commun.3
2014 Image Cosegmentation via Multi-task Learning
Qiang Zhang 0021, Yilin Wang 0002, Jieping Ye, Baoxin Li
BMVC5
2014 Predicting Multiple Attributes via Relative Multi-task Learning
abstract
Relative attributes learning aims to learn ranking functions describing the relative strength of attributes. Most of current learning approaches learn ranking functions for each attribute independently without considering possible intrinsic relatedness among the attributes. For a problem involving multiple attributes, it is reasonable to assume that utilizing such relatedness among the attributes would benefit learning, especially when the number of labeled training pairs are very limited. In this paper, we proposed a relative multi-attribute learning framework that integrates relative attributes into a multi-task learning scheme. The formulation allows us to exploit the advantages of the state-of-the-art regularization-based multi-task learning for improved attribute learning. In particular, using joint feature learning as the case studies, we evaluated our framework with both synthetic data and two real datasets. Experimental results suggest that the proposed framework has clear performance gain in ranking accuracy and zero-shot learning accuracy over existing methods of independent relative attributes learning and multi-task learning.
Lin Chen 0010, Qiang Zhang 0021, Baoxin Li
CVPR3
2014 Instructive Video Retrieval Based on Hybrid Ranking and Attribute Learning: A Case Study on Surgical Skill Training
abstract
Video-based systems have been increasingly used in various training tasks in applications like sports, dancing, and surgery. One key task to add automation to such systems is to automatically select reference videos for a given training video of a trainee. In this paper, we formulate a new problem of instructive video retrieval and propose a solution using both attribute learning and learning to rank. The method first evaluates a user's skill attributes by relative attribute learning. Then, the most critical skill attribute in need of improvement is selected and reported to the user. Finally, a hybrid ranking learning to rank method is employed to retrieve instructive videos from a dataset, which serve as reference for the user. Two main technical problems are solved in this method. First, we combine both skill and visual feature to characterize skill superiority and context similarity. Second, we propose a hybrid ranking approach that works with both pair-wise and point-wise labels of the data. The benefit of the proposed method over other heuristic methods is demonstrated by both objective and subjective experiments, using surgical training videos as a case study.
Lin Chen 0010, Peng Zhang 0017, Baoxin Li
ACM Multimedia3
2014 Max-Margin Multiattribute Learning With Low-Rank Constraint
abstract
Attribute learning has attracted a lot of interests in recent years for its advantage of being able to model high-level concepts with a compact set of midlevel attributes. Real-world objects often demand multiple attributes for effective modeling. Most existing methods learn attributes independently without explicitly considering their intrinsic relatedness. In this paper, we propose max margin multiattribute learning with low-rank constraint, which learns a set of attributes simultaneously, using only relative ranking of the attributes for the data. By learning all the attributes simultaneously through low-rank constraint, the proposed method is able to capture their intrinsic correlation for improved learning; by requiring only relative ranking, the method avoids restrictive binary labels of attributes that are often assumed by many existing techniques. The proposed method is evaluated on both synthetic data and real visual data including a challenging video data set. Experimental results demonstrate the effectiveness of the proposed method.
Qiang Zhang 0021, Lin Chen 0010, Baoxin Li
IEEE Trans. Image Process.3
2013 On predicting Twitter trend: factors and models
abstract
In this paper, we predict hashtag trend in Twitter network with two basic issues under investigation, i.e. trend factors and prediction models. To address the first issue, we consider different content and context factors by designing features from tweet messages, network topology, user behavior, etc. To address the second issue, we adopt prediction models that have different combinations of the two basic model properties, i.e. linearity and state-space. Experiments on large Twitter dataset show that both content and context factors can help trend prediction. However, the most relevant factors are derived from user behaviors on the specific trend. Non-linear models are significantly better than their linear counterparts, which can be further slightly improved by the adoption of state-space models.
Peng Zhang 0017, Xufei Wang, Baoxin Li
ASONAM3
2013 Relative Hidden Markov Models for Evaluating Motion Skill
abstract
This paper is concerned with a novel problem: learning temporal models using only relative information. Such a problem arises naturally in many applications involving motion or video data. Our focus in this paper is on videobased surgical training, in which a key task is to rate the performance of a trainee based on a video capturing his motion. Compared with the conventional method of relying on ratings from senior surgeons, an automatic approach to this problem is desirable for its potential lower cost, better objectiveness, and real-time availability. To this end, we propose a novel formulation termed Relative Hidden Markov Model and develop an algorithm for obtaining a solution under this model. The proposed method utilizes only a relative ranking (based on an attribute of interest) between pairs of the inputs, which is easier to obtain and often more consistent, especially for the chosen application domain. The proposed algorithm effectively learns a model from the training data so that the attribute under consideration is linked to the likelihood of the inputs under the learned model. Hence the model can be used to compare new sequences. Synthetic data is first used to systematically evaluate the model and the algorithm, and then we experiment with real data from a surgical training system. The experimental results suggest that the proposed approach provides a promising solution to the real-world problem of motion skill evaluation from video.
Qiang Zhang 0021, Baoxin Li
CVPR2
2013 Detecting text in floor maps using Histogram of Oriented Gradients
abstract
Automatic detection of text labels in maps is essential for applications requiring automatic map understanding. This task is challenging due to factors such as varying font size and style, slanted words/phrases, and interfering graphics that are similar to text. This paper presents an approach for text detection in indoor floor maps. We exploit the difference in spatial frequency of edge orientations between text and non-text regions through Histogram of Oriented Gradients (HOG) features, and design a gradient-filtered Support Vector Machine (SVM) classifier based on such features. Special care was taken in conditioning the data for proper training of the classifier. The proposed approach was evaluated on a data set that had been collected and manually labeled. Experimental results show that the proposed method attained improved performance, outperforming a couple of reference methods/systems.
Hima Bindu Maguluri, Qiongjie Tian, Baoxin Li
ICASSP3
2013 Supporting navigation of outdoor shopping complexes for visuallyimpaired users through multi-modal data fusion
abstract
Outdoor shopping complexes (OSC) are extremely difficult for people with visual impairment to navigate. Existing GPS devices are mostly designed for roadside navigation and seldom transition well into an OSC-like setting. We report our study on the challenges faced by a blind person in navigating OSC through developing a new mobile application named iExplore. We first report an exploratory study aiming at deriving specific design principles for building this system by learning the unique challenges of the problem. Then we present a methodology that can be used to derive the necessary information for the development of iExplore, followed by experimental validation of the technology by a group of visually impaired users in a local outdoor shopping center. User feedback and other performance metrics collected from the experiments suggest that iExplore, while at its very initial phase, has the potential of filling a practical gap in existing assistive technologies for the visually impaired.
Devi Archana Paladugu, Parag S. Chandakkar, Peng Zhang 0017, Baoxin Li
ICME4
2013 Towards Predicting the Best Answers in Community-based Question-Answering Services
Qiongjie Tian, Peng Zhang 0017, Baoxin Li
ICWSM3
2012 Development of a Video-based System for Surgical Skill Training and Assessment
Gazi Islam, Kanav Kahol, Baoxin Li
AMIA3
2012 Automated description generation for indoor floor maps
abstract
People with visual impairment generally suffer from diminished freedom in navigating an environment. A practical need is to navigate through unfamiliar indoor environments such as school buildings, hotels, etc., for which commonly-used existing tools like canes, seeing-eye dogs and GPS devices cannot provide adequate support. We demonstrate a prototype system that aims at addressing this practical need. The input to the system is the name of the building/establishment supplied by a user, which is used by a web crawler to determine the availability of a floor map on the corresponding website. If available, the map is downloaded and used by the proposed system to generate a verbal description giving an overview of the locations of key landmarks inside the map with respect to one another. Our preliminary survey and experiments indicate that this is a promising direction to pursue in supporting indoor navigation for the visually impaired.
Devi Archana Paladugu, Hima Bindu Maguluri, Qiongjie Tian, Baoxin Li
ASSETS4
2012 Active learning for tag recommendation utilizing on-line photos lacking tags
abstract
Recommending text tags for on-line photos is useful for Internet photo services. Typical solutions to this problem require analysis of the correlation among different attributes of the photos, including the correlation between the textual features and visual features computed from a photo. However, most on-line photos have very few tags or even no tags, and thus they contribute little or none to the analysis of tag-photo correlation, which is a key component in those schemes that rely on such analysis for tag recommendation. To address this practical challenge, we propose an active learning method for incorporating photos with no or few tags so as to enhance the correlation analysis for improved performance in tag recommendation. We demonstrate the effectiveness of the proposed approach using a dataset of more than 33,000 photos collected from Flickr.
Yajun Gao, Baoxin Li
ICIP2
2012 Mining discriminative components with low-rank and sparsity constraints for face recognition
abstract
This paper introduces a novel image decomposition approach for an ensemble of correlated images, using low-rank and sparsity constraints. Each image is decomposed as a combination of three components: one common component, one condition component, which is assumed to be a low-rank matrix, and a sparse residual. For a set of face images of Nsubjects, the decomposition finds N common components, one for each subject, K low-rank components, each capturing a different global condition of the set (e.g., different illumination conditions), and a sparse residual for each input image. Through this decomposition, the proposed approach recovers a clean face image (the common component) for each subject and discovers the conditions (the condition components and the sparse residuals) of the images in the set. The set of N+K images containing only the common and the low-rank components form a compact and discriminative representation for the original images. We design a classifier using only these N+K images. Experiments on commonly-used face data sets demonstrate the effectiveness of the approach for face recognition through comparing with the leading state-of-the-art in the literature. The experiments further show good accuracy in classifying the condition of an input image, suggesting that the components from the proposed decomposition indeed capture physically meaningful features of the input.
Qiang Zhang 0021, Baoxin Li
KDD2
2012 Fast and independent access to map directions for people who are blind
abstract
This article presents an automatic approach, complete with a prototype system, to supporting fast and independent access to online maps for local navigation by people with visual impairment. With user-inputted start and end addresses from a keyboard, the approach first queries MapQuest (http://www.mapquest.com) for obtaining the walking directions and the corresponding map image. Then, it automatically converts the obtained information in a form that can be reproduced immediately through a tactile printer, and subsequently generates an SVG (Scalable Vector Graphics) file, which associates textual descriptions of the directions with a recreated tactile map. The tactile hard copy can be placed on a touchpad which is connected to a computer. With the generated SVG file opened in the computer, a user can explore the tactile map by hands, receiving instant audio feedback of the directions by pressing certain regions with special tactile patterns. This approach supports instant queries of walking directions without requiring tedious manual conversion by a sighted professional. The audio-tactile patterns, the adaptive representation scheme and the blind-friendly user interface are specifically designed for the visually-impaired users. Results from experimental evaluation based on a group of users with visual impairment suggest that the proposed approach is effective for providing blind computer users with independent access to geographic directions.
Zheshen Wang, Baoxin Li
Interact. Comput.3
2012 Understanding Compressive Sensing and Sparse Representation-Based Super-Resolution
abstract
Recently, compressive sensing (CS) has emerged as a powerful tool for solving a class of inverse/underdetermined problems in computer vision and image processing. In this paper, we investigate the application of CS paradigms on single image super-resolution (SR) problems that are considered to be the most challenging in this class. In light of recent promising results, we propose novel tools for analyzing sparse representation-based inverse problems using redundant dictionary basis. Further, we provide novel results establishing tighter correspondence between SR and CS. As such, we gain insights into questions concerning regularizing the solution to the underdetermined problem, such as follows. 1) Is sparsity prior alone sufficient? 2) What is a good dictionary? 3) What is the practical implication of noncompliance with theoretical CS hypothesis? Unlike in other underdetermined problems that assume random down-projections, the low-resolution image formation model employed in CS-based SR is a deterministic down-projection that may not necessarily satisfy some critical assumptions of CS. We further investigate the impact of such projections in concern to the above questions.
Naveen Kulkarni, Pradeep Nagesh, Rahul Gowda, Baoxin Li
IEEE Trans. Circuits Syst. Video Technol.4
2011 Discriminative affine sparse codes for image classification
abstract
Images in general are captured under a diverse set of conditions. An image of the same object can be captured with varied poses, illuminations, scales, backgrounds and probably different camera parameters. The task of image classification then lies in forming features of the input images in a representational space where classifiers can be better supported in spite of the above variations. Existing methods have mostly focused on obtaining features which are invariant to scale and translation, and thus they generally suffer from performance degradation on datasets which consist of images with varied poses or camera orientations. In this paper we present a new framework for image classification, which is built upon a novel way of feature extraction that generates largely affine-invariant features called affine sparse codes. This is achieved through learning a compact dictionary of features from affine-transformed input images. Analysis and experiments indicate that this novel feature is highly discriminative in addition to being largely affine-invariant. A classifier using AdaBoost is then designed using the affine sparse codes as the input. Extensive experiments with standard databases demonstrate that the proposed approach can obtain the state-of-the-art results, outperforming existing leading approaches in the literature.
Naveen Kulkarni, Baoxin Li
CVPR2
2011 TactileFace: a system for enabling access to face photos by visually-impaired people
abstract
Face photos/Portraits play an important role in people's social and emotional life. Unfortunately, this type of media is inaccessible to people with visual impairment. We propose a novel, prototypical system that was designed to demonstrate the feasibility of bridging this accessibility gap through automatic and realtime conversion of face images into their tactile counterparts. Unlike conventional and existing tactile conversion efforts, which are largely done by human specialists, the proposed system serves to provide an intelligent interface between blind computer users and this important type of media, human face images.
Zheshen Wang, Jesus Yuriar, Baoxin Li
IUI4
2011 Extracting key frames from consumer videos using bi-layer group sparsity
abstract
Compared to well-edited videos with predefined structures (e.g., news or sports videos), extracting key frames from unconstrained consumer videos remains a much more challenging problem due to their extremely diverse contents (no pre-imposed structure) and uncontrolled video quality (e.g., due to poor lighting or camera shake). In order to exploit spatio-temporal correlation present in the video for key frame extraction, we propose a bi-layer group sparse representation in which the input video frames are first segmented into homogeneous patches and group sparsity is imposed at two levels simultaneously: (i) patch-to-frame, and (ii) frame-to-sequence. The grouped sparse coefficients are further combined with frame quality scores to generate key frames. Extensive experiments are performed on videos from actual end users. Results obtained by the proposed approach compare favorably with existing methods to confirm its effectiveness.
Zheshen Wang, Mrityunjay Kumar, Jiebo Luo 0001, Baoxin Li
ACM Multimedia4
2011 Multifactor feature extraction for human movement recognition
Bo Peng 0005, Gang Qian, Yunqian Ma, Baoxin Li
Comput. Vis. Image Underst.4
2011 Rapid modeling of cones and cylinders from a single calibrated image using minimum 2D control points
Jin Zhou 0005, Baoxin Li
Mach. Vis. Appl.2
2010 YouTubeCat: Learning to categorize wild web videos
abstract
Automatic categorization of videos in a Web-scale unconstrained collection such as YouTube is a challenging task. A key issue is how to build an effective training set in the presence of missing, sparse or noisy labels. We propose to achieve this by first manually creating a small labeled set and then extending it using additional sources such as related videos, searched videos, and text-based webpages. The data from such disparate sources has different properties and labeling quality, and thus fusing them in a coherent fashion is another practical challenge. We propose a fusion framework in which each data source is first combined with the manually-labeled set independently. Then, using the hierarchical taxonomy of the categories, a Conditional Random Field (CRF) based fusion strategy is designed. Based on the final fused classifier, category labels are predicted for the new videos. Extensive experiments on about 80K videos from 29 most frequent categories in YouTube show the effectiveness of the proposed method for categorizing large-scale wild Web videos.
Zheshen Wang, Ming Zhao 0003, Yang Song 0009, Sanjiv Kumar, Baoxin Li
CVPR5
2010 Discriminative K-SVD for dictionary learning in face recognition
abstract
In a sparse-representation-based face recognition scheme, the desired dictionary should have good representational power (i.e., being able to span the subspace of all faces) while supporting optimal discrimination of the classes (i.e., different human subjects). We propose a method to learn an over-complete dictionary that attempts to simultaneously achieve the above two goals. The proposed method, discriminative K-SVD (D-KSVD), is based on extending the K-SVD algorithm by incorporating the classification error into the objective function, thus allowing the performance of a linear classifier and the representational power of the dictionary being considered at the same time by the same optimization procedure. The D-KSVD algorithm finds the dictionary and solves for the classifier using a procedure derived from the K-SVD algorithm, which has proven efficiency and performance. This is in contrast to most existing work that relies on iteratively solving sub-problems with the hope of achieving the global optimal through iterative approximation. We evaluate the proposed method using two commonly-used face databases, the Extended YaleB database and the AR database, with detailed comparison to 3 alternative approaches, including the leading state-of-the-art in the literature. The experiments show that the proposed method outperforms these competing methods in most of the cases. Further, using Fisher criterion and dictionary incoherence, we also show that the learned dictionary and the corresponding classifier are indeed better-posed to support sparse-representation-based recognition.
Qiang Zhang 0021, Baoxin Li
CVPR2
2010 Fast GPU implementation of large scale dictionary and sparse representation based vision problems
abstract
Recently, Computer Vision problems like Face Recognition and Super-Resolution solved using sparse representation based methods with large dictionaries have shown state-of-the-art results. However such methods are computationally prohibitive for typical CPUs, especially for a large dictionary size. We present fast implementation of these methods by exploiting the massively parallel processing capabilities of a GPU within a CUDA framework, owing to its easy off-the-shelf availability and programmer friendliness. We provide details of system level design, memory management and implementation strategies. Further, we integrate the solution to the preferred scientific computational platform - MATLAB.
Pradeep Nagesh, Rahul Gowda, Baoxin Li
ICASSP3
2010 Tensor completion for on-board compression of hyperspectral images
abstract
We present a new image compression scheme for hyperspectral images based on the newly-emerged matrix/tensor completion theory. Unlike typical transform-coding based methods, the proposed approach does not require any transform to be performed by the imaging sensor when doing on-board compression. Only a small set of pixels on a sparse set of locations on the imaging sensor needs to be captured and transmitted for each image. The decoder side relies on matrix/tensor completion for reconstructing the original images. Hence the scheme can drastically reduce the computation and bandwidth requirements on the on-board imaging sensors. Experiments show that the proposed method is able to obtain compression performance close to JPEG2000 while enjoying the afore-mentioned unique benefits.
Baoxin Li
ICIP2
2010 Joint Sparsity Model with Matrix Completion for an ensemble of face images
abstract
An ensemble of correlated signals are often encountered in many applications of image processing, such as a set of face images of the same subject. In this paper, we propose a new model, called Joint Sparsity Model with Matrix Completion (JSM-MC), which extracts a common component, an innovation component, and a low-rank component from an ensemble of face images. These components have their respective physical significance in terms of representing different types of information in the original ensemble, hence facilitating an analysis task such as recognition. An algorithm is proposed under the model to solve for the components, based on Block Coordinate Descent and Singular Value Thresholding. Experimental results show that the proposed method has unique advantages over existing methods in dealing with challenging face images with extreme illumination conditions or occlusions.
Qiang Zhang 0021, Baoxin Li
ICIP2
2010 On implementing motion-based Region of Interest detection on multi-core CELL
Avin Kumar Kannur, Baoxin Li
Comput. Vis. Image Underst.2
2010 A Bayesian Approach to Automated Creation of Tactile Facial Images
abstract
Portrait photos (facial images) play important social and emotional roles in our life. This type of visual media is unfortunately inaccessible by users with visual impairment. This paper proposes a systematic approach for automatically converting human facial images into a tactile form that can be printed on a tactile printer and explored by a user who is blind. We propose a deformable Bayesian Active Shape Model (BASM), which integrates anthropometric priors with shape and appearance information learnt from a face dataset. We design an inference algorithm under this model for processing new face images to create an input-adaptive face sketch. Further, the model is enhanced by input-specific details through semantic-aware processing. We report experiments on evaluating the accuracy of face alignment using the proposed method, with comparison with other state-of-the-art results. Furthermore, subjective evaluations of the produced tactile face images were performed by 17 persons including six visually-impaired users, confirming the effectiveness of the proposed approach in conveying via haptics vital visual information in a face image.
Zheshen Wang, Baoxin Li
IEEE Trans. Multim.2
2009 Instant tactile-audio map: enabling access to digital maps for people with visual impairment
abstract
In this paper, we propose an automatic approach, complete with a prototype system, for supporting instant access to maps for local navigation by people with visual impairment. The approach first detects and segments texts from a map image and recreates the remaining graphical parts in a tactile form which can be reproduced immediately through a tactile printer. Then, it generates an SVG (Scalable Vector Graphics) file, which integrates both text and graphical information. The tactile hardcopy and the SVG file together are used to provide a user with interactive access to the map image through a touchpad, resulting in a tactile-audio representation of the original input image. This supports real-time access to the map without tedious conversion by a sighted professional. Evaluations with six users who are blind show that the created tactile-audio maps from our prototype system convey the most important map information and are deemed as potentially useful for local navigation.
Zheshen Wang, Baoxin Li, Terri Hedgpeth, Teresa Haven
ASSETS2
2009 A compressive sensing approach for expression-invariant face recognition
abstract
We propose a novel technique based on compressive sensing for expression-invariant face recognition. We view the different images of the same subject as an ensemble of intercorrelated signals and assume that changes due to variation in expressions are sparse with respect to the whole image. We exploit this sparsity using distributed compressive sensing theory, which enables us to grossly represent the training images of a given subject by only two feature images: one that captures the holistic (common) features of the face, and the other that captures the different expressions in all training samples. We show that a new test image of a subject can be fairly well approximated using only the two feature images from the same subject. Hence we can drastically reduce the storage space and operational dimensionality by keeping only these two feature images or their random measurements. Based on this, we design an efficient expression-invariant classifier. Furthermore, we show that substantially low dimensional versions of the training features, such as (i) ones extracted from critically-downsampled training images, or (ii) low-dimensional random projection of original feature images, still have sufficient information for good classification. Extensive experiments with publically-available databases show that, on average, our approach performs better than the state-of-the-art despite using only such super-compact feature representation.
Pradeep Nagesh, Baoxin Li
CVPR2
2009 Power-aware content-adaptive H.264 video encoding
abstract
H.264 is a computationally intensive video codec striving for achieving the best quality for the compressed video. The computational complexity poses as a challenge for power-constrained applications. We present a system level complexity reduction for H.264 video encoding by allocating resources based on computational complexity and quality trade-off. We develop a framework which allocates the computational power of the encoder adaptive to video contents and also scales with the available battery power using a ROI classification method. Analysis is done to profile the key modules of the encoder which can be power-optimized while allocating resources. The results of the encoder module analysis are combined with the motion content analysis to obtain a power efficient encoder parameter set which reduces the computations and hence the power consumed. Our simulation results on the JM H.264 framework confirm our hypothesis and computational savings of more than 50% with quality degradation less than 1% is achieved thereby extending it's feasibility for battery powered wireless devices.
Avin Kumar Kannur, Baoxin Li
ICASSP2
2009 Compressive imaging of color images
abstract
In this paper, we propose a novel compressive imaging framework for color images. We first introduce an imaging architecture based on combining the existing single-pixel Compressive Sensing (CS) camera with a Bayer color filter, thereby enabling acquisition of compressive color measurements. Then we propose a novel CS reconstruction algorithm that employs joint sparsity models in simultaneously recovering the R, G, B channels from the compressive measurements. Experiments simulating the imaging and reconstruction procedures demonstrate the feasibility of the proposed idea and the superior quality in reconstruction.
Pradeep Nagesh, Baoxin Li
ICASSP2
2009 Human Activity Encoding and Recognition Using Low-level Visual Features
Zheshen Wang, Baoxin Li
IJCAI2
2009 Adapting quantization offset in multiple description coding for error resilient video transmission
Viswesh Parameswaran, Avin Kumar Kannur, Baoxin Li
J. Vis. Commun. Image Represent.3
2009 Virtual View Specification and Synthesis for Free Viewpoint Television
abstract
Free viewpoint television (FTV) is a new concept that aims at giving viewers the flexibility to select a novel viewpoint by employing multiple video streams as the input. Current proposed solutions for FTV include those based on ray-space resampling which demand at least dozens of cameras and large storage and transmission resources for those video streams. Image-based rendering (IBR) methods that rely on dense depth map estimation also face practical difficulties since accurate depth map estimation remains a challenging problem. This paper proposes a framework for FTV based on IBR that relieves the need for an accurate depth map by introducing a hybrid virtual view synthesis method. The framework also includes an intuitive method for virtual view specification in uncalibrated views. We present both simulation and real data experiments to validate the proposed framework and the component algorithms.
Wenfeng Li 0004, Jin Zhou 0005, Baoxin Li, M. Ibrahim Sezan
IEEE Trans. Circuits Syst. Video Technol.3
2009 Exploiting Motion Correlations in 3-D Articulated Human Motion Tracking
abstract
In 3-D articulated human motion tracking, the curse of dimensionality renders commonly-used particle-filter-based approaches inefficient. Also, noisy image measurements and imperfect feature extraction call for strong motion prior. We propose to learn the correlation between the right-side and the left-side human motion using partial least square (PLS) regression. The correlation effectively constrains the sampling of the proposal distribution to portions of the parameter space that correspond to plausible human motions. The learned correlation is then used as motion prior in designing a Rao-Blackwellized particle filter algorithm, RBPF-PLS, which estimates only one group of state variables using the Monte Carlo method, leaving the other group being exactly computed through an analytical filter that utilizes the learned motion correlation. We quantitatively assessed the accuracy of the proposed algorithm with challenging HumanEva-I/II data set. Experiments with comparison with both the annealed particle filter and the standard particle filter show that the proposed method achieves lower estimation error in processing challenging real-world data of 3-D human motion. In particular, the experiments demonstrate that the learned motion correlation model generalizes well to motions outside of the training set and is insensitive to the choice of the training subjects, suggesting the potential wide applicability of the method.
Baoxin Li
IEEE Trans. Image Process.2
2008 Joint Conditional Random Field of multiple views with online learning for image-based rendering
abstract
There are many applications, such as image-based rendering, where multiple views of a scene are considered simultaneously for improved analysis through employing strong correlation among the set of pixels corresponding to the same physical scene point. While being a useful tool for modeling pixel interactions, Markov random field (MRF) models encounter challenges in such cases since they assume strong independence of the observed data for tractability, rendering it difficult to take advantage of having multiple correlated views. In this paper we propose joint conditional random field (CRF) for multiple views in the context of virtual view synthesis in image-based rendering. The model is enabled by the adoption of steerable spatial filters for capturing not only the pixel dependence in a single image but also their correlations among multiple views. Furthermore, a novel on-line learning scheme is proposed for the CRF model, which learns the CRF parameters from the same input data for synthesizing virtual views. This effectively makes the model adaptive to the input and thus optimal results can be expected. Experiments are designed to validate the proposed approach and its effectiveness.
Wenfeng Li 0004, Baoxin Li
CVPR2
2008 Bayesian tactile face
abstract
Computer users with visual impairment cannot access the rich graphical contents in print or digital media unless relying on visual-to-tactile conversion, which is done primarily by human specialists. Automated approaches to this conversion are an emerging research field, in which currently only simple graphics such as diagrams are handled. This paper proposes a systematic method for automatically converting a human portrait image into its tactile form. We model the face based on deformable active shape model (ASM) (Cootes et al., 1995), which is enriched by local appearance models in terms of gradient profiles along the shape. The generic face model including the appearance components is learnt from a set of training face images. Given a new portrait image, the prior model is updated through Bayesian inference. To facilitate the incorporation of a pose-dependent appearance model, we propose a statistical sampling scheme for the inference task. Furthermore, to compensate for the simplicity of the face model, edge segments of a given image are used to enrich the basic face model in generating the final tactile printout. Experiments are designed to evaluate the performance of the proposed method.
Zheshen Wang, Baoxin Li
CVPR3
2008 Multi-target tracking based on KLD mixture particle filter with radial basis function support
abstract
The two major difficulties associated with practical real time multi-target tracking are accuracy and speed. A new technique is proposed for multi-target tracking based on multi-modal particle filter with fast tracking capability and improved accuracy. The speed in tracking is achieved by a KLD sampling stage while the accuracy is improved by an additional stage that uses radial basis functions (RBF) for interpolating the sparse particles. Test results of the proposed multi-target tracking approach on both synthetic and real video data demonstrate the improved performance.
Jayanth Madapura, Baoxin Li
ICASSP2
2008 A two-stage approach to saliency detection in images
abstract
Researches in psychology, perception and related fields show that there may be a two-stage process involved in human vision. In this paper, we propose an approach by following a two-stage framework for saliency detection. In the first stage, we extend an existing spectrum residual model for better locating visual pop-outs, while in the second stage we make use of coherence based propagation for further refinement of the results from the first step. For evaluation of the proposed approach, 300 images with diverse contents were manually and accurately labeled. Experiments show that our approach achieves much better performance than that from the existing state-of-art.
Zheshen Wang, Baoxin Li
ICASSP2
2008 An enhanced rate control scheme with motion assisted slice grouping for low bit rate coding in H.264
abstract
This paper presents an enhanced rate control scheme for H.264 using motion detection and motion analysis. PSNR based measure is used for determining motion complexity and detecting scene changes. Motion Vectors and Rate Distortion measures are further used to classify regions with motion inside a frame and Macro-blocks are then grouped into slices using the Flexible Macro-block Ordering feature of H.264. We present a filtered slice grouping technique to maintain a stable region of motion. The proposed method provides smoother quality videos with better average PSNR. The bit rates are within bound and fewer frames are skipped in sequences with scene changes at low bit rates.
Avin Kumar Kannur, Baoxin Li
ICIP2
2008 Virtual view synthesis with heuristic spatial motion
abstract
Probabilistic methods have been used in image-based rendering for solving the virtual view synthesis problem with Bayesian inference. To work well, the inference process requires the input views to be consistent to yield reasonable result, which in turn constrains the cameras to be very close to each other. Many approaches to relieving such constraint focus on the prior model. In this paper, we present a method which treats the virtual view as the outcome of a spatial motion from one real view. A sequence of images is generated heuristically to preserve textures with the aid of steerable filters. Interim results are further refined with texture-based Markov random field prior model. Experiments show that the synthesized view can have satisfactory image quality with only a few input images from wide baseline cameras.
Wenfeng Li 0004, Baoxin Li
ICIP2
2007 Evaluating Multi-class Multiple-Instance Learning for Image Categorization
Baoxin Li
ACCV (2)2
2007 Exploiting Vertical Lines in Vision-Based Navigation for Mobile Robot Platforms
abstract
Vertical line-like structures are abundant in man-made environments, such as the vertical edges of buildings in an outdoor environment and door frames and vertical edges of furniture in an indoors environment. These ubiquitous line or line-like structures may provide useful information for mobile robot navigation. In this paper, we propose an approach to estimating the camera orientation based on detected vertical lines. Further we use the estimated camera orientation in one key problem of mobile robot navigation: ground/obstacle detection. Our method is theoretically well-grounded and yet simple and practical in implementation.
Jin Zhou 0005, Baoxin Li
ICASSP (1)2
2007 Learning Motion Correlation for Tracking Articulated Human Body with a Rao-Blackwellised Particle Filter
abstract
Inference in 3D articulated human body tracking is challenging due to the high dimensionality and nonlinearity of the parameter-space. We propose a particle filter with Rao-Blackwellisation which marginalizes part of the state variables by exploiting the correlation between the right-side and the left-side joint Euler angles. The correlation is naturally induced by the symmetric and repetitive patterns in specific human activities. A novel algorithm is proposed to learn the correlation from the training data using partial least square regression. The learned correlation is then used as motion prior in designing the Rao-Blackwellised particle filter, which estimates only one group of state variables using the Monte Carlo method, leaving the other group being exactly computed through an analytical filter that utilizes the learned motion correlation. We evaluate the effectiveness of the motion correlation for 3D articulated human body tracking. The accuracy of the proposed 3D tracker is quantitatively assessed based on the distance between the true and the estimated marker positions. Extensive experiments with multi-camera walking sequences from the HumanEva-I/II data set show that (i) the proposed tracker achieves significantly lower estimation error than both the annealed particle filter and the standard particle filter; and (ii) the learned motion correlation generalizes well to motion performed by subjects other than the training subject.
Baoxin Li
ICCV2
2007 MAP Estimation of Epipolar Geometry by EM Algorithm and Local Diffusion
abstract
Finding epipolar geometry for two images is a fundamental problem in computer vision. While this typically relies on feature point correspondence, the epipolar constraint can also be used for improving the accuracy of correspondence. We propose a probabilistic framework for estimating the epiploar geometry, in which the geometry and the feature correspondence are estimated iteratively at the same time. Using the EM algorithm to maximize a posteriori, our approach updates feature correspondence with estimated epipolar geometry. The correspondence is further improved with local diffusion on a prior Markov Random Field model. In turn, more accurate epipolar geometry is recovered. Experiments show this approach produces more accurate fundamental matrix compared with typical methods and can handle some challenging situations such as view rotation and scale changes.
Wenfeng Li 0004, Baoxin Li
ICIP (5)2
2007 3D Articulated Human Body Tracking using KLD-Annealed Rao-Blackwellised Particle Filter
abstract
The difficulties introduced by large degrees of freedom are still a challenge in articulated human body tracking. In this paper, an efficient tracker is proposed based on the integration of a set of statistical techniques including KLD sampling, Rao-Blackwellisation, and Particle filtering. This results in a KLD-Annealed Particle filter with Rao-Blackwellisation, which can address the key issues in 3D human tracking, such as accuracy, stability, and speed simultaneously. Both synthetic and real data were used in our experiments to demonstrate the improved performance of the proposed tracker.
Jayanth Madapura, Baoxin Li
ICME2
2007 Video Inpainting for Largely Occluded Moving Human
abstract
In this paper, a video inpainting approach is proposed, which targets at repairing a video containing moving humans that are largely or completely occluded or missing for some of the frames. The proposed approach first categorizes typically periodic human motion in a video into a set of temporal states (called motion states), and then estimates the motion states for the frames with missing humans so as to repair the missing parts using other undamaged frames with the same motion states. This deviates from common approaches that directly repair the pixels of the damaged parts. Experiments demonstrate that the proposed method can well repair the damaged video sequences without introducing strong artifacts that exist in many existing techniques.
Haomian Wang, Houqiang Li, Baoxin Li
ICME3
2007 Adaptive Rao-Blackwellized Particle Filter and Its Evaluation for Tracking in Surveillance
abstract
Particle filters can become quite inefficient when being applied to a high-dimensional state space since a prohibitively large number of samples may be required to approximate the underlying density functions with desired accuracy. In this paper, by proposing an adaptive Rao-Blackwellized particle filter for tracking in surveillance, we show how to exploit the analytical relationship among state variables to improve the efficiency and accuracy of a regular particle filter. Essentially, the distributions of the linear variables are updated analytically using a Kalman filter which is associated with each particle in a particle filtering framework. Experiments and detailed performance analysis using both simulated data and real video sequences reveal that the proposed method results in more accurate tracking than a regular particle filter.
Baoxin Li
IEEE Trans. Image Process.2
2006 Robust Ground Plane Detection with Normalized Homography in Monocular Sequences from a Robot Platform
abstract
We present a homography-based approach to detect the ground plane from monocular sequences captured by a robot platform. By assuming that the camera is fixed on the robot platform and can at most rotate horizontally, we derive the constraints that the homograph of the ground plane must satisfy and then use these constraints to design algorithms for detecting the ground plane. Due to the reduced degree of freedom, the resultant algorithm is not only more efficient and robust, but also able to avoid false detection due to virtual planes. We present experiments with real data from a robot platform to validate the proposed approaches.
Jin Zhou 0005, Baoxin Li
ICIP2
2006 Homography-based Ground Detection for a Mobile Robot Platform using a Single Camera
abstract
This paper presents a practical approach to ground detection in mobile robot applications based on a monocular sequence captured by an on-board camera. We formulate the problem of ground plane detection as one of estimating the dominant homography between two frames taken from the sequence, and then design an efficient algorithm for the estimation. In particular, we analyze a problem inherent to any homography-based approach to the given task, and show how the proposed approach can address this problem to a large degree. Although not explicitly discussed, the proposed method can be used to guide the maneuver of the robot, as the detected ground plane can in turn be used in obstacle avoidance
Jin Zhou 0005, Baoxin Li
ICRA2
2005 Head Tracking Using Particle Filter with Intensity Gradient and Color Histogram
abstract
This paper presents a method for tracking human head using a particle filter to naturally integrate two complementary cues: intensity gradient and color histogram. The shape of the head is modeled as an ellipse, along which an intensity gradient is estimated, while the interior appearance is modeled using a color histogram. These two cues play complementary roles in tracking a human head with free rotation on a cluttered background. To evaluate the tracker performance, we test with both synthetic image sequences and real sequences. Experiments show that the tracker is robust to 360-degree rotation of the head on a cluttered background
Baoxin Li
ICME2
2005 Non-linear image enhancement for digital TV applications using Gabor filters
abstract
We propose a non-linear image enhancement method based on Gabor filters, which allows selective enhancement based on the contrast sensitivity function of the human visual system. We also propose an evaluation method for measuring the performance of the algorithm and for comparing it with existing approaches. The selective enhancement of the proposed approach is especially suitable for digital television applications to improve the perceived visual quality of the images when the source image contains less satisfactory amount of high frequencies due to various reasons, including interpolation that is used to convert standard definition sources into high-definition images.
Baoxin Li
ICME2
2005 Automatic Generation of Pencil-Sketch Like Drawings from Personal Photos
abstract
In this paper, we present an algorithm for automatically generating pencil-sketch like drawings from personal photos. On top of the core step of gradient computation, some proper transformations are introduced to achieve the best visual effects. It is found that the popular Sobel operator and Laplacian operator are both not desirable for this task, and thus we propose a computationally simple algorithm for gradient estimation. Experimental results show that the proposed method can generate visually appealing pencil-sketch like images from personal photos.
Jin Zhou 0005, Baoxin Li
ICME2
2004 Automatic keystone correction for smart projectors with embedded camera
Baoxin Li, M. Ibrahim Sezan
ICIP1
2004 Bridging the semantic gap in sports video retrieval and summarization
Baoxin Li, James H. Errico, Hao Pan 0002, M. Ibrahim Sezan
J. Vis. Commun. Image Represent.1
2003 A general framework for sports video summarization with its application to soccer
abstract
We propose a general framework for indexing and summarizing sports broadcast programs and its specific application to soccer. The framework is based on a high-level model of sports broadcast video using the concept of an event, defined according to domain-specific knowledge for different types of sports. Thus it covers most sports including those that have the action-and-stop pattern (e.g., baseball) and those containing continuous actions (e.g., soccer). In particular, within this general framework, using soccer as an example, we propose a novel approach to automatic event detection, which is based on automatic analysis of the visual and aural signals in the media. An MPEG-7 compliant prototype browsing system has been implemented to demonstrate the results.
Baoxin Li, Hao Pan 0002, M. Ibrahim Sezan
ICASSP (3)1
2003 Semantic sports video analysis: approaches and new applications
abstract
In this paper, we show that event-based semantic analysis of sports video can be achieved through two major classes of methods: approaches using deterministic reasoning and approaches depending on probabilistic inference. Deterministic approaches depend on an explicit description of an event and an explicit representation of the event in terms of low-level aural-visual features that can be automatically detected. Probabilistic approaches are based on some proper state/transition models whose parameters are learnt through labeled training sequences. We first discuss the pros and cons of the two classes of approaches. Next, we present algorithms of these two types for a new application, automatic analysis of coaching video (which is different from sports broadcast programs that have been subject of most prior work).
Baoxin Li, Ibrahim Sezan
ICIP (1)1
2002 Bayesian methods for face recognition from video
abstract
Face recognition (FR) from video necessitates simultaneously solving two asks, recognition and tracking. To accommodate the video, a time series state space model is introduced in a Bayesian approach. Given this model, the goal reduces to estimating the posterior distribution of the state vector given the observations up to the present. The Sequential Importance Sampling (SIS) technique is invoked to generate a numerical solution to this model. However, the ultimate goal is to estimate the posterior distribution of the identity of humans for recognition purposes. Presented here are two methods to approximate the above distribution under different experimental scenarios.
Rama Chellappa, Shaohua Kevin Zhou, Baoxin Li
ICASSP3
2002 Automatic detection of replay segments in broadcast sports programs by detection of logos in scene transitions
abstract
In broadcast sports, replays provide viewers another look at interesting events. We propose an automatic algorithm for replay segment detection by detecting frames containing logos in the special scene transitions that sandwich replays. Detected replays are utilized in efficient navigation, indexing, and summarization of sports programs. The proposed algorithm first automatically determines the logo template from frames surrounding slow motion segments, where slow motion segments are automatically detected using [4]. Then, it locates all the similar frames in the video using the logo template. Finally the algorithm identifies the replay segments by grouping the detected logo frames and slow-motion segments. Our algorithm accurately detects replays, with or without slow motion.
Hao Pan 0002, Baoxin Li, M. Ibrahim Sezan
ICASSP2
2002 A generic approach to simultaneous tracking and verification in video
abstract
In this paper, a generic approach to simultaneous tracking and verification in video data is presented. The approach is based on posterior density estimation using sequential Monte Carlo methods. Visual tracking, which is in essence a temporal correspondence problem, is solved through probability density propagation, with the density being defined over a proper state space characterizing the object configuration. Verification is realized through hypothesis testing using the estimated posterior density. In its most basic form, verification can be performed as follows. Given a measurement vector Z and two hypotheses H1 and H0, we first estimate posterior probabilities P(H0/Z) and P(H1/Z), and then choose the one with the larger posterior probability as the true hypothesis. Several applications of the approach are illustrated by experiments devised to evaluate its performance. The idea is first tested on synthetic data, and then experiments with real video sequences are presented, illustrating vehicle tracking and verification, human (face) tracking and verification, facial feature tracking, and image sequence stabilization.
Baoxin Li, Rama Chellappa
IEEE Trans. Image Process.1
2001 Adaptive Video Background Replacement
abstract
A video background replacement algorithm is proposed, which is based on background subtraction with adaptive background modeling. Like one early work ([1]) in this direction, it does not require a blue screen but needs only a pre-recorded background scene image. The problem is formulated as one that detects statistical outliers with respect to the given background. A two-pass process, which refines initial segmentation based on the statistics on a pixel’s neighborhood, is adopted in order to suppress false positives in the background region while increasing detection rate for the foreground object. Experiments with real image sequences are presented, along with comparisons with some other existing methods, illustrating the advantages of the proposed algorithm. 1.
Baoxin Li, M. Ibrahim Sezan
ICME1
2001 Experimental Evaluation of FLIR ATR Approaches - A Comparative Study
Baoxin Li, Rama Chellappa, Qinfen Zheng, Sandor Z. Der, Nasser M. Nasrabadi, LipChen Alex Chan, Lin-Cheng Wang
Comput. Vis. Image Underst.1
2001 Model-based temporal object verification using video
abstract
An approach to model-based dynamic object verification and identification using video is proposed. From image sequences containing the moving object, we compute its motion trajectory. Then we estimate its three-dimensional (3-D) pose at each time step. Pose estimation is formulated as a search problem, with the search space constrained by the motion trajectory information of the moving object and assumptions about the scene structure. A generalized Hausdorff (1962) metric, which is more robust to noise and allows a confidence interpretation, is suggested for the matching procedure used for pose estimation as well as the identification and verification problem. The pose evolution curves are used to assist in the acceptance or rejection of an object hypothesis. The models are acquired from real image sequences of the objects. Edge maps are extracted and used for matching. Results are presented for both infrared and optical sequences containing moving objects involved in complex motions.
Baoxin Li, Rama Chellappa, Qinfen Zheng, Sandor Z. Der
IEEE Trans. Image Process.1
2000 Simultaneous Tracking and Verification via Sequential Posterior Estimation
abstract
An approach to simultaneous tracking and verification in video data is presented. The approach is based on posterior estimation using sequential Monte Carlo methods. Visual tracking, which is in essence a temporal correspondence problem, is solved through probability density propagation, with the density being defined over a proper state space characterizing the object configuration. Verification is realized through hypothesis testing using the estimated posterior density. In its most basic form, verification can be performed as follows. Given measurement Z and two hypothesis H/sub 1/ and H/sub 0/, we first estimate posterior probabilities P(H/sub 0/|Z) and P(H/sub 1/|Z); and choose the one with the larger posterior probability as the true hypothesis. Applications of the approach are illustrated with experiments devised to evaluated the performance. The idea is first tested on synthetic data, and then experiments with real video sequences are presented.
Baoxin Li, Rama Chellappa
CVPR1
2000 Gabor Attributes Tracking for Face Verification
abstract
A method based on sequential importance sampling is proposed for tracking facial features on a grid with Gabor attributes. The motion of facial feature points is modeled as a global 2-D affine transformation (accounting for head motion) plus a local deformation (accounting for residual motion due to inaccuracies in 2-D affine modeling and other factors such as facial expression). Motion of both types is estimated simultaneously by the tracker: global motion is tracked by importance sampling, and residual motion is handled by incorporating local deformation into the measurement likelihood in computing the weight of a sample. While it has other applications in facial analysis, the method is particularly applicable to face verification because of a novel parametrization.
Baoxin Li, Rama Chellappa
ICIP1
1999 Dynamic object identification and verification using video
abstract
We introduce the concepts of dynamic object identification and verification using video. A generalized Hausdorff metric, which is more robust to noise and allows a confidence interpretation, is suggested for the identification and verification problem. Parameters from sensor motion compensation procedure are incorporated into the search step such that the Hausdorff metric based matching can be achieved efficiently under more complex transformation groups. An algorithm is proposed for identification/verification based on edge map matching using the generalized Hausdorff metric. Experiments on infrared video sequences are provided.
Baoxin Li, Rama Chellappa, Qinfen Zheng, Sandor Z. Der
ICASSP1
1999 Building pattern classifiers using convolutional neural networks
abstract
Pattern classification is the core task of many applications such as image segmentation. This paper studies the possibility of building pattern classifiers for text/picture segmentation and text detection problems using convolutional neural networks (CNNs). By using CNNs, explicit feature extraction is avoided-the feature detectors are learned from the training data. More importantly, CNNs can directly operate on grey level images, making its application straightforward. Addressed are practical issues such as kernel size, convergence speed, etc. Experiments on Chinese text/picture segmentation and text detection are presented.
Bao-Qing Li, Baoxin Li
IJCNN2