Xiuwen Liu 0001

dblp:89/3077-1 · DBLP profile ↗
← Back
103ranked-venue papers
18as first author
16since 2021 · last 2026
0000-0002-9320-3872ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 12 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-authorHuman-computer interaction and ubiquitous computing · 7Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 2Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Adversarial attacks on large language models using regularized relaxation
Samuel Jacob Chacko, Sajib Biswas, Chashi Mahiul Islam, Fatema Tabassum Liza, Xiuwen Liu 0001
Inf. Sci.5
2025 Data-Driven Fairness Generalization for Deepfake Detection
Uzoamaka Ezeakunne, Chrisantus Eze, Xiuwen Liu 0001
ICAART (3)3
2025 Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
abstract
As Large Language Models (LLMs) are widely used, understanding them systematically is key to improving their safety and realizing their full potential. Although many models are aligned using techniques such as reinforcement learning from human feedback (RLHF), they are still vulnerable to jailbreaking attacks. Some of the existing adversarial attack methods search for discrete tokens that may jailbreak a target model while others try to optimize the continuous space represented by the tokens of the model’s vocabulary. While techniques based on the discrete space may prove to be inefficient, optimization of continuous token embeddings requires projections to produce discrete tokens, which might render them ineffective. To fully utilize the constraints and the structures of the space, we develop an intrinsic optimization technique using exponentiated gradient descent with the Bregman projection method to ensure that the optimized one-hot encoding always stays within the probability simplex. We prove the convergence of the technique and implement an efficient algorithm that is effective in jailbreaking several widely used LLMs. We demonstrate the efficacy of the proposed technique using five open-source LLMs on four openly available datasets. The results show that the technique achieves a higher success rate with great efficiency compared to three other state-of-the-art jailbreaking techniques. The source code for our implementation is available at: https://github.com/sbamit/Exponentiated-Gradient-Descent-LLM-Attack
Sajib Biswas, Mao Nishino, Samuel Jacob Chacko, Xiuwen Liu 0001
IJCNN4
2025 Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
abstract
Transformer-based models have dominated natural language processing and other areas in the last few years due to their superior (zero-shot) performance on benchmark datasets. However, these models are poorly understood due to their complexity and size. While probing-based methods are widely used to understand specific properties, the structures of the representation space are not systematically characterized; consequently, it is unclear how such models generalize and overgeneralize to new inputs beyond datasets. In this paper, based on our recently proposed gradient descent optimization method, we are able to explore the embedding space of a commonly used vision-language model. Using the Imagenette dataset, we show that while the model achieves over 99% zero-shot classification performance, it fails systematic evaluations completely. Using a linear approximation, we provide a framework to explain the striking differences. We have also obtained similar results using a different model to support that our results are applicable to other transformer models with continuous inputs. We also propose a robust way to detect the modified images.
Shaeke Salman, Montasir Shams, Xiuwen Liu 0001, Lingjiong Zhu
IJCNN3
2025 Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
abstract
Utilizing a shared embedding space, emerging multimodal models exhibit unprecedented zero-shot capabilities. However, the shared embedding space could lead to new vulnerabilities if different modalities can be misaligned. In this paper, we extend and utilize a recently developed effective gradient-based procedure that allows us to match the embedding of a given text by minimally modifying an image. Using the procedure, we show that we can align the embeddings of distinguishable texts to any image with unnoticeable changes in joint image-text models, revealing that semantically unrelated images can have embeddings of identical texts and at the same time visually indistinguishable images can be matched to the embeddings of very different texts. Our technique achieves 100% success rate when it is applied to text datasets and images from multiple sources. Without overcoming the vulnerability, multimodal models cannot robustly align inputs from different modalities in a semantically meaningful way. Warning: the text data used in this paper are toxic in nature and may be offensive to some readers.
Shaeke Salman, Montasir Shams, Mao Nishino, Xiuwen Liu 0001
IJCNN4
2025 Nonlinear Correct and Smooth for Graph-Based Semi-Supervised Learning
abstract
Graph-based semi-supervised learning (GSSL) has achieved significant success across various applications by leveraging the graph structure and labeled samples for classification tasks. In the field of GSSL, Label Propagation (LP) and Graph Neural Networks (GNNs) are two complementary methods, in which LP iteratively propagates and updates node labels through connected nodes, whereas GNNs aggregate node features by incorporating information from their neighbors. Recently, the complementary nature of LP and GNNs has been utilized to improve performance through the combination of two approaches. However, the utilization of higher-order graph structures within these combined approaches, such as triangles, is still under-explored. Therefore, to advance understanding in this ongoing research, we first model GSSL as a two-step feature-label process. Then, we introduce Nonlinear Correct and Smooth (NLCS) in the post-processing step, a combined method that incorporates nonlinearity and higher-order structures into the residual propagation to handle intricate node relationships effectively. We propose a new synthetic graph generator to deepen the analysis and broaden the experimentation, providing insights into the mechanisms that enable NLCS to handle intricate node relationships effectively. Our systematic evaluations across six synthetic graphs show that NLCS outperforms base predictions by an average of 12.44% and the existing state-of-the-art post-processing method by 8.04%. Furthermore, on six commonly used real-world datasets, NLCS demonstrates a 10.9% improvement over six base prediction models and a 1.6% over the state-of-the-art post-processing method. Our comparisons and analyses reveal that NLCS substantially enhances the prediction accuracy of nodes within complex graph structures by effectively utilizing higher-order structures of graphs.
Yuanhang Shao, Xiuwen Liu 0001
ACM Trans. Knowl. Discov. Data2
2024 FaCTQA:Detecting and Localizing Factual Errors in Generated Summaries Through Question and Answering from Heterogeneous Models
abstract
With the advancements of pre-trained large language models, it has become easier to generate fluent abstractive summaries automatically. However, these generated summaries often suffer from factual inconsistencies, or incorrect information, known as hallucinations. Even though there are several methods on identifying hallucinations and hallucinated quantities in machine-generated summaries, detecting hallucination accurately and precisely is still challenging. In this paper, we propose a new pipeline, Factual Check Through Question-Answering (FaCTQA) to detect hallucinations and localize hallucinated quantities for automatically evaluating machine-generated summaries. State of the art (sota) language models fine-tuned with benchmark datasets have better performance on the application problem. Hence, we use question-answering techniques with heterogeneous fine-tuned models to interrogate a summary and its corresponding text document to verify its factual consistencies and extract the probable causes of hallucinations. We have also implemented a variant of the pipeline, FaCTQA* where all the components are replaced by GPT 4. Our experimental results show that FaCTQA outperforms the previous sota models for detecting and localizing hallucination. FaCTQA achieves 90.3% accuracy on the benchmark data XSum and 82.62% accuracy on CNN/DM. FaCTQA has 12.5% better accuracy than the sota models and 4% better performance that FaCTQA*.
Trina Dutta, Xiuwen Liu 0001
IJCNN2
2024 Nonlinear Correct and Smooth for Semi-Supervised Learning
abstract
Graph-based semi-supervised learning (GSSL) has been used successfully in various applications, with existing methods leveraging the graph structure and labeled samples for classification. Label Propagation (LP) and Graph Neural Networks (GNNs) both iteratively pass messages on graphs, where LP propagates and updates node labels across the graph and GNN aggregates node features from their neighboring nodes. Recently, combining LP and GNN has led to improvements in performance, yet the joint utilization of labels and features in higher-order structures of graphs, such as triangles, remains unexplored. Therefore, we introduce Nonlinear Correct and Smooth (NLCS), a combined post-processing method that incorporates non-linearity and higher-order representation into the residual propagation to address intricate node relationships effectively. Systematic evaluations show that our approach achieves remarkable average improvements of 13.2% over base prediction and 2.1% over the state-of-the-art post-processing method on six commonly used datasets. Comparisons and analyses reveal that our method enhances the prediction accuracy of nodes with complex architecture by effectively utilizing triangle relationships within graphs.
Yuanhang Shao, Xiuwen Liu 0001
IJCNN2
2024 Inductive Link Prediction in Knowledge Graphs using Path-based Neural Networks
abstract
Link prediction is a crucial research area in knowledge graphs, with many downstream applications. In many real- world scenarios, inductive link prediction is required, where predictions have to be made among unseen entities. Embeddingbased models usually need fine-tuning on new entity embeddings, and hence are difficult to be directly applied to inductive link prediction tasks. Logical rules captured by rule-based models can be directly applied to new entities with the same graph typologies, but the captured rules are discrete and usually lack generosity. Graph neural networks (GNNs) can generalize topological information to new graphs taking advantage of deep neural networks, which however may still need fine-tuning on new entity embeddings. In this paper, we propose SiaILP, a path-based model for inductive link prediction using siamese neural networks. Our model only depends on relation and path embeddings, which can be generalized to new entities without fine-tuning. Experiments show that our model achieves several new state-of-the-art performances in link prediction tasks using inductive versions of WN18RR, FB15k-237, and Nell995. Our code is available at https://github.com/canlinzhang/SiaILP.
Canlin Zhang, Xiuwen Liu 0001
IJCNN2
2024 Malicious Path Manipulations via Exploitation of Representation Vulnerabilities of Vision-Language Navigation Systems
abstract
Building on the unprecedented capabilities of large language models for command understanding and zero-shot recognition of multi-modal vision-language transformers, visual language navigation (VLN) has emerged as an effective way to address multiple fundamental challenges toward a natural language interface to robot navigation. However, such vision-language models are inherently vulnerable due to the lack of semantic meaning of the underlying embedding space. Using a recently developed gradient-based optimization procedure, we demonstrate that images can be modified imperceptibly to match the representation of totally different images and unrelated texts for a vision-language model. Building on this, we develop algorithms that can adversarially modify a minimal number of images so that the robot will follow a route of choice for commands that require a number of landmarks. We demonstrate that experimentally using a recently proposed VLN system; for a given navigation command, a robot can be made to follow drastically different routes. We also develop an efficient algorithm to detect such malicious modifications reliably based on the fact that the adversarially modified images have much higher sensitivity to added Gaussian noise than the original images.
Chashi Mahiul Islam, Shaeke Salman, Montasir Shams, Xiuwen Liu 0001
IROS4
2023 Facial Deepfake Detection Using Gaussian Processes
Uzoamaka Ezeakunne, Xiuwen Liu 0001
PSIVT2
2022 Geometric Analysis and Metric Learning of Instruction Embeddings
abstract
Embeddings for instructions have been shown to be essential for software reverse engineering and automated program analysis. However, due to the complexity of dependencies and inherent variability of instructions, instruction embeddings using models that are successful for natural language processing may not be effective. In this paper, we perform geometric analysis of instruction embeddings at the token level and instruction family level, showing much greater variability and leading to degraded performance on intrinsic analyses. Then we propose to use metric learning to improve the relationships among instructions using triplet loss. Our results on a large dataset of instruction groups shows significant improvements. We also provide a theoretical analysis of the instruction embeddings by looking at the BERT components and characteristics of inner-product matrices for attention in the transformer blocks. The code will be available publicly after the paper is accepted for publication.
Sajib Biswas, Timothy Barao, John Lazzari, Jeret McCoy, Xiuwen Liu 0001, Alexander Kostandarithes
IJCNN5
2022 Dense Embeddings Preserving the Semantic Relationships in WordNet
abstract
In this paper, we provide a novel way to generate low dimensional vector embeddings for the noun and verb synsets in WordNet, where the hypernym-hyponym relationship is preserved in the embeddings. We call this embedding the Sense Spectrum (and Sense Spectra for embeddings). In order to create suitable labels for the training of sense spectra, we designed a new similarity measurement for noun and verb synsets in WordNet. We call this similarity measurement the Hypernym Intersection Similarity (HIS), since it compares the common and unique hypernyms between two synsets. Our experiments show that on the noun and verb pairs of the SimLex-999 dataset, HIS outperforms the three similarity measurements in WordNet. Moreover, to the best of our knowledge, the sense spectra provide the first dense synset embeddings that preserve the semantic relationships in WordNet.
Canlin Zhang, Xiuwen Liu 0001
IJCNN2
2021 Optimism/Pessimism Prediction of Twitter Messages and Users Using BERT with Soft Label Assignment
abstract
Being able to accurately predict users' outlooks on social media platforms is important for developing educational public health interventions. In this paper, utilizing the contextualized representations provided by BERT, we propose new models to predict optimism/pessimism by fine-tuning BERT. By paying attention to the negation and other syntactic patterns, the self-attention mechanism via the transformer in BERT leads to more accurate models. For example, using the commonly used dataset for optimism/pessimism prediction with the proposed Soft Label Assignment (SLA), we have achieved 100% prediction accuracy at the user level and 97.10% at the tweet message level on the test set excluding the neutral ones. Furthermore, utilizing available training sample scores, we assign labels softly, which improves the generalization of the models and therefore their performance further. We additionally analyze the attention heads to illustrate the mechanisms of our models to classify different messages and demonstrate the connections between emotions and optimism/pessimism prediction.
Ali Alshahrani, Meysam Ghaffari, Kobra Amirizirtol, Xiuwen Liu 0001
IJCNN4
2021 How Can the [MASK] Know? The Sources and Limitations of Knowledge in BERT
abstract
We explore the idea of using the pre-trained BERT as a source of factual knowledge, analyze which components of the model are responsible for its ability to answer questions requiring factual knowledge, and study the transferability of the knowledge to downstream tasks. Our experiments show that the Language Modeling Head is indispensable for predicting facts, implying that transferability of any knowledge captured in the model is limited. While the dominant approach to researching how knowledge is stored in language models focuses on tailoring question formulation to optimize the retrieval quality, we find question patterns easily understood by humans that confuse BERT to the point that the answer does not make sense. The nature of the found patterns implies that the stored knowledge is fragile and based on token co-occurrence in the training set used during pre-training, rather than generalization or inference. Moreover, using a novel, hand-crafted dataset we show that BERT is vulnerable to common misconceptions, which could have fatal effects on downstream applications. Overall, we conclude that BERT offers low and unreliable performance out of the box. Jupyter notebooks with experiments are available on GitHub.11Project URL: https://github.com/tenpercent/knowledge-and-confusion
Maksim Podkorytov, Daniel Bis, Xiuwen Liu 0001
IJCNN3
2021 Too Much in Common: Shifting of Embeddings in Transformer Language Models and its Implications
abstract
The success of language models based on the Transformer architecture appears to be inconsistent with observed anisotropic properties of representations learned by such models.We resolve this by showing, contrary to previous studies, that the representations do not occupy a narrow cone, but rather drift in common directions.At any training step, all of the embeddings except for the ground-truth target embedding are updated with gradient in the same direction.Compounded over the training set, the embeddings drift and share common components, manifested in their shape in all the models we have empirically tested.Our experiments show that isotropy can be restored using a simple transformation. 1
Daniel Bis, Maksim Podkorytov, Xiuwen Liu 0001
NAACL-HLT3
2020 Identifying Optimism and Pessimism in Twitter Messages Using XLNet and Deep Consensus
abstract
Modeling optimism and pessimism accurately in social media has important applications to personal health individually and society wellness collectively. In this paper, we predict optimism and pessimism in Twitter messages by building multiple models on top of XLNet, an integrated model using multiple auto-regressive language models to capture left and right contexts jointly in sentences. Utilizing multiple-head self attentions via multi-layer transformers, XLNet models are able to model negations and other semantic relationships by paying attentions to crucial and important words, leading to more accurate predictive models for optimism and pessimism. For example, using XLNet models, we have improved the state of the art accuracy of 90.32% to 96.45%, a 63.32% error reduction on a benchmark dataset. Based on the observations that all deep models should generalize to new messages based on the same training samples, we train multiple predictive models and use the consensus to further improve the accuracy on subsets of the test samples. We also demonstrate that positive emotions and sentiments in optimistic messages are much more common while negative emotions and sentiments are more so in pessimistic ones using XLNet models finetuned for emotion classification and sentiment analysis. The proposed models could be used for understanding optimism and pessimism in social media.
Ali Alshahrani, Meysam Ghaffari, Kobra Amirizirtol, Xiuwen Liu 0001
IJCNN4
2020 Explaining AI for Malware Detection: Analysis of Mechanisms of MalConv
abstract
In recent years, machine learning has been used in a very wide variety of applications and malware detection is no exception. Because of its fast and widespread adaptation to various diverse fields, machine learning can, and often is, treated as a black box. The disadvantage of doing so is that the decisions can often be difficult to interpret which can be especially challenging in the field of malware detection. Training deep neural networks also requires a vast amount of data from all classes which can be quite challenging in the field of proprietary software, specially for smaller research labs. In this paper, we introduce a framework which interpolates between samples of different classes at different layers to see how a deep network architecture generalizes to samples that are not in the training set, explaining the results of deep networks in real-world testing. Using this framework, we attempt to demystify the mechanisms behind the MalConv architecture [1] by analyzing the weights and gradients of multiple layers in its architecture and decipher what the architecture learns by analyzing raw bytes from the binary. For this architecture, our analysis shows that the network assigns much higher weights to specific portions of the executable Indicating that these portions contribute significantly more to the classification than other portions of the executable. Through the proposed framework, we can explain the mechanisms behind machine learning algorithms and explain their decisions better. In addition, the analyses will allow us to look inside existing networks without training them from scratch.
Shamik Bose, Timothy Barao, Xiuwen Liu 0001
IJCNN3
2020 Effects of Architecture and Training on Embedding Geometry and Feature Discriminability in BERT
abstract
Natural language processing has improved substantially in the last few years due to the increased computational power and availability of text data. Bidirectional Encoder Representations from Transformers (BERT) have further improved the performance by using an auto-encoding model that incorporates larger bidirectional contexts. However, the underlying mechanisms of BERT for its effectiveness are not well understood. In the paper we investigate how the BERT architecture and its pretraining protocol affect the geometry of its embeddings and the effectiveness of its features for classification tasks. As an autoencoding model, during pre-training, it produces representations that are context dependent and at the same time must be able to "reconstruct" the original input sentences. The complex interactions of the two via transformers lead to interesting geometric properties of the embeddings and subsequently affect the inherent discriminability of the resulting representations. Our experimental results illustrate that the BERT models do not produce "effective" contextualized representations for words and their improved performance may mainly be due to fine-tuning or classifiers that model the dependencies explicitly by encoding syntactic patterns in the training data.
Maksim Podkorytov, Daniel Bis, Jinglun Cai, Kobra Amirizirtol, Xiuwen Liu 0001
IJCNN5
2020 DeepConsensus: Consensus-based Interpretable Deep Neural Networks with Application to Mortality Prediction
abstract
Deep neural networks have achieved remarkable success in various challenging tasks. However, the black-box nature of such networks is not acceptable to critical applications, such as healthcare. In particular, the existence of adversarial examples and their overgeneralization to irrelevant, out-ofdistribution inputs with high confidence makes it difficult, if not impossible, to explain decisions by such networks. In this paper, we analyze the underlying mechanism of generalization of deep neural networks and propose an (n, k) consensus algorithm which is insensitive to adversarial examples and can reliably reject out- of-distribution samples. Furthermore, the consensus algorithm is able to improve classification accuracy by using multiple trained deep neural networks. To handle the complexity of deep neural networks, we cluster linear approximations of individual models and identify highly correlated clusters among different models to capture feature importance robustly, resulting in improved interpretability. Motivated by the importance of building accurate and interpretable prediction models for healthcare, our experimental results on an ICU dataset show the effectiveness of our algorithm in enhancing both the prediction accuracy and the interpretability of deep neural network models on one-year patient mortality prediction. In particular, while the proposed method maintains similar interpretability as conventional shallow models such as logistic regression, it improves the prediction accuracy significantly.
Shaeke Salman, Seyedeh Neelufar Payrovnaziri, Xiuwen Liu 0001, Pablo Rengifo-Moreno, Zhe He 0001
IJCNN3
2020 Towards Quantifying Intrinsic Generalization of Deep ReLU Networks
abstract
Understanding the underlying mechanisms that enable the empirical successes of deep neural networks is essential for further improving their performance and explaining such networks. Towards this goal, a specific question is how to explain the "surprising" behavior of the same over-parametrized deep neural networks that can generalize well on real datasets and at the same time "memorize" training samples when the labels are randomized. In this paper, we demonstrate that deep ReLU networks generalize from training samples to new points via piece-wise linear interpolation. We provide a quantified analysis on the generalization ability of a deep ReLU network: Given a fixed point x and a fixed direction in the input space $\mathcal{S}$, there is always a segment such that any point on the segment will be classified the same as the fixed point x. We call this segment the generalization interval. We show that the generalization intervals of a ReLU network behave similarly along pairwise directions between samples of the same label in both real and random cases on the MNIST and CIFAR-10 datasets. This result suggests that the same interpolation mechanism is used in both cases. Additionally, for datasets using real labels, such networks provide a good approximation of the underlying manifold in the data, where the changes are much smaller along tangent directions than along normal directions. Our systematic experiments demonstrate for the first time that such deep neural networks generalize through the same interpolation and explain the differences between their performance on datasets with real and random labels.
Shaeke Salman, Canlin Zhang, Xiuwen Liu 0001, Washington Mio
IJCNN3
2020 Explainable artificial intelligence models using real-world electronic health record data: a systematic scoping review
abstract
OBJECTIVE: To conduct a systematic scoping review of explainable artificial intelligence (XAI) models that use real-world electronic health record data, categorize these techniques according to different biomedical applications, identify gaps of current studies, and suggest future research directions. MATERIALS AND METHODS: We searched MEDLINE, IEEE Xplore, and the Association for Computing Machinery (ACM) Digital Library to identify relevant papers published between January 1, 2009 and May 1, 2019. We summarized these studies based on the year of publication, prediction tasks, machine learning algorithm, dataset(s) used to build the models, the scope, category, and evaluation of the XAI methods. We further assessed the reproducibility of the studies in terms of the availability of data and code and discussed open issues and challenges. RESULTS: Forty-two articles were included in this review. We reported the research trend and most-studied diseases. We grouped XAI methods into 5 categories: knowledge distillation and rule extraction (N = 13), intrinsically interpretable models (N = 9), data dimensionality reduction (N = 8), attention mechanism (N = 7), and feature interaction and importance (N = 5). DISCUSSION: XAI evaluation is an open issue that requires a deeper focus in the case of medical applications. We also discuss the importance of reproducibility of research work in this field, as well as the challenges and opportunities of XAI from 2 medical professionals' point of view. CONCLUSION: Based on our review, we found that XAI evaluation in medicine has not been adequately and formally practiced. Reproducibility remains a critical concern. Ample opportunities exist to advance XAI research in medicine.
Seyedeh Neelufar Payrovnaziri, Zhaoyi Chen, Pablo Rengifo-Moreno, Tim Miller 0001, Jiang Bian 0001, Jonathan H. Chen, Xiuwen Liu 0001, Zhe He 0001
J. Am. Medical Informatics Assoc.7
2019 High-resolution home location prediction from tweets using deep learning with dynamic structure
abstract
Timely and high-resolution estimates of the home locations of a sufficiently large subset of the population are critical for applications such as disaster response and public health. However, conventional data sources, such as census and surveys, have a substantial time lag and cannot capture seasonal trends. Recently, the large user-base and real-time nature of social media data have been leveraged to address this problem. However, inherent sparsity and noise, along with large estimation uncertainty in home locations, have limited their effectiveness. In this paper, we develop a deep-learning solution that deals with the sparsity and noise of social media data. We obtained over 90% accuracy for large subsets on a commonly used dataset. Systematic comparisons show that our method gives the highest accuracy both for the entire sample and for subsets.
Meysam Ghaffari, Ashok Srinivasan, Xiuwen Liu 0001
ASONAM3
2019 Next-generation high-resolution vector-borne disease risk assessment
abstract
Vector-borne diseases cause more than 1 million deaths annually. Estimates of epidemic risk at high spatial resolutions can enable effective public health interventions. Our goal is to identify the risk of importation of such diseases into vulnerable cities at the granularity of neighborhoods. Conventional models cannot achieve such spatial resolution, especially in real-time. Besides, they lack real-time data on demographic heterogeneity, which is vital for accurate risk estimation. Social media, such as Twitter, promise data from which demographic and spatial information could be inferred in real-time. On the other hand, such data can be noisy and inaccurate. Our novel approach leverages Twitter data, using machine learning techniques at multiple spatial scales to overcome its limitations, to deliver results at the desired resolution. We validate our method against the Zika outbreak in Florida in 2016. Our main contribution lies in proposing a novel approach that uses machine learning on social media data to identify the risk of vector-borne disease importation at a sufficiently fine spatial resolution to permit effective intervention. It will lead to a new generation of epidemic risk assessment models, promising to transform public health by identifying specific locations for targeted intervention.
Meysam Ghaffari, Ashok Srinivasan, Anuj Mubayi, Xiuwen Liu 0001, Krishnan Viswanathan
ASONAM4
2019 Sparsity as the Implicit Gating Mechanism for Residual Blocks
abstract
Neural networks are the core component in the recent empirical successes of deep learning techniques in challenging tasks. Residual network (ResNet) architectures have been instrumental in improving performance in object recognition and other tasks by enabling training much deeper neural networks. Studies of residual networks reveal that they are robust to removing layers. However, it is still an open question of why residual networks behave well and how they make it feasible to train networks with many layers. In this paper, we show that sparsity of the residual blocks acts as the implicit gating mechanism. When a neuron is inactive, it behaves as a node in an information highway, allowing the information from the previous layer to pass to the next layer unchanged. As the identity function has a derivative of 1, it avoids the exploding or vanishing gradient problem that is known to contribute to the difficulty of training deep neural networks. When a neuron is active, it captures input-output relationships that are necessary to achieve good performance. By using the ReLu activation functions, residual blocks produce sparse outputs for typical inputs. We perform systematic experimental analysis on the residual blocks of trained ResNet models and show that sparsity acts as the implicit gate for deep residual networks.
Shaeke Salman, Xiuwen Liu 0001
IJCNN2
2019 An Analysis on the Learning Rules of the Skip-Gram Model
abstract
To improve the generalization of the representations for natural language processing tasks, words are commonly represented using vectors, where distances among the vectors are related to the similarity of the words. While word2vec, the state-of-the-art implementation of the skip-gram model, is widely used and improves the performance of many natural language processing tasks, its mechanism is not yet well understood. In this work, we derive the learning rules for the skip-gram model and establish their close relationship to competitive learning. In addition, we provide the global optimal solution constraints for the skip-gram model and validate them by experimental results.
Canlin Zhang, Xiuwen Liu 0001, Daniel Bis
IJCNN2
2019 Biomedical word sense disambiguation with bidirectional long short-term memory and attention-based neural networks
abstract
BACKGROUND: In recent years, deep learning methods have been applied to many natural language processing tasks to achieve state-of-the-art performance. However, in the biomedical domain, they have not out-performed supervised word sense disambiguation (WSD) methods based on support vector machines or random forests, possibly due to inherent similarities of medical word senses. RESULTS: In this paper, we propose two deep-learning-based models for supervised WSD: a model based on bi-directional long short-term memory (BiLSTM) network, and an attention model based on self-attention architecture. Our result shows that the BiLSTM neural network model with a suitable upper layer structure performs even better than the existing state-of-the-art models on the MSH WSD dataset, while our attention model was 3 or 4 times faster than our BiLSTM model with good accuracy. In addition, we trained "universal" models in order to disambiguate all ambiguous words together. That is, we concatenate the embedding of the target ambiguous word to the max-pooled vector in the universal models, acting as a "hint". The result shows that our universal BiLSTM neural network model yielded about 90 percent accuracy. CONCLUSION: Deep contextual models based on sequential information processing methods are able to capture the relative contextual information from pre-trained input word embeddings, in order to provide state-of-the-art results for supervised biomedical WSD tasks.
Canlin Zhang, Daniel Bis, Xiuwen Liu 0001, Zhe He 0001
BMC Bioinform.3
2018 Layered Multistep Bidirectional Long Short-Term Memory Networks for Biomedical Word Sense Disambiguation
Daniel Bis, Canlin Zhang, Xiuwen Liu 0001, Zhe He 0001
BIBM3
2018 Land cover classification from multi-temporal, multi-spectral remotely sensed imagery using patch-based recurrent neural networks
Atharva Sharma, Xiuwen Liu 0001
Neural Networks2
2017 An exploration of semantic relations in neural word embeddings using extrinsic knowledge
abstract
In the recent few years, neural-network-based word embeddings have been widely used in text mining. However, the dense representations of word embeddings act as a black box and lack interpretability. Even though word embeddings are able to capture semantic regularities in free text documents, it is not clear what kinds of semantic relations can be represented by word embeddings and how semantically-related terms can be retrieved from word embeddings. In this study, we propose a novel approach to explore the semantic relations in neural embeddings using extrinsic knowledge from WordNet and Unified Medical Language System (UMLS). We trained multiple word embeddings using health-related articles in Wikipedia. We then evaluated the performance of the different word embeddings in semantic relation term retrieval tasks. This study shows that word embeddings can group terms with diverse semantic relations together.
Zhe He 0001, Xiuwen Liu 0001, Jiang Bian 0001
BIBM3
2017 A patch-based convolutional neural network for remote sensing image classification
Atharva Sharma, Xiuwen Liu 0001
Neural Networks2
2016 Fast Nearest Neighbor Search with Transformed Residual Quantization
abstract
Product quantization (PQ) and residual quantization (RQ) have been successfully used to solve fast nearest neighbor search problems thanks to their exponentially reduced complexities of both storage and computation with respect to the codebook size, Recent efforts have been focused on employing optimization strategies and seeking more effective models. Based on the observation that randomness typically increases in subsequent residual spaces, we propose a new strategy, called, transformed RQ (TRQ), that jointly learns a local transformation per residual cluster with an ultimate goal to further reduce overall quantization errors. Additionally we propose a hybrid approximate nearest search method based on the proposed TRQ and PQ. We show that our methods achieve significantly better accuracy on nearest neighbor search than both the original and the optimized PQ on several benchmark datasets.
Jiangbo Yuan, Xiuwen Liu 0001
ICMLA2
2016 TBAS: Enhancing Wi-Fi Authentication by Actively Eliciting Channel State Information
abstract
In this paper, we propose Time Bounded Anti- Spoofing (TBAS), a novel spoof detection method for W-Fi networks. TBAS is based on the simple idea of comparing the channel states embedded in packets received within a short interval and discarding the packets if the channel state change is above a reasonable level for the interval. We leverage the packet transmission policies in Wi-Fi to actively elicit packets for the channel state check without modifying the Wi-Fi protocol, such as sending a dummy data packet which will automatically elicit an acknowledgement packet. Comparing to the existing methods that wait passively for the packets from the packet sender for the channel state information, TBAS has the full control of time to collect the channel state, such that it can use stricter rules to classify the packets and achieve higher security. We solve key problems in TBAS, including a Channel State Check (CSC) procedure which determines if two packets are from the same sender, as well as a simple protocol wrapped around CSC to reduce the overhead. We implement TBAS on the Sora software defined ratio and our trace-driven emulations show that TBAS can achieve desirable False Positive and False Negative ratios with very low overhead.
Muye Liu, Avishek Mukherjee, Xiuwen Liu 0001
SECON4
2015 A path based approach to quantifying the progression of Alzheimer's disease
abstract
Histological studies suggest that different pathological changes occur in the subfields of hippocampi in aging and in Alzheimer0s disease. The pathological changes due to Alzheimer's disease follows a neural system and are consistent with subjects. The aim of this study is to study the path in vivo the changes in normal control, mild cognitive impairment and Alzheimers disease patients using high resolution MRI scanned at 3 Tesla. T1 weighted images were obtained from the ADNI database. The dataset consists of 11 control patients, 13 MCI patients, and 9 AD patients. The hippocampal subfields were segmented using the Freesurfer image-analysis suite (Version 6.0). The volume of the subfields and the path along the medial axis of the subfields were studied. The results suggest that the intensity along the medial axis shows more variation on the CA1 region than other subfields in cases of Alzheimer disease, MCI and normal control patients. The changes are more prominent in the early section of the CA1 suggesting the progressive nature of the disease. The mean intensity value along the medial axis of CA1 subfield shows change in AD, MCI and NC but similar change are not seen in other subfields.
Prabesh Kanel, Xiuwen Liu 0001
BIBM2
2015 Product tree quantization for approximate nearest neighbor search
abstract
The product quantization (PQ) performance degrades on read-world data due to the severity of dependence between feature groups. Meanwhile, tree structured vector quantization (TSVQ) often supply lower distortion than other structured vector quantizers; yet it is prohibitive to learning compact codes like PQ does considering its codebook storage. In this paper, we propose a hybrid model dubbed as product tree quantization (PTQ) that aims to relax the PQ constraints while to retain the tree-structured codebooks with reasonable size. We first show that our methods can achieve significantly better quantization performance on several large scale benchmarks. We then demonstrate the advantage for very large scale ANN search; for instance, on a 1-billion scale dataset, we have achieved on average 4% improvement in accuracy than the existing state of the art methods.
Jiangbo Yuan, Xiuwen Liu 0001
ICIP2
2015 Nano-scale context-sensitive semantic segmentation
abstract
Nano-scale imaging technologies make it possible to visualize objects at nanometer resolutions. To investigate structures and functions of interest, there is an intrinsic demand for explicit models to extract them from nano-scale data. Segmentation is one of the most critical steps in processing pipelines. However, existing segmentation methods often fail due to extremely low signal-to-noise ratio, low contrast and large data size. In this paper we propose a new context-sensitive method for segmenting three-dimensional volumes. As our method efficiently narrows the search space by using robust context cues, we achieve tractable and reliable nano-scale semantic segmentation. We demonstrate our method on a tomogram of microvilli spikes, for which our method is able to yield accurate spike segmentation and in comparison the state-of-the-art semantic segmentation methods fail due to their inability to handle signal-to-noise ratio and low contrast volumes.
Chaity Banerjee 0001, Xiuwen Liu 0001
ICIP3
2015 Liar, Liar, IM on Fire: Deceptive language-action cues in spontaneous online communication
abstract
With an increasing number of online users, the potential danger of online deception grows accordingly - as does the importance of better understanding human behavior online to mitigate these risks. One critical element to address such online threat is to identify intentional deception in spontaneous online communication. For this study, we designed an interactive online game that creates player scenarios to encourage deception. Data was collected and analyzed in October 2014 to identify certain deceptive cues. Players' interactive dialogue was analyzed using linear regression analysis. The results reveal that certain language features are highly significant predictors of deception in synchronous, spontaneous online communication.
Shuyuan Mary Ho, Jeffrey T. Hancock, Cheryl Booth, Xiuwen Liu 0001, Shashanka Surya Timmarajus, Mike Burmester
ISI4
2012 A Trusted Computing Architecture for Secure Substation Automation
David Guidry, Mike Burmester, Xiuwen Liu 0001, Jonathan Jenkins, Sean Easton, Xin Yuan 0001
CRITIS3
2012 A novel index structure for large scale image descriptor search
abstract
This paper presents a k-means based algorithm for approximate nearest neighbor search. The proposed Embedded k-Means algorithm is a two-level clustered index structure which consists of two groups of centroids; additionally, an inverted file is used for recording of the assignments. The coarse-to-fine structure achieves high search efficiency using multi-assignment operations on the coarse level. At the query stage, pruning strategies are utilized to balance the trade-off between search qualities and speeds. Our algorithm is able to explore the neighborhood space of a given query efficiently. Due to its good recall/selectivity and memory efficiency, the proposed algorithm is scalable and is able to process very large databases. Experimental results on SIFT and GIST image descriptor datasets show search performance better and comparable to the state-of-the-art approaches with lower memory usage and complexity.
Jiangbo Yuan, Xiuwen Liu 0001
ICIP2
2011 A hybrid PCA-LDA model for dimension reduction
abstract
Several variants of Linear Discriminant Analysis (LDA) have been investigated to address the vanishing of the within-class scatter under projection to a low-dimensional subspace in LDA. However, some of these proposals are ad hoc and some others do not address the problem of generalization to new data. Meanwhile, even though LDA is preferred in many application of dimension reduction, it does not always outperform Principal Component Analysis (PCA). In order to optimize discrimination performance in a more generative way, a hybrid dimension reduction model combining PCA and LDA is proposed in this paper. We also present a dimension reduction algorithm correspondingly and illustrate the method with several experiments. Our results have shown that the hybrid model outperform PCA, LDA and the combination of them in two separate stages.
Washington Mio, Xiuwen Liu 0001
IJCNN3
2011 FISH Finder: a high-throughput tool for analyzing FISH images
abstract
MOTIVATION: Fluorescence in situ hybridization (FISH) is used to study the organization and the positioning of specific DNA sequences within the cell nucleus. Analyzing the data from FISH images is a tedious process that invokes an element of subjectivity. Automated FISH image analysis offers savings in time as well as gaining the benefit of objective data analysis. While several FISH image analysis software tools have been developed, they often use a threshold-based segmentation algorithm for nucleus segmentation. As fluorescence signal intensities can vary significantly from experiment to experiment, from cell to cell, and within a cell, threshold-based segmentation is inflexible and often insufficient for automatic image analysis, leading to additional manual segmentation and potential subjective bias. To overcome these problems, we developed a graphical software tool called FISH Finder to automatically analyze FISH images that vary significantly. By posing the nucleus segmentation as a classification problem, compound Bayesian classifier is employed so that contextual information is utilized, resulting in reliable classification and boundary extraction. This makes it possible to analyze FISH images efficiently and objectively without adjustment of input parameters. Additionally, FISH Finder was designed to analyze the distances between differentially stained FISH probes. AVAILABILITY: FISH Finder is a standalone MATLAB application and platform independent software. The program is freely available from: http://code.google.com/p/fishfinder/downloads/list.
James W. Shirley, Sereyvathana Ty, Shin-ichiro Takebayashi, Xiuwen Liu 0001, David M. Gilbert
Bioinform.4
2011 Efficient Path-Based Stereo Matching With Subpixel Accuracy
abstract
This paper presents an efficient algorithm to achieve accurate subpixel matchings for calculating correspondences between stereo images based on a path-based matching algorithm. Compared with point-by-point stereo-matching algorithms, path-based algorithms resolve local ambiguities by maximizing the cross correlation (or other measurements) along a path, which can be implemented efficiently using dynamic programming. An effect of the global matching criterion is that cross correlations at all pixels contribute to the criterion; since cross correlation can change significantly even with subpixel changes, to achieve subpixel accuracy, it is no longer sufficient to first find the path that maximizes the criterion at integer pixel locations and then refine to subpixel accuracy. In this paper, by writing bilinear interpolation using integral images, we show that cross correlations at all subpixel locations can be computed efficiently and, thus, lead to a subpixel accuracy path-based matching algorithm. Our results show the feasibility of the method and illustrate significant improvement over existing path-based matching methods.
Arturo Donate, Xiuwen Liu 0001, Emmanuel G. Collins Jr.
IEEE Trans. Syst. Man Cybern. Part B2
2010 A Model of Volumetric Shape for the Analysis of Longitudinal Alzheimer's Disease Data
Xiuwen Liu 0001, Yonggang Shi, Paul M. Thompson, Washington Mio
ECCV (3)2
2010 Scale-Space Spectral Representation of Shape
abstract
We construct a scale space of shape of closed Riemannian manifolds, equipped with metrics derived from spectral representations and the Hausdorff distance. The representation depends only on the intrinsic geometry of the manifolds, making it robust to pose and articulation. The computation of shape distance involves an optimization problem over the 2p-element group of all p-bit strings, which is approached with Markov chain Monte Carlo techniques. The methods are applied to cluster surfaces in 3D space.
Jonathan Bates, Xiuwen Liu 0001, Washington Mio
ICPR2
2010 A Computational Model of Multidimensional Shape
abstract
We develop a computational model of shape that extends existing Riemannian models of curves to multidimensional objects of general topological type. We construct shape spaces equipped with geodesic metrics that measure how costly it is to interpolate two shapes through elastic deformations. The model employs a representation of shape based on the discrete exterior derivative of parametrizations over a finite simplicial complex. We develop algorithms to calculate geodesics and geodesic distances, as well as tools to quantify local shape similarities and contrasts, thus obtaining a formulation that accounts for regional differences and integrates them into a global measure of dissimilarity. The Riemannian shape spaces provide a common framework to treat numerous problems such as the statistical modeling of shapes, the comparison of shapes associated with different individuals or groups, and modeling and simulation of shape dynamics. We give multiple examples of geodesic interpolations and illustrations of the use of the models in brain mapping, particularly, the analysis of anatomical variation based on neuroimaging data.
Xiuwen Liu 0001, Yonggang Shi, Ivo D. Dinov, Washington Mio
Int. J. Comput. Vis.1
2009 Linear Representation Learning Using Sphere Factor Analysis
abstract
Representation learning is a fundamental challenge for feature selection and plays an important role in applications such as dimension reduction, data mining and object recognition. Traditional linear representation methods, such as principal component analysis (PCA), independent component analysis (ICA) and linear discriminate analysis (LDA), have good performance on certain applications based on corresponding criteria. However, these linear representation methods are not optimal in general. Sphere factor analysis (SFA) is a recently proposed method which provides a general framework for optimization problems. In term of object recognition, SFA seeks to optimize the discriminant ability of the nearest neighbor classifier for data classification and labeling. Based on the geometry structure of the search space, a gradient search algorithms have been applied to obtain an optimal basis. A detail presentation of these algorithm is given in this paper. Furthermore, to speed up the search procedure of SFA, a two-stage strategy is proposed, which we called two-stage SFA. We illustrate the effectiveness of the original SFA and two-stage SFA methods on UCI data sets and two face data sets.
Yiming Wu 0010, Xiuwen Liu 0001, Washington Mio
ICMLA2
2009 Image retrieval based on intrinsic spectral histogram representation
abstract
The spectral histogram features are not invariant to images' scale transformation. We investigate in the technique of scale-invariant feature extraction. An approach is proposed to get the characteristic scales based on the reliable keypoints which are detected as local extrema in combination of normalized derivatives. Making use of characteristic scale of image content, which reflects characteristic length of a corresponding image structure, we are able to contribute in eliminating the effect of image transformation. In our content based image retrieval process, images are firstly resized by the characteristic scale and then represented based on the statistics of their spectral components and a linear dimension reduction technique optimizing class differentiation with respect to cross-correlation metrics of spectral histograms. Our retrieval consists of a preliminary classification step to index images in dataset and a following step of class by class retrieval. Experiments are performed on the Corel database and the outcome is compared with those of some existing work.
Yuhua Zhu, Xiuwen Liu 0001, Washington Mio
IJCNN2
2009 Optimal dimension reduction for image retrieval with correlation metrics
abstract
We investigate content-based image retrieval employing a representation of images based on the statistics of their spectral components and a new linear dimension reduction technique. This linear dimension reduction technique is designed to optimize class separation with respect to metrics derived from cross-correlation of spectral histograms. Our approach to retrieval involves a preliminary classification step to index images in a database followed by a class-by-class retrieval step. We carry out several experiments with the Corel database and compare the outcome with several results previously reported in the literature.
Yuhua Zhu, Washington Mio, Xiuwen Liu 0001
IJCNN3
2009 Rut Tracking and Steering Control for Autonomous Rut Following
abstract
Ruts formed as a result of vehicle traversal on soft ground are used by expert off road drivers because they can improve vehicle safety on turns and slopes thanks to the extra lateral force they provide to the vehicle. In this paper we propose a rut detection and tracking algorithm for autonomous ground vehicles (AGVs) equipped with a laser range finder. The proposed algorithm utilizes an extended Kalman filter (EKF) to recursively estimate the parameters of the rut and the relative position and orientation of the vehicle with respect to the ruts. Simulation results show that the approach is promising for future implementation.
Camilo Ordonez, Oscar Chuy, Emmanuel G. Collins Jr., Xiuwen Liu 0001
SMC4
2009 Shape of Elastic Strings in Euclidean Space
Washington Mio, John Christopher Bowers, Xiuwen Liu 0001
Int. J. Comput. Vis.3
2008 Efficient and accurate subpixel path based stereo matching
abstract
This paper presents an efficient algorithm to achieve accurate subpixel matchings for calculating correspondences between stereo images based on a path-based matching algorithm. Compared to point-by-point stereo matching algorithms, path-based algorithms resolve local ambiguities by maximizing the cross correlation (or other measurements) along a path, which can be implemented efficiently using dynamic programming. An effect of the global matching criterion is that the cross correlation at all pixels can contribute to the criterion; since cross correlation can change significantly even with subpixel changes, to achieve subpixel accuracy, it is no longer sufficient to first find the path that maximizes the criterion and then refine to subpixel accuracy. In this paper, by writing bilinear interpolation using integral images, we show that cross correlations at all subpixel locations can be computed efficiently and thus lead to a subpixel accuracy path based matching algorithm. Our results show the feasibility of the method and illustrate the significant improvements over the original path-based matching method.
Arturo Donate, Xiuwen Liu 0001, Emmanuel G. Collins Jr.
ICPR3
2008 Kernel functions for robust 3D surface registration with spectral embeddings
abstract
Registration of 3D surfaces is a critical step for shape analysis. Recent studies show that spectral representations based on intrinsic pairwise geodesic distances between points on surfaces are effective for registration and alignment due to their invariance under rigid transformations and articulations. Kernel functions are often applied to the pairwise geodesic distances to make the registration process based on spectral embedding robust to elastic deformations. The Gaussian kernel is most commonly used, but the effect of the choice of the kernel function has not been studied in the previous works. In this paper, we compare the results obtained with several different choices and show empirically that significant improvements can be achieved in shape registration with appropriate choices.
Xiuwen Liu 0001, Arturo Donate, Matthew Jemison, Washington Mio
ICPR1
2008 Transductive optimal component analysis
abstract
We propose a new transductive learning algorithm for learning optimal linear representations that utilizes unlabeled data. We pose the problem of learning linear representations as an optimization one on the underlying nonlinear manifold. An additional term is used to prefer representations with large ldquomarginsrdquo when classifying unlabeled data in the nearest classifier sense, a generalization of transductive support vector machines to learning representations. Experimental results of the proposed algorithm on face recognition data sets show the potential significant improvement for classification accuracy on test sets.
Yuhua Zhu, Yiming Wu 0010, Xiuwen Liu 0001, Washington Mio
ICPR3
2008 Models of Normal Variation and Local Contrasts in Hippocampal Anatomy
Washington Mio, Yonggang Shi, Ivo D. Dinov, Xiuwen Liu 0001, Natasha Leporé, Franco Lepore, Madeleine Fortin, Patrice Voss, Maryse Lassonde, Paul M. Thompson
MICCAI (2)5
2008 Two-stage optimal component analysis
Yiming Wu 0010, Xiuwen Liu 0001, Washington Mio, Kyle A. Gallivan
Comput. Vis. Image Underst.2
2008 Learning representations for object classification using multi-stage optimal component analysis
Yiming Wu 0010, Xiuwen Liu 0001, Washington Mio
Neural Networks2
2007 Modeling Brain Anatomy with 3D Arrangements of Curves
abstract
We employ 3D arrangements of curves to represent and analyze biological shapes, in particular, the anatomy of the human brain. The arrangements of curves may vary from fairly sparse - such as a collection of sulcal lines that coarsely approximates the global shape of the brain - to very dense decompositions of the cortical surface into space curves. A space of shapes of such arrangements is constructed equipped with geodesic metrics that can be used in conjunction with curve registration techniques to quantify shape resemblance or dissimilarity, as well as to identify the regions where anatomical differences are most pronounced. The metric is applied to the panellation and labeling of configurations associated with the left and right hemispheres of the brain. Examples are also given of geodesic interpolations between decompositions into space curves of surfaces representing the entire left hemisphere of the brain.
Washington Mio, John Christopher Bowers, Monica K. Hurdal, Xiuwen Liu 0001
ICCV4
2007 Content-Based Image Categorization and Retrieval using Neural Networks
abstract
We propose a neural network based method for organizing images for content-based image retrieval. We use spectral histogram features, the histograms of filtered images to capture the spatial relationship among pixels as well as global appearance of images. We then find the optimal combination of spectral histogram features using optimal factor analysis to reduce the dimension of features and maximize the discrimination. The reduced features are then used as input to a multiple layer perceptron, which is trained to categorize images based on content using back propagation. For a query image, images are retrieved from different classes based on the categorization probability for the query image. Experimental results on a subset of Corel dataset demonstrate the effectiveness of the proposed method and comparisons show that the proposed method gives significant improvement over other methods.
Yuhua Zhu, Xiuwen Liu 0001, Washington Mio
ICME2
2007 Scalable optimal linear representation for face and object recognition
abstract
Optimal component analysis (OCA) is a linear method for feature extraction and dimension reduction. It has been widely used in many applications such as face and object recognitions. The optimal basis of OCA is obtained through solving an optimization problem on a Grassmann manifold. However, one limitation of OCA is the computational cost becoming heavy when the number of training data is large, which prevents OCA from efficiently applying in many real applications. In this paper, a scalable OCA (S-OCA) that uses a two-stage strategy is developed to bridge this gap. In the first stage, we cluster the training data using K-means algorithm and the dimension of data is reduced into a low dimensional space. In the second stage, OCA search is performed in the reduced space and the gradient is updated using an numerical approximation. In the process of OCA gradient updating, instead of choosing the entire training data, S-OCA randomly chooses a small subset of the training images in each class to update the gradient. This achieves stochastic gradient updating and at the same time reduces the searching time of OCA in orders of magnitude. Experimental results on face and object datasets show efficiency of the S-OCA method, in term of both classification accuracy and computational complexity.
Yiming Wu 0010, Xiuwen Liu 0001, Washington Mio
ICMLA2
2007 Multi-Stage Optimal Component Analysis
abstract
Optimal component analysis (OCA) uses a stochastic gradient optimization process to find optimal representations for general criteria and shows good performance in object recognition applications. However, OCA often requires extensive computation for gradient estimation and linear representation updating. To significantly reduce the required computation, in this paper, a multi-stage learning process is proposed which decomposes the original optimization problem into several levels. As the learning process at each level starts with a good initial point obtained from next level, the multistage OCA algorithm can speed up the original algorithm significantly and make OCA learning feasible for many applications. We illustrate the effectiveness of the proposed method on the application of face classification.
Yiming Wu 0010, Xiuwen Liu 0001, Washington Mio
IJCNN2
2006 Recognition using Rapid Classification Tree
abstract
This paper proposes a method to achieve object classification with high throughput and accuracy using a rapid classification tree. To achieve this, we decouple the training and test stages. During the training stage, we learn optimal discriminatory features from the training set and then train a classifier with high accuracy. Then we create a classification tree, where each node uses a lookup table to store the solutions, resulting high throughput at the test stage. To make the lookup tables feasible for applications, we learn a projection matrix through stochastic optimization. We illustrate the effectiveness of the proposed method using several datasets; our results show the proposed method achieves often several orders of magnitudes of improvement in throughput while maintaining a similar accuracy.
Keith Haynes, Xiuwen Liu 0001, Washington Mio
ICIP2
2006 Splitting Factor Analysis and Multi-Class Boosting
abstract
We develop splitting factor analysis (SFA), a novel linear model selection technique for dimension reduction that seeks to optimize the discriminative ability of the nearest neighbor classifier for data classification and labeling. We also discuss methodology for data kernelization that can be used in conjunction with any model selection technique. Applied to SFA, it leads to KSFA, a powerful new technique for the analysis of datasets with essential nonlinearities underlying their structures. For computational efficiency in the analysis of large datasets, we combine weak KSFA classifiers with multi-class boosting techniques. Several applications to image-based classification are discussed.
Xiuwen Liu 0001, Washington Mio
ICIP1
2006 Landmark Representation of Shapes and Fisher-Rao Geometry
abstract
We develop computational strategies to calculate geodesies and geodesic distances between plane shapes represented by mixture of Gaussians centered at landmark points with a fixed variance with respect to the information-theoretic Fisher-Rao metric. This representation and metric have been investigated recently by Peter and Rangarajan, but a feasible computational approach was not provided. The algorithms developed are applied to shape clustering and the results are compared to those obtained with other methods.
Washington Mio, Xiuwen Liu 0001
ICIP2
2006 Rotation Invariant Face Detection using Spectral Histograms and Support Vector Machines
abstract
This paper presents a face detection method that detects faces with arbitrary rotation in the image plane. In this method, images are represented using a spectral histogram representation consisting of marginal distributions of filtered images. A support vector machine with an R.B.F. kernel is chosen as the classifier, which is trained on 4500 face and 8000 non-face images. The choice of filters allows a large degree of rotation invariance and by shuffling the marginals of certain filters, invariance to arbitrary rotation is achieved. A distinctive advantage of our method is that the invariance is achieved largely through the underlying representation while in other methods the invariance is typically achieved by detecting faces at a large number of different angles. The proposed method is tested on standard data sets and comparisons with other methods show that our method gives the best detection performance with respect to detection rate and false positives.
Christopher A. Waring, Xiuwen Liu 0001
ICIP2
2006 Two-Stage Optimal Component Analysis
abstract
Linear representations are widely used to reduce dimension in applications involving high dimensional data. While specialized procedures exist for certain optimality criteria, such as principle component analysis (PCA) and Fisher discriminant analysis (FDA), they can not be generalized for more general criteria. To overcome this fundamental limitation, optimal component analysis (OCA) uses a stochastic gradient optimization procedure intrinsic to the manifold giving by the constraints of applications and therefore gives a procedure for finding optimal representations for general criteria. However, due to its generality nature, OCA often requires extensive computation for gradient estimation and updating. To significantly reduce the required computation, in this paper, we propose a two-stage method by first reducing the dimension of input to a smaller one (but larger than the final resulting dimension) using a computationally efficient method and then performing OCA in the reduced space. This reduces the computation time from days to minutes on widely used databases, making OCA learning feasible for many applications. Additionally, since the reduced space is much smaller, the stochastic gradient optimization tends to be more efficient. We illustrate the effectiveness of the proposed method on face classification.
Yiming Wu 0010, Xiuwen Liu 0001, Washington Mio, Kyle A. Gallivan
ICIP2
2006 Contour Inferences for Image Understanding
Washington Mio, Anuj Srivastava, Xiuwen Liu 0001
Int. J. Comput. Vis.3
2006 Face recognition using optimal linear components of range images
Anuj Srivastava, Xiuwen Liu 0001, Curt Hesher
Image Vis. Comput.2
2005 Tools for application-driven linear dimension reduction
Anuj Srivastava, Xiuwen Liu 0001
Neurocomputing2
2005 Statistical Shape Analysis: Clustering, Learning, and Testing
abstract
Using a differential-geometric treatment of planar shapes, we present tools for: 1) hierarchical clustering of imaged objects according to the shapes of their boundaries, 2) learning of probability models for clusters of shapes, and 3) testing of newly observed shapes under competing probability models. Clustering at any level of hierarchy is performed using a mimimum variance type criterion criterion and a Markov process. Statistical means of clusters provide shapes to be clustered at the next higher level, thus building a hierarchy of shapes. Using finite-dimensional approximations of spaces tangent to the shape space at sample means, we (implicitly) impose probability models on the shape space, and results are illustrated via random sampling and classification (hypothesis testing). Together, hierarchical clustering and hypothesis testing provide an efficient framework for shape retrieval. Examples are presented using shapes and images from ETH, Surrey, and AMCOM databases.
Anuj Srivastava, Shantanu H. Joshi, Washington Mio, Xiuwen Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2005 Face detection using spectral histograms and SVMs
abstract
We present a face detection method using spectral histograms and support vector machines (SVMs). Each image window is represented by its spectral histogram, which is a feature vector consisting of histograms of filtered images. Using statistical sampling, we show systematically the representation groups face images together; in comparison, commonly used representations often do not exhibit this necessary and desirable property. By using an SVM trained on a set of 4500 face and 8000 nonface images, we obtain a robust classifying function for face and non-face patterns. With an effective illumination-correction algorithm, our system reliably discriminates face and nonface patterns in images under different kinds of conditions. Our method on two commonly used data sets give the best performance among recent face-detection ones. We attribute the high performance to the desirable properties of the spectral histogram representation and good generalization of SVMs. Several further improvements in computation time and in performance are discussed.
Christopher A. Waring, Xiuwen Liu 0001
IEEE Trans. Syst. Man Cybern. Part B2
2004 Hierarchical Organization of Shapes for Efficient Retrieval
Shantanu H. Joshi, Anuj Srivastava, Washington Mio, Xiuwen Liu 0001
ECCV (3)4
2004 Learning and Bayesian Shape Extraction for Object Recognition
Washington Mio, Anuj Srivastava, Xiuwen Liu 0001
ECCV (4)3
2004 Optimal Linear Representations of Images for Object Recognition
abstract
Although linear representations are frequently used in image analysis, their performances are seldom optimal in specific applications. This paper proposes a stochastic gradient algorithm for finding optimal linear representations of images for use in appearance-based object recognition. Using the nearest neighbor classifier, a recognition performance function is specified and linear representations that maximize this performance are sought. For solving this optimization problem on a Grassmann manifold, a stochastic gradient algorithm utilizing intrinsic flows is introduced. Several experimental results are presented to demonstrate this algorithm.
Xiuwen Liu 0001, Anuj Srivastava, Kyle A. Gallivan
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Optimal Linear Representations of Images for Object Recognition
abstract
Simplicity of linear representations (of images) makes them a popular tool in imaging analysis applications such as object recognition and image classification. Although several linear representations, namely PCA (principal component analysis), ICA, and FDA (Fisher discriminant analysis), have frequently been used, these representations are generally far from optimal in terms of actual application performance. We argue that representations should be chosen with respect to the application and the databases involved. Fixing an application, say object recognition, and assuming that recognition performance is computable for any linear basis (given a classifier and a database), we propose a Monte Carlo simulated annealing method that leads to optimal linear representations by maximizing the recognition performance over all fixed-rank subspaces. We illustrate this method on two popular databases.
Xiuwen Liu 0001, Anuj Srivastava, Kyle A. Gallivan
CVPR (1)1
2003 Sparse linear representations for recognition
abstract
Recently, it has been argued that sparse coding is an important principle for recognition, which has been used effectively to derive filters with desirable properties. However there is no effective algorithm to link the sparse coding principle to the recognition performance. Our experiments show that commonly used sparse bases often give worse recognition performance compared to other linear bases. In this paper, we propose a criterion consisting of weighted combination of recognition performance and sparseness. Using a Monte Carlo simulated annealing algorithm, we obtain linear bases with sparse representation as well as good recognition performance. We also find an interesting relationship among commonly used linear representations by comparing their sparseness and recognition performance.
Xiuwen Liu 0001
IJCNN2
2003 Integrated learning of linear representations
abstract
While the importance of representations for recognition has been widely recognized, in practice the choice of representations is often limited and applications are forced to choose relatively the best one among the available. In this paper, we advocate an integrated learning framework where the representation is learned with respect to a chosen performance criterion. For linear representations, this problem is posed as an optimization one on the underlying manifold determined by the constraints of the application, where the manifolds related to typical applications are given. To develop computationally effective algorithms, the underlying geometric structures are exploited. We demonstrate the feasibility and effectiveness of the proposed framework by finding optimal linear filters for recognition with other additional properties.
Xiuwen Liu 0001, Aruj Srivastava
IJCNN1
2003 On intrinsic generalization of low dimensional representations of images for recognition
abstract
Low dimensional representations of images impose equivalence relations in the image space; the induced equivalence class of an image is named as its intrinsic generalization. The intrinsic generalization of a representation provides a novel way to measure its generalization and leads to more fundamental insights than the commonly used recognition performance, which is heavily influenced by the choice of training and test data. We demonstrate the limitations of linear subspace representations by sampling their intrinsic generalization, and propose a nonlinear representation that overcomes these limitations. The proposed representation projects images nonlinearly into the marginal densities of their filter responses, followed by linear projections of the marginals. We have used experiments on large datasets to show that the representations that have better intrinsic generalization also lead to a better recognition performance.
Xiuwen Liu 0001, Anuj Srivastava, DeLiang Wang
IJCNN1
2003 Spectral histogram based face detection
abstract
This paper adopts a generic feature representation and applies it to the task of face detection as an appearance-based case. The distribution of faces and non-faces are modeled from the marginal distribution of filter responses. The face detection algorithm proposed here uses the spectral representation of a 21/spl times/21 image window as input to a multiple layer perceptron for classification. The classifier is trained with the backpropagation learning rule. A simple method to correct nonuniform illuminance is used to normalize all training and test images. Testing is done on a standard data set and compared to the work of others.
Christopher A. Waring, Xiuwen Liu 0001
IJCNN2
2003 Hierarchical learning of optimal linear representations
abstract
Due to their efficiency, linear representations are widely used in appearance-based recognition. However, frequently used ones, such as PCA, ICA, and FDA, do not provide optimal performance as empirical studies have reported contradictory conclusions in the literature. To overcome this problem and provide an algorithm for finding the optimal linear representations for different applications, a Monte Carlo Markov chain based optimization algorithm was recently proposed and its effectiveness has been demonstrated on a number of datasets. By formulating the problem on Grassmann manifolds, the algorithm is computationally efficient when the image size is relatively small. When images in typical applications are used, the optimization process is time consuming. In this paper, to speed up the algorithm, we propose a hierarchical learning one. The proposed algorithm decomposes the optimization in the given image space into several stages organized according to hierarchical layers. Given an image space, first its dimension is reduced using a shrinkage matrix and the optimization is then performed in the reduced space. By expanding the obtained optimal subspace in the reduced one in a specified way, we show analytically that the performance is maintained. By applying the decomposition procedure recursively, a hierarchy of layers can be formed. This speeds up the original algorithm significantly as the search is done mainly in reduced spaces. The effectiveness of hierarchical learning is illustrated on a popular database, where the computation time is reduced by 600,000 factors compared to the original algorithm.
Xiuwen Liu 0001
IJCNN2
2003 Geometric Analysis of Constrained Curves
abstract
We present a geometric approach to statistical shape analysis of closed curves in images. The basic idea is to specify a space of closed curves satisfying given constraints, and exploit the differential geometry of this space to solve optimization and inference problems. We demonstrate this approach by: (i) defining and computing statistics of observed shapes, (ii) defining and learning a parametric probability model on shape space, and (iii) designing a binary hypothesis test on this space.
Anuj Srivastava, Xiuwen Liu 0001, Washington Mio, Eric Klassen
NIPS2
2003 Statistical hypothesis pruning for identifying faces from infrared images
Anuj Srivastava, Xiuwen Liu 0001
Image Vis. Comput.2
2003 Intrinsic generalization analysis of low dimensional representations
Xiuwen Liu 0001, Anuj Srivastava, DeLiang Wang
Neural Networks1
2003 Texture classification using spectral histograms
abstract
Based on a local spatial/frequency representation,we employ a spectral histogram as a feature statistic for texture classification. The spectral histogram consists of marginal distributions of responses of a bank of filters and encodes implicitly the local structure of images through the filtering stage and the global appearance through the histogram stage. The distance between two spectral histograms is measured using chi(2)-statistic. The spectral histogram with the associated distance measure exhibits several properties that are necessary for texture classification. A filter selection algorithm is proposed to maximize classification performance of a given dataset. Our classification experiments using natural texture images reveal that the spectral histogram representation provides a robust feature statistic for textures and generalizes well. Comparisons show that our method produces a marked improvement in classification performance. Finally we point out the relationships between existing texture features and the spectral histogram, suggesting that the latter may provide a unified texture feature.
Xiuwen Liu 0001, DeLiang Wang
IEEE Trans. Image Process.1
2002 Analytical Image Models and Their Applications
Anuj Srivastava, Xiuwen Liu 0001, Ulf Grenander
ECCV (1)2
2002 Spaces and subspaces of images for recognition
abstract
In this paper we study and compare the recognition performance of subspaces in two different spaces, namely the image space and spectral histogram space. In image space, each image is represented as a long vector and in the spectral histogram space, each image is represented by its histograms of the convolved images with a chosen bank of filters. Spectral histogram space is a nonlinear transformation of the image space. First principal components and independent components in the spaces are studied. Then we study different subspaces by connecting the known subspaces through geodesic curves in the projection space. Our preliminary results show the recognition performance depends more on which space to use than the different subspaces in a given space. This suggests the need to study different spaces for recognition purpose.
Xiuwen Liu 0001, Anuj Srivastava
ICIP (3)1
2002 Universal Analytical Forms for Modeling Image Probabilities
abstract
Seeking probability models for images, we employ a spectral approach where the images are decomposed using bandpass filters and probability models are imposed on the filter outputs (also called spectral components). We employ a (two-parameter) family of probability densities, called Bessel K forms, for modeling the marginal densities of the spectral components, and demonstrate their fit to the observed histograms for video, infrared, and range images. Motivated by object-based models for image analysis, a relationship between the Bessel parameters and the imaged objects is established. Using L/sup 2/-metric on the set of Bessel K forms, we propose a pseudometric on the image space for quantifying image similarities/differences. Some applications, including clutter classification and pruning of hypotheses for target recognition, are presented.
Anuj Srivastava, Xiuwen Liu 0001, Ulf Grenander
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Learning in Gibbsian Fields: How Accurate and How Fast Can It Be?
abstract
Gibbsian fields or Markov random fields are widely used in Bayesian image analysis, but learning Gibbs models is computationally expensive. The computational complexity is pronounced by the recent minimax entropy (FRAME) models which use large neighborhoods and hundreds of parameters. In this paper, we present a common framework for learning Gibbs models. We identify two key factors that determine the accuracy and speed of learning Gibbs models: The efficiency of likelihood functions and the variance in approximating partition functions using Monte Carlo integration. We propose three new algorithms. In particular, we are interested in a maximum satellite likelihood estimator, which makes use of a set of precomputed Gibbs models called "satellites" to approximate likelihood functions. This algorithm can approximately estimate the minimax entropy model for textures in seconds in a HP workstation. The performances of various learning algorithms are compared in our experiments.
Song-Chun Zhu, Xiuwen Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Scene analysis by integrating primitive segmentation and associative memory
abstract
Scene analysis is a major aspect of perception and continues to challenge machine perception. This paper addresses the scene-analysis problem by integrating a primitive segmentation stage with a model of associative memory. The model is a multistage system that consists of an initial primitive segmentation stage, a multimodule associative memory, and a short-term memory (STM) layer. Primitive segmentation is performed by a locally excitatory globally inhibitory oscillator network (LEGION), which segments the input scene into multiple parts that correspond to groups of synchronous oscillations. Each segment triggers memory recall and multiple recalled patterns then interact with one another in the STM layer. The STM layer projects to the LEGION network, giving rise to memory-based grouping and segmentation. The system achieves scene analysis entirely in phase space, which provides a unifying mechanism for both bottom-up analysis and top-down analysis. The model is evaluated with a systematic set of three-dimensional (3-D) line drawing objects, which are arranged in an arbitrary fashion to compose input scenes that allow object occlusion. Memory-based organization is responsible for a significant improvement in performance. A number of issues are discussed, including input-anchored alignment, top-down organization, and the role of STM in producing context sensitivity of memory recall.
DeLiang Wang, Xiuwen Liu 0001
IEEE Trans. Syst. Man Cybern. Part B2
2001 Image segmentation using local spectral histograms
abstract
We propose a new algorithm for image segmentation. We use the spectral histogram, which is a vector consisting of marginal distributions of responses from chosen filters as a generic feature for texture as well as intensity images. Motivated by a new segmentation energy functional, we derive an iterative and deterministic approximation algorithm for segmentation. Based on the relationships between different scales and neighboring windows, we also develop an algorithm which can automatically detect homogeneous regions in an input image, which may consist of texture regions. To reduce the boundary uncertainty due to the large spatial window used for spectral histograms, we propose a novel local feature by building precise probability models based on current segmentation results. We have applied our algorithm to intensity, texture, and natural images and obtained good results with accurate texture boundaries.
Xiuwen Liu 0001, DeLiang Wang, Anuj Srivastava
ICIP (1)1
2001 Analytical models for reduced spectral representations of images
abstract
Spectral components, obtained via bandpass filtering of images, have become important tools in capturing image variability. We present a two-parameter family of probability densities, called the Bessel forms, to model the marginal densities of the spectral components. The two parameters, shape and scale parameters, are used to characterize each spectral component of the image. We derive an L/sup 2/-metric on the space of Bessel forms that leads to a metric on the space of natural images. The strength of these forms/metric is demonstrated via a study of natural clutter images.
Anuj Srivastava, Xiuwen Liu 0001, Ulf Grenander
ICIP (1)2
2001 Extraction of hydrographic regions from remote sensing images using an oscillator network with weight adaptation
abstract
The authors propose a framework for object extraction with accurate boundaries. A multilayer perceptron is used to identify seed points through examples, and regions are extracted and localized using a locally coupled network with weight adaptation. A functional system has been developed and applied to hydrographic region extraction from Digital Orthophoto Quarter-Quadrangle images.
Xiuwen Liu 0001, Ke Chen 0001, DeLiang Wang
IEEE Trans. Geosci. Remote. Sens.1
2000 Learning in Gibbsian Fields: How Accurate and How Fast Can It Be?
abstract
In this article, we present a unified framework for learning Gibbs models from training images. We identify two key factors that determine the accuracy and speed of learning Gibbs models: (1). Fisher information, and (2). The accuracy of Monte Carlo estimate for partition functions. We propose three new learning algorithms under the unified framework. (I). The maximum partial likelihood estimator. (II). The maximum patch likelihood estimator, and (III). The maximum satellite likelihood estimator. The first two algorithms can speed up the minimax entropy algorithm by about 2D times without losing much accuracy. The third one makes use of a set of known Gibbs models as references-dubbed "satellites" and can approximately estimate the minimax entropy model in the speed of 10 seconds.
Song-Chun Zhu, Xiuwen Liu 0001
CVPR2
2000 Equivalence of Julesz Ensembles and FRAME Models
Ying Nian Wu, Song-Chun Zhu, Xiuwen Liu 0001
Int. J. Comput. Vis.3
2000 Exploring Texture Ensembles by Efficient Markov Chain Monte Carlo-Toward a 'Trichromacy' Theory of Texture
abstract
Presents a mathematical definition of texture, the Julesz ensemble /spl Omega/(h), which is the set of all images (defined on Z/sup 2/) that share identical statistics h. Then texture modeling is posed as an inverse problem: Given a set of images sampled from an unknown /spl Omega/(h/sub */), we search for the statistics h/sub */ which define the ensemble. A /spl Omega/(h) has an associated probability distribution q(I; h), which is uniform over the images in /spl Omega/(h) and has zero probability outside. The authors previously (1999) showed q(I; h) to be the limit distribution of the FRAME (filter, random field, and minimax entropy) model, as the image lattice /spl Lambda//spl rarr/Z/sup 2/. This conclusion establishes the intrinsic link between the scientific definition of texture on Z/sup 2/ and the mathematical models of texture on finite lattices. It brings two advantages: the practice of texture image synthesis by matching statistics is put on a mathematical foundation; and we need not learn the expensive FRAME model in feature pursuit, model selection and texture synthesis. An efficient Markov chain Monte Carte algorithm is proposed for sampling Julesz ensembles. It generates random texture images by moving along the directions of filter coefficients and, thus, extends the traditional single site Gibbs sampler. We compare four popular statistical measures in the literature, in terms of their descriptive abilities. Our experiments suggest that a small number of bins in marginal histograms are sufficient for capturing a variety of texture patterns.
Song-Chun Zhu, Xiuwen Liu 0001, Ying Nian Wu
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Boundary detection by contextual non-linear smoothing
Xiuwen Liu 0001, DeLiang Wang, J. Raul Ramirez
Pattern Recognit.1
2000 Weight adaptation and oscillatory correlation for image segmentation
abstract
We propose a method for image segmentation based on a neural oscillator network. Unlike previous methods, weight adaptation is adopted during segmentation to remove noise and preserve significant discontinuities in an image. Moreover, a logarithmic grouping rule is proposed to facilitate grouping of oscillators representing pixels with coherent properties. We show that weight adaptation plays the roles of noise removal and feature preservation. In particular, our weight adaptation scheme is insensitive to termination time and the resulting dynamic weights in a wide range of iterations lead to the same segmentation results. A computer algorithm derived from oscillatory dynamics is applied to synthetic and real images and simulation results show that the algorithm yields favorable segmentation results in comparison with other recent algorithms. In addition, the weight adaptation scheme can be directly transformed to a novel feature-preserving smoothing procedure. We also demonstrate that our nonlinear smoothing algorithm achieves good results for various kinds of images.
Ke Chen 0001, DeLiang Wang, Xiuwen Liu 0001
IEEE Trans. Neural Networks Learn. Syst.3
1999 Equivalence of Julesz and Gibbs Texture Ensembles
abstract
Research on texture has been pursued along two different lines. The first line of research, pioneered by Julesz (1962), seeks the essential ingredients in terms of features and statistics in human texture perception. This leads us to a mathematical definition of texture as a Julesz ensemble. A Julesz ensemble is the maximum set of images that share the same value of some basic feature statistics as the image lattice /spl Lambda//spl rarr/Z/sup 2/, or equivalently it is a uniform distribution on this set. The second line of research studies statistical models, in particular, Markov random field (MRF) and FRAME models (Zhu et al., 1997), to characterize texture patterns locally. In this article, we bridge the two lines by the fundamental principle of equivalence of ensembles in statistical mechanics (Gibbs, 1902). We prove that 1) the conditional probability of an arbitrary image patch given its environment, under the Julesz ensemble or the uniform model, is inevitably a FRAME (MRF) model, and 2) the limit of the FRAME (MRF) model, which we called the Gibbs ensemble, is equivalent to a Julesz ensemble as /spl Lambda//spl rarr/Z/sup 2/. Thus the advantages of the two methodologies can be fully utilized.
Ying Nian Wu, Song-Chun Zhu, Xiuwen Liu 0001
ICCV3
1999 A boundary-pair representation for perception modeling
abstract
It is widely accepted that responses from on- and off-center cells give rise to edges and are equivalent to edge detectors. In this paper, we point out that on- and off-center cell responses provide more information than edges. We show that an edge-based representation makes the ownership of boundaries ambiguous and requires a combinatorial search to model perceptual grouping. By analyzing the differences between edges and responses from on- and off-center cells, we propose a boundary-pair representation, which makes the ownership of boundaries explicit and eliminates the need of a combinatorial search computationally. Each boundary in the boundary-pair representation is associated with regional attributes. We show that this representation is equivalent to a surface representation through a local diffusion. This provides a unified representation for perception modeling. Based on this representation, a figure-ground segregation network is constructed to demonstrate the capabilities of the model in explaining many perceptual phenomena.
Xiuwen Liu 0001, DeLiang Wang
IJCNN1
1999 Perceptual organization based on temporal dynamics
abstract
This paper presents a computational model for perceptual organization. A figure-ground segregation network is proposed based on a novel boundary pair representation. The system solves the figure-ground segregation problem through temporal evolution. Gestalt-like grouping rules are incorporated by modulating connections, which determines the temporal behavior and thus the perception of the system. The results are then fed to a surface completion module based on local diffusion. Different perceptual phenomena, such as modal and a modal completion, virtual contours, grouping and shape decomposition are explained by the model with a fixed set of parameters. Computationally, the system eliminates combinatorial optimization, which is common to many existing computational approaches. It also accounts for more examples that are consistent with psychological experiments. In addition, the boundary-pair representation is consistent with well-known on- and off-center cell responses and thus biologically more plausible.
Xiuwen Liu 0001, DeLiang Wang
IJCNN1
1999 Perceptual Organization Based on Temporal Dynamics
Xiuwen Liu 0001, DeLiang Wang
NIPS1
1999 Range image segmentation using a relaxation oscillator network
abstract
A locally excitatory globally inhibitory oscillator network (LEGION) is constructed and applied to range image segmentation, where each oscillator has excitatory lateral connections to the oscillators in its local neighborhood as well as a connection with a global inhibitor. A feature vector, consisting of depth, surface normal, and mean and Gaussian curvatures, is associated with each oscillator and is estimated from local windows at its corresponding pixel location. A context-sensitive method is applied in order to obtain more reliable and accurate estimations. The lateral connection between two oscillators is established based on a similarity measure of their feature vectors. The emergent behavior of the LEGION network gives rise to segmentation. Due to the flexible representation through phases, our method needs no assumption about the underlying structures in image data and no prior knowledge regarding the number of regions. More importantly, the network is guaranteed to converge rapidly under general conditions. These unique properties may lead to a real-time approach for range image segmentation in machine perception.
Xiuwen Liu 0001, DeLiang Wang
IEEE Trans. Neural Networks1
1998 Oriented Statistical Nonlinear Smoothing Filter
abstract
This paper presents a nonlinear smoothing method which is based on an orientation-sensitive probability measure. By incorporating geometrical constraints through the coupling structure, we obtain a robust nonlinear smoothing algorithm. Even when noise is substantial the proposed smoothing algorithm can still preserve salient boundaries. Compared with anisotropic diffusive approaches, the proposed nonlinear algorithm not only performs better in preserving boundaries but also has a non-uniform stable state, whereby reliable results are available within a fixed number of iterations independent of images. A system using the proposed method and LEGION network has been developed and applied in noisy image segmentation and hydrographic feature extraction from digital ortho-photo quadrangles. Experimental results using synthetic and real images are provided.
Xiuwen Liu 0001, DeLiang Wang, J. Raul Ramirez
ICIP (2)1
1997 Visualization of plant growth
abstract
The measurement, analysis and visualization of plant growth is of primary interest to plant biologists. We are developing software tools to support such investigations. There are two parts in this investigation, namely growth visualization of (i) a plant root and (ii) a plant stem. For both domains, the input data is a stream of images taken by cameras. The tools being developed make it possible to measure various time-varying quantities, such as differential growth. For both domains, the plant is modeled by using flexible templates to represent non-rigid motions.
Jeremy J. Loomis, Xiuwen Liu 0001, Zhaohua Ding, Kikuo Fujimura, Michael L. Evans, Hideo Ishikawa
IEEE Visualization2