EDBT 2026 Demo / reviewers in the wild / expert
Yalin Wang 0001
dblp:88/128-1
· DBLP profile ↗
90ranked-venue papers
23as first author
33since 2021 · last 2026
0000-0002-6241-735XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 13 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 39 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 36 · 10 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model AdaptationabstractVision-language models (VLMs) such as CLIP demonstrate strong performance but struggle when adapted to downstream tasks. Prompt learning has emerged as an efficient and effective strategy to adapt VLMs while preserving their pre-trained knowledge. However, existing methods still lead to overfitting and degrade zero-shot generalization. To address this challenge, we propose an optimal transport (OT)-guided prompt learning framework that mitigates forgetting by preserving the structural consistency of feature distributions between pre-trained and fine-tuned models. Unlike conventional point-wise constraints, OT naturally captures cross-instance relationships and expands the feasible parameter space for prompt tuning, allowing a better tradeoff between adaptation and generalization. Our approach enforces joint constraints on both vision and text representations, ensuring a holistic feature alignment. Extensive experiments on benchmark datasets demonstrate that our simple yet effective method outperforms existing prompt learning strategies in base-to-novel generalization, cross-dataset evaluation, and domain generalization, without requiring additional augmentation or ensemble techniques. Xiwen Chen, Peijie Qiu, Hao Wang 0176, Haiyu Wu, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi |
WACV | 8 |
| 2026 | EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language ModelsabstractInstructed Visual Segmentation (IVS) tasks require segmenting objects in images or videos based on natural language instructions. While recent multimodal large language models (MLLMs) have achieved strong performance on IVS, their inference cost remains a major bottleneck, particularly in video. We empirically analyze visual token sampling in MLLMs and observe a strong correlation between subset token coverage and segmentation performance. This motivates our design of a simple and effective token pruning method that selects a compact yet spatially representative subset of tokens to accelerate inference. In this paper, we introduce a novel visual token pruning method for IVS, called EVTP-IV, which builds upon the k-center by integrating spatial information to ensure better coverage. We further provide an information-theoretic analysis to support our design. Experiments on standard IVS benchmarks show that our method achieves up to 5X speed-up on video tasks and 3.5X on image tasks, while maintaining comparable accuracy using only 20% of the tokens. Our method also consistently outperforms state-of-the-art pruning baselines under varying pruning ratios. Xiwen Chen, Shao Tang, Xuanzhao Dong, Rajat Koner, Yalin Wang 0001 |
WACV | 8 |
| 2025 | Sequence Complementor: Complementing Transformers for Time Series Forecasting with Learnable SequencesabstractSince its introduction, the transformer has shifted the development trajectory away from traditional models (e.g., RNN, MLP) in time series forecasting, which is attributed to its ability to capture global dependencies within temporal tokens. Follow-up studies have largely involved altering the tokenization and self-attention modules to better adapt Transformers for addressing special challenges like non-stationarity, channel-wise dependency, and variable correlation in time series. However, we found that the expressive capability of sequence representation is a key factor influencing Transformer performance in time forecasting after investigating several representative methods, where there is an almost linear relationship between sequence representation entropy and mean square error, with more diverse representations performing better. In this paper, we propose a novel attention mechanism with Sequence Complementors and prove feasible from an information theory perspective, where these learnable sequences are able to provide complementary information beyond current input to feed attention. We further enhance the Sequence Complementors via a diversification loss that is theoretically covered. The empirical evaluation of both long-term and short-term forecasting has confirmed its superiority over the recent state-of-the-art methods. Xiwen Chen, Peijie Qiu, Hao Wang 0176, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi |
AAAI | 7 |
| 2025 | Multimodal Variational Autoencoder: A Barycentric ViewabstractMultiple signal modalities, such as vision and sounds, are naturally present in real-world phenomena. Recently, there has been growing interest in learning generative models, in particular variational autoencoder (VAE), for multimodal representation learning especially in the case of missing modalities. The primary goal of these models is to learn a modality-invariant and modality-specific representation that characterizes information across multiple modalities. Previous attempts at multimodal VAEs approach this mainly through the lens of experts, aggregating unimodal inference distributions with a product of experts (PoE), a mixture of experts (MoE), or a combination of both. In this paper, we provide an alternative generic and theoretical formulation of multimodal VAE through the lens of barycenter. We first show that PoE and MoE are specific instances of barycenters, derived by minimizing the asymmetric weighted KL divergence to unimodal inference distributions. Our novel formulation extends these two barycenters to a more flexible choice by considering different types of divergences. In particular, we explore the Wasserstein barycenter defined by the 2-Wasserstein distance, which better preserves the geometry of unimodal distributions by capturing both modality-specific and modality-invariant representations compared to KL divergence. Empirical studies on three multimodal benchmarks demonstrated the effectiveness of the proposed method. Peijie Qiu, Sayantan Kumar, Xiwen Chen, Abolfazl Razi, Yalin Wang 0001, Aristeidis Sotiras |
AAAI | 8 |
| 2025 | Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image AnalysisabstractWhile multiple instance learning (MIL) has shown to be a promising approach for histopathological whole slide image (WSI) analysis, its reliance on permutation invariance significantly limits its capacity to effectively uncover semantic correlations between instances within WSIs. Based on our empirical and theoretical investigations, we argue that approaches that are not permutation-invariant but better capture spatial correlations between instances can offer more effective solutions. In light of these findings, we propose a novel alternative to existing MIL for WSI analysis by learning to restore the order of instances from their randomly shuffled arrangement. We term this task as cracking an instance jigsaw puzzle problem, where semantic correlations between instances are uncovered. To tackle the instance jigsaw puzzles, we propose a novel Siamese network solution, which is theoretically justified by optimal transport theory. We validate the proposed method on WSI classification and survival prediction tasks, where the proposed method outperforms the recent state-of-the-art MIL competitors. The code is available at https://github.com/xiwenc1/MIL-JigsawPuzzles. Xiwen Chen, Peijie Qiu, Hao Wang 0176, Xuanzhao Dong, Yalin Wang 0001, Abolfazl Razi, Aristeidis Sotiras |
ICCV | 9 |
| 2025 | FIC-TSC: Learning Time Series Classification with Fisher Information ConstraintabstractAnalyzing time series data is crucial to a wide spectrum of applications, including economics, online marketplaces, and human healthcare. In particular, time series classification plays an indispensable role in segmenting different phases in stock markets, predicting customer behavior, and classifying worker actions and engagement levels. These aspects contribute significantly to the advancement of automated decision-making and system optimization in real-world applications. However, there is a large consensus that time series data often suffers from domain shifts between training and test sets, which dramatically degrades the classification performance. Despite the success of (reversible) instance normalization in handling the domain shifts for time series regression tasks, its performance in classification is unsatisfactory. In this paper, we propose $\textit{FIC-TSC}$, a training framework for time series classification that leverages Fisher information as the constraint. We theoretically and empirically show this is an efficient and effective solution to guide the model converges toward flatter minima, which enhances its generalizability to distribution shifts. We rigorously evaluate our method on 30 UEA multivariate and 85 UCR univariate datasets. Our empirical results demonstrate the superiority of the proposed method over 14 recent state-of-the-art methods. Xiwen Chen, Peijie Qiu, Hao Wang 0176, Yalin Wang 0001, Aristeidis Sotiras, Abolfazl Razi |
ICML | 7 |
| 2025 | How Effective Can Dropout Be in Multiple Instance Learning ?abstractMultiple Instance Learning (MIL) is a popular weakly-supervised method for various applications, with a particular interest in histological whole slide image (WSI) classification. Due to the gigapixel resolution of WSI, applications of MIL in WSI typically necessitate a two-stage training scheme: first, extract features from the pre-trained backbone and then perform MIL aggregation. However, it is well-known that this suboptimal training scheme suffers from "noisy" feature embeddings from the backbone and inherent weak supervision, hindering MIL from learning rich and generalizable features. However, the most commonly used technique (i.e., dropout) for mitigating this issue has yet to be explored in MIL. In this paper, we empirically explore how effective the dropout can be in MIL. Interestingly, we observe that dropping the top-k most important instances within a bag leads to better performance and generalization even under noise attack. Based on this key observation, we propose a novel MIL-specific dropout method, termed MIL-Dropout, which systematically determines which instances to drop. Experiments on five MIL benchmark datasets and two WSI datasets demonstrate that MIL-Dropout boosts the performance of current MIL methods with a negligible computational cost. The code is available at https://github.com/ChongQingNoSubway/MILDropout. Peijie Qiu, Xiwen Chen, Zhangsihao Yang, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001 |
ICML | 7 |
| 2025 | CUNSB-RFIE: Context-Aware Unpaired Neural Schrödinger Bridge in Retinal Fundus Image EnhancementabstractRetinal fundus photography is significant in diagnosing and monitoring retinal diseases. However, systemic imperfections and operator/patient-related factors can hinder the acquisition of high-quality retinal images. Previous efforts in retinal image enhancement primarily relied on GANs, which are limited by the trade-off between training stability and output diversity. In contrast, the Schrödinger Bridge (SB), offers a more stable solution by utilizing Optimal Transport (OT) theory to model a stochastic differential equation (SDE) between two arbitrary distributions. This allows SB to effectively transform low-quality retinal images into their high-quality counterparts. In this work, we leverage the SB framework to propose an image-to-image translation pipeline for retinal image enhancement. Additionally, previous methods often fail to capture fine struc tural details, such as blood vessels. To address this, we enhance our pipeline by introducing Dynamic Snake Convolution, whose tortuous receptive field can better preserve tubular structures. We name the resulting retinal fundus image enhancement framework the Context-aware Unpaired Neural Schrödinger Bridge (CUNSB-RFIE). To the best of our knowledge, this is the first endeavor to use the SB approach for retinal image enhancement. Experimental results on a large-scale dataset demonstrate the advantage of the proposed method compared to several state-of-the-art supervised and unsupervised methods in terms of image quality and performance on downstream tasks.The code is available at https://github.com/Retinal-Research/CUNSB-RFIE. Xuanzhao Dong, Vamsi Krishna Vasa, Peijie Qiu, Xiwen Chen, Yi Su 0004, Yujian Xiong, Zhangsihao Yang, Yanxi Chen 0002, Yalin Wang 0001 |
WACV | 10 |
| 2025 | A Recipe for Geometry-Aware 3D Mesh TransformersabstractUtilizing patch-based transformers for unstructured geometric data such as polygon meshes presents significant challenges, primarily due to the absence of a canonical ordering and variations in input sizes. Prior approaches to handling 3D meshes and point clouds have either relied on computationally intensive node-level tokens for large objects or resorted to resampling to standardize patch size. Moreover, these methods generally lack a geometry-aware, stable Structural Embedding (SE), often depending on simplistic absolute SEs such as 3D coordinates, which compromise isometry invariance essential for tasks like semantic segmentation. In our study, we meticulously examine the various components of a geometry-aware 3D mesh transformer, from tokenization to structural encoding, assessing the contribution of each. Initially, we introduce a spectral-preserving tokenization rooted in algebraic multi-grid methods. Subsequently, we detail an approach for embedding features at the patch level, accommodating patches with variable node counts. Through comparative analyses against a baseline model employing simple point-wise Multi-Layer Perceptrons (MLP), our research highlights critical insights: 1) the importance of structural and positional embeddings facilitated by heat diffusion in general 3D mesh transformers; 2) the effectiveness of novel components such as geodesic masking and feature interaction via cross-attention in enhancing learning; and 3) the superior performance and efficiency of our proposed methods in challenging segmentation and classification tasks. Mohammad Farazi, Yalin Wang 0001 |
WACV | 2 |
| 2025 | Context-Aware Optimal Transport Learning for Retinal Fundus Image EnhancementabstractRetinal fundus photography offers a non-invasive way to diagnose and monitor a variety of retinal diseases, but is prone to inherent quality glitches arising from systemic imperfections or operator/patient-related factors. However, high-quality retinal images are crucial for carrying out accurate diagnoses and automated analyses. The fundus image enhancement is typically formulated as a distribution alignment problem, by finding a one-to-one mapping between a low-quality image and its high-quality counterpart. This paper proposes a context-informed optimal transport (OT) learning framework for tackling unpaired fundus image enhancement. In contrast to standard generative image enhancement methods, which struggle with handling contextual information (e.g., over-tampered local structures and unwanted artifacts), the proposed context-aware OT learning paradigm better preserves local structures and minimizes unwanted artifacts. Leveraging deep contextual features, we derive the proposed context-aware OT using the earth mover's distance and show that the proposed context-OT has a solid theoretical guarantee. Experimental results on a large-scale dataset demonstrate the superiority of the proposed method over several state-of-the-art supervised and unsupervised methods in terms of signal-to-noise ratio, structural similarity index, as well as two downstream tasks. The code is available at https://github.com/Retinal-Research/Contextual-OT. Vamsi Krishna Vasa, Peijie Qiu, Yujian Xiong, Oana M. Dumitrascu, Yalin Wang 0001 |
WACV | 6 |
| 2025 | Schizophrenia Detection Based on Morphometry of Hippocampus and AmygdalaabstractSchizophrenia (SZ) is a severe mental disorder characterized by hallucinations, delusions, cognitive impairments, and social withdrawal. It leads to a series of brain abnormalities, particularly the deformation of the hippocampus and amygdala, which are highly associated with emotion, memory, and motivation. Most previous studies have used the hippocampal and amygdaloid volume, whereas surface-based morphometry reflects nuclear deformation more finely, but it is unclear the hippocampal and amygdaloid morphometry relates to schizophrenic pathology and its potential as a biomarker. In this study, we extracted individual multivariate morphometry statistics (MMS) of hippocampus and amygdala from MRI images and analyzed the morphometric differences between groups. After dictionary learning and max pooling, we obtain reduced dimensional features and use machine learning algorithms for individual diagnosis. The results showed that the hippocampus of the schizophrenia group was significantly atrophied bilaterally and the atrophied areas were symmetrical. Subregions of the amygdala are both atrophied and expanded, and in particular, the right amygdala shows a greater degree and extent of deformation. Using the random forest classifier, the accuracy of classification using hippocampal and amygdaloid morphometric features are 94.52% and 94.57%, respectively, and the accuracy of classification combining the two morphometric features reached 96.57%. Our study demonstrates the efficacy of MMS in identifying morphometric differences of the hippocampus and amygdala between healthy controls and schizophrenic, and these findings emphasize the potential of MMS as a reliable biomarker for the diagnosis of schizophrenia. Qunxi Dong, Yuhang Sheng, Junru Zhu, Jingyu Liu 0002, Yalin Wang 0001, Bin Hu 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | OmniMotionGPT: Animal Motion Generation with Limited DataabstractOur paper aims to generate diverse and realistic ani-mal motion sequences from textual descriptions, without a large-scale animal text-motion dataset. While the task of text-driven human motion synthesis is already extensively studied and benchmarked, it remains challenging to transfer this success to other skeleton structures with limited data. In this work, we design a model architecture that imitates Generative Pretraining Transformer (GPT), utilizing prior knowledge learned from human data to the animal domain. We jointly train motion autoencoders for both animal and human motions and at the same time optimize through the similarity scores among human motion encoding, animal motion encoding, and text CLIP embedding. Presenting the first solution to this problem, we are able to generate animal motions with high diversity and fidelity, quantitatively and qualitatively outperforming the results of training human motion generation baselines on animal data. Additionally, we introduce AnimalML3D, the first text-animal motion dataset with 1240 animation sequences spanning 36 different ani-mal identities. We hope this dataset would mediate the data scarcity problem in text-driven animal motion generation, providing a new playground for the research community. Zhangsihao Yang, Mingyuan Zhou, Mengyi Shan, Bingbing Wen, Ziwei Xuan, Mitch Hill, Guo-Jun Qi, Yalin Wang 0001 |
CVPR | 9 |
| 2024 | DGR-MIL: Exploring Diverse Global Representation in Multiple Instance Learning for Whole Slide Image Classification
Xiwen Chen, Peijie Qiu, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001 |
ECCV (38) | 6 |
| 2024 | TimeMIL: Advancing Multivariate Time Series Classification via a Time-aware Multiple Instance LearningabstractDeep neural networks, including transformers and convolutional neural networks (CNNs), have significantly improved multivariate time series classification (MTSC). However, these methods often rely on supervised learning, which does not fully account for the sparsity and locality of patterns in time series data (e.g., quantification of diseases-related anomalous points in ECG and abnormal detection in signal). To address this challenge, we formally discuss and reformulate MTSC as a weakly supervised problem, introducing a novel multiple-instance learning (MIL) framework for better localization of patterns of interest and modeling time dependencies within time series. Our novel approach, TimeMIL, formulates the temporal correlation and ordering within a time-aware MIL pooling, leveraging a tokenized transformer with a specialized learnable wavelet positional token. The proposed method surpassed 26 recent state-of-the-art MTSC methods, underscoring the effectiveness of the weakly supervised TimeMIL in MTSC. The code is available https://github.com/xiwenc1/TimeMIL. Xiwen Chen, Peijie Qiu, Hao Wang 0176, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi |
ICML | 7 |
| 2024 | SelfReg-UNet: Self-Regularized UNet for Medical Image Segmentation
Xiwen Chen, Peijie Qiu, Mohammad Farazi, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001 |
MICCAI (8) | 7 |
| 2024 | MGM-AE: Self-Supervised Learning on 3D Shape Using Mesh Graph Masked AutoencodersabstractThe challenges of applying self-supervised learning to 3D mesh data include difficulties in explicitly modeling and leveraging geometric topology information and designing appropriate pretext tasks and augmentation methods for irregular mesh topology. In this paper, we propose a novel approach for pre-training models on large-scale, unlabeled datasets using graph masking on a mesh graph composed of faces. Our method, Mesh Graph Masked Autoencoders (MGM-AE), utilizes masked autoencoding to pre-train the model and extract important features from the data. Our pre-trained model outperforms prior state-of-the-art mesh encoders in shape classification and segmentation benchmarks, achieving 90.8% accuracy on ModelNet40 and 78.5 mIoU on ShapeNet. The best performance is obtained when the model is trained and evaluated under different masking ratios. Our approach demonstrates effectiveness in pretraining models on large-scale, unlabeled datasets and its potential for improving performance on downstream tasks. Zhangsihao Yang, Kaize Ding, Huan Liu 0001, Yalin Wang 0001 |
WACV | 4 |
| 2024 | Adaptive large neighborhood search algorithm with reinforcement search strategy for solving extended cooperative multi task assignment problem of UAVs
Yougang Xiao, Huan Liu 0001, Yingguo Chen, Yalin Wang 0001, Guohua Wu 0001 |
Inf. Sci. | 5 |
| 2024 | Adaptive smoothing of retinotopic maps based on Teichmüller parametrization
Yanshuai Tu, Xin Li 0201, Zhonglin Lu, Yalin Wang 0001 |
Medical Image Anal. | 4 |
| 2023 | Bidirectional Mapping with Contrastive Learning on Multimodal Neuroimaging Data
Kai Ye 0002, Haoteng Tang, Siyuan Dai, Lei Guo 0028, Johnny Yuehan Liu, Yalin Wang 0001, Alex D. Leow, Paul M. Thompson, Heng Huang 0001, Liang Zhan |
MICCAI (3) | 6 |
| 2023 | Keypoint-Augmented Self-Supervised Learning for Medical Image Segmentation with Limited AnnotationabstractPretraining CNN models (i.e., UNet) through self-supervision has become a powerful approach to facilitate medical image segmentation under low annotation regimes. Recent contrastive learning methods encourage similar global representations when the same image undergoes different transformations, or enforce invariance across different image/patch features that are intrinsically correlated. However, CNN-extracted global and local features are limited in capturing long-range spatial dependencies that are essential in biological anatomy. To this end, we present a keypoint-augmented fusion layer that extracts representations preserving both short- and long-range self-attention. In particular, we augment the CNN feature map at multiple scales by incorporating an additional input that learns long-range spatial self-attention among localized keypoint features. Further, we introduce both global and local self-supervised pretraining for the framework. At the global scale, we obtain global representations from both the bottleneck of the UNet, and by aggregating multiscale keypoint features. These global features are subsequently regularized through image-level contrastive objectives. At the local scale, we define a distance-based criterion to first establish correspondences among keypoints and encourage similarity between their features. Through extensive experiments on both MRI and CT segmentation tasks, we demonstrate the architectural advantages of our proposed method in comparison to both CNN and Transformer-based UNets, when all architectures are trained with randomly initialized weights. With our proposed pretraining strategy, our method further outperforms existing SSL methods by producing more robust self-attention and achieving state-of-the-art segmentation results. The code is available at https://github.com/zshyang/kaf.git. Zhangsihao Yang, Mengwei Ren, Kaize Ding, Guido Gerig, Yalin Wang 0001 |
NeurIPS | 5 |
| 2023 | Anisotropic Multi-Scale Graph Convolutional Network for Dense Shape CorrespondenceabstractThis paper studies 3D dense shape correspondence, a key shape analysis application in computer vision and graphics. We introduce a novel hybrid geometric deep learning-based model that learns geometrically meaningful and discretization-independent features. The proposed framework has a U-Net model as the primary node feature extractor, followed by a successive spectral-based graph convolutional network. To create a diverse set of filters, we use anisotropic wavelet basis filters, being sensitive to both different directions and band-passes. This filter set overcomes the common over-smoothing behavior of conventional graph neural networks. To further improve the model's performance, we add a function that perturbs the feature maps in the last layer ahead of fully connected layers, forcing the network to learn more discriminative features overall. The resulting correspondence maps show state-of-the-art performance on the benchmark datasets based on average geodesic errors and superior robustness to discretization in 3D meshes. Our approach provides new insights and practical solutions to the dense shape correspondence research. Mohammad Farazi, Zhangsihao Yang, Yalin Wang 0001 |
WACV | 4 |
| 2023 | Signed graph representation learning for functional-to-structural brain network mapping
Haoteng Tang, Lei Guo 0028, Xiyao Fu, Yalin Wang 0001, Scott Mackin, Olusola Ajilore, Alex D. Leow, Paul M. Thompson, Heng Huang 0001, Liang Zhan |
Medical Image Anal. | 4 |
| 2023 | Correlation Studies of Hippocampal Morphometry and Plasma NFL Levels in Cognitively Unimpaired SubjectsabstractAlzheimer's disease(AD) is being the burden of society and family. Applying computing-aided strategies to reveal its pathology is one of the research highlights. Plasma neurofilament light (NFL) is an emerging noninvasive and economic biomarker for AD molecular pathology. It is valuable to reveal the correlations between the plasma NFL levels and neurodegeneration, especially hippcampal deformations at the preclinical stage. The negative correlation between plasma NFL levels and hippocampal volumes has been documented. However, the relationship between the plasma NFL levels and the hippocampal morphometry details at the preclinical stage is still elusive. This study seeks to demonstrate the capacity of our proposed surface-based hippocampal morphometry system to discern the plasma NFL positive (NFL+>41.9 pg/L) level and plasma NFL negative (NFL-<41.9pg/L) level and illustrate its superiority to the hippocampal volume measurement by drawing the cohort of 154 CU middle aged and elderly adults. We also apply this morphometry measure and a proposed sparse coding based classification algorithm to classify CU individuals with NFL+ and NFL- levels. Experimental results show that the proposed hippocampal morphometry system offers stronger statistical power to discriminate CU subjects with NFL+ and NFL- levels, comparing with the hippocampal volume measure. Furthermore, this system can discriminate plasma NFL levels in CU individuals (Accuracy=0.86). Both the group level and individual level analysis results indicate that the association between plasma NFL levels and the hippocampal shapes can be mapped at the preclinical stage. Qunxi Dong, Kewei Chen 0001, Yi Su 0004, Richard J. Caselli, Eric Reiman, Yalin Wang 0001, Jian Shen 0004 |
IEEE Trans. Comput. Soc. Syst. | 9 |
| 2022 | Geometry-Aware Hierarchical Bayesian Learning on ManifoldsabstractBayesian learning with Gaussian processes demonstrates encouraging regression and classification performances in solving computer vision tasks. However, Bayesian methods on 3D manifold-valued vision data, such as meshes and point clouds, are seldom studied. One of the primary challenges is how to effectively and efficiently aggregate geometric features from the irregular inputs. In this paper, we propose a hierarchical Bayesian learning model to address this challenge. We initially introduce a kernel with the properties of geometry-awareness and intra-kernel convolution. This enables geometrically reasonable inferences on manifolds without using any specific hand-crafted feature descriptors. Then, we use a Gaussian process regression to organize the inputs and finally implement a hierarchical Bayesian network for the feature aggregation. Furthermore, we incorporate the feature learning of neural networks with the feature aggregation of Bayesian models to investigate the feasibility of jointly learning on manifolds. Experimental results not only show that our method outperforms existing Bayesian methods on manifolds but also demonstrate the prospect of coupling neural networks with Bayesian networks. Yonghui Fan, Yalin Wang 0001 |
WACV | 2 |
| 2022 | Quantitative characterization of the human retinotopic map based on quasiconformal mappingabstractThe retinotopic map depicts the cortical neurons' response to visual stimuli on the retina and has contributed significantly to our understanding of human visual system. Although recent advances in high field functional magnetic resonance imaging (fMRI) have made it possible to generate the in vivo retinotopic map with great detail, quantifying the map remains challenging. Existing quantification methods do not preserve surface topology and often introduce large geometric distortions to the map. In this study, we developed a new framework based on computational conformal geometry and quasiconformal Teichmüller theory to quantify the retinotopic map. Specifically, we introduced a general pipeline, consisting of cortical surface conformal parameterization, surface-spline-based cortical activation signal smoothing, and vertex-wise Beltrami coefficient-based map description. After correcting most of the violations of the topological conditions, the result was a "Beltrami coefficient map" (BCM) that rigorously and completely characterizes the retinotopic map by quantifying the local quasiconformal mapping distortion at each visual field location. The BCM provided topological and fully reconstructable retinotopic maps. We successfully applied the new framework to analyze the V1 retinotopic maps from the Human Connectome Project (n=181), the largest state of the art retinotopy dataset currently available. With unprecedented precision, we found that the V1 retinotopic map was quasiconformal and the local mapping distortions were similar across observers. The new framework can be applied to other visual areas and retinotopic maps of individuals with and without eye diseases, and improve our understanding of visual cortical organization in normal and clinical populations. Duyan Ta, Yanshuai Tu, Zhonglin Lu, Yalin Wang 0001 |
Medical Image Anal. | 4 |
| 2022 | Corrigendum to 'Quantitative Characterization of the Human Retinotopic Map Based on Quasiconformal Mapping' [Medical Image Analysis Volume 75 (2022) 102230]
Duyan Ta, Yanshuai Tu, Zhonglin Lu, Yalin Wang 0001 |
Medical Image Anal. | 4 |
| 2021 | Cortical Surface Shape Analysis Based on Alexandrov PolyhedraabstractShape analysis has been playing an important role in early diagnosis and prognosis of neurodegenerative diseases such as Alzheimer's diseases (AD). However, obtaining effective shape representations remains challenging. This paper proposes to use the Alexandrov polyhedra as surface-based shape signatures for cortical morphometry analysis. Given a closed genus-0 surface, its Alexandrov polyhedron is a convex representation that encodes its intrinsic geometry information. We propose to compute the polyhedra via a novel spherical optimal transport (OT) computation. In our experiments, we observe that the Alexandrov polyhedra of cortical surfaces between pathology-confirmed AD and cognitively unimpaired individuals are significantly different. Moreover, we propose a visualization method by comparing local geometry differences across cortical surfaces. We show that the proposed method is effective in pinpointing regional cortical structural changes impacted by AD. Min Zhang 0069, Na Lei, Xiaoyin Xu, Yalin Wang 0001, Xianfeng Gu |
ICCV | 7 |
| 2021 | Topological Receptive Field Model for Human Retinotopic Mapping
Yanshuai Tu, Duyan Ta, Zhonglin Lu, Yalin Wang 0001 |
MICCAI (7) | 4 |
| 2021 | Predicting future cognitive decline with hyperbolic stochastic coding
Jie Zhang 0026, Qunxi Dong, Jie Shi 0001, Qingyang Li 0001, Cynthia M. Stonnington, Boris Gutman, Kewei Chen 0001, Eric Reiman, Richard J. Caselli, Paul M. Thompson, Jieping Ye, Yalin Wang 0001 |
Medical Image Anal. | 12 |
| 2021 | Tetrahedral spectral feature-Based bayesian manifold learning for grey matter morphometry: Findings from the Alzheimer's disease neuroimaging initiative
Yonghui Fan, Gang Wang 0029, Qunxi Dong, Natasha Leporé, Yalin Wang 0001 |
Medical Image Anal. | 6 |
| 2021 | Developing univariate neurodegeneration biomarkers with low-rank and sparse subspace decomposition
Gang Wang 0029, Qunxi Dong, Yi Su 0004, Kewei Chen 0001, Qingtang Su, Xiaofeng Zhang 0003, Jinguang Hao, Li Liu 0035, Caiming Zhang 0001, Richard J. Caselli, Eric Reiman, Yalin Wang 0001 |
Medical Image Anal. | 14 |
| 2021 | Topology-preserving smoothing of retinotopic mapsabstractRetinotopic mapping, i.e., the mapping between visual inputs on the retina and neuronal activations in cortical visual areas, is one of the central topics in visual neuroscience. For human observers, the mapping is obtained by analyzing functional magnetic resonance imaging (fMRI) signals of cortical responses to slowly moving visual stimuli on the retina. Although it is well known from neurophysiology that the mapping is topological (i.e., the topology of neighborhood connectivity is preserved) within each visual area, retinotopic maps derived from the state-of-the-art methods are often not topological because of the low signal-to-noise ratio and spatial resolution of fMRI. The violation of topological condition is most severe in cortical regions corresponding to the neighborhood of the fovea (e.g., < 1 degree eccentricity in the Human Connectome Project (HCP) dataset), significantly impeding accurate analysis of retinotopic maps. This study aims to directly model the topological condition and generate topology-preserving and smooth retinotopic maps. Specifically, we adopted the Beltrami coefficient, a metric of quasiconformal mapping, to define the topological condition, developed a mathematical model to quantify topological smoothing as a constrained optimization problem, and elaborated an efficient numerical method to solve the problem. The method was then applied to V1, V2, and V3 simultaneously in the HCP dataset. Experiments with both simulated and real retinotopy data demonstrated that the proposed method could generate topological and smooth retinotopic maps. Yanshuai Tu, Duyan Ta, Zhonglin Lu, Yalin Wang 0001 |
PLoS Comput. Biol. | 4 |
| 2021 | Multi-Resemblance Multi-Target Low-Rank Coding for Prediction of Cognitive Decline With Longitudinal Brain ImagesabstractAn effective presymptomatic diagnosis and treatment of Alzheimer's disease (AD) would have enormous public health benefits. Sparse coding (SC) has shown strong potential for longitudinal brain image analysis in preclinical AD research. However, the traditional SC computation is time-consuming and does not explore the feature correlations that are consistent over the time. In addition, longitudinal brain image cohorts usually contain incomplete image data and clinical labels. To address these challenges, we propose a novel two-stage Multi-Resemblance Multi-Target Low-Rank Coding (MMLC) method, which encourages that sparse codes of neighboring longitudinal time points are resemblant to each other, favors sparse code low-rankness to reduce the computational cost and is resilient to both source and target data incompleteness. In stage one, we propose an online multi-resemblant low-rank SC method to utilize the common and task-specific dictionaries in different time points to immune to incomplete source data and capture the longitudinal correlation. In stage two, supported by a rigorous theoretical analysis, we develop a multi-target learning method to address the missing clinical label issue. To solve such a multi-task low-rank sparse optimization problem, we propose multi-task stochastic coordinate coding with a sequence of closed-form update steps which reduces the computational costs guaranteed by a theoretical convergence proof. We apply MMLC on a publicly available neuroimaging cohort to predict two clinical measures and compare it with six other methods. Our experimental results show our proposed method achieves superior results on both computational efficiency and predictive accuracy and has great potential to assist the AD prevention. Jie Zhang 0026, Qingyang Li 0001, Richard J. Caselli, Paul M. Thompson, Jieping Ye, Yalin Wang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Regularized Wasserstein Means for Aligning Distributional DataabstractWe propose to align distributional data from the perspective of Wasserstein means. We raise the problem of regularizing Wasserstein means and propose several terms tailored to tackle different problems. Our formulation is based on the variational transportation to distribute a sparse discrete measure into the target domain. The resulting sparse representation well captures the desired property of the domain while reducing the mapping cost. We demonstrate the scalability and robustness of our method with examples in domain adaptation, point set registration, and skeleton layout. Liang Mi, Wen Zhang 0010, Yalin Wang 0001 |
AAAI | 3 |
| 2020 | Convolutional Bayesian Models for Anatomical Landmarking on Multi-dimensional Shapes
Yonghui Fan, Yalin Wang 0001 |
MICCAI (4) | 2 |
| 2020 | Optimizing Visual Cortex Parameterization with Error-Tolerant Teichmüller Map in Retinotopic Mapping
Yanshuai Tu, Duyan Ta, Zhonglin Lu, Yalin Wang 0001 |
MICCAI (7) | 4 |
| 2020 | Deep Representation Learning for Multimodal Brain Networks
Wen Zhang 0010, Liang Zhan, Paul M. Thompson, Yalin Wang 0001 |
MICCAI (7) | 4 |
| 2020 | Regularize, Expand and Compress: NonExpansive Continual LearningabstractContinual learning (CL), the problem of lifelong learning where tasks arrive in sequence, has attracted increasing attention in the computer vision community lately. The goal of CL is to learn new tasks while maintaining the performance on the previously learned tasks. There are two major obstacles for CL of deep neural networks: catastrophic forgetting and limited model capacity. Inspired by the recent breakthroughs in automatically learning good neural network architectures, we develop a nonexpansive AutoML framework for CL termed Regularize, Expand and Compress (REC) to solve the above issues. REC is a unified framework with three highlights: 1) a novel regularized weight consolidation (RWC) algorithm to avoid forgetting, where accessing the data seen in the previously learned tasks is not required; 2) an automatic neural architecture search (AutoML) engine to expand the network to increase model capability; 3) smart compression of the expanded model after a new task is learned to improve the model efficiency. The experimental results on four different image recognition datasets demonstrate the superior performance of the proposed REC over other CL algorithms. Jie Zhang 0026, Junting Zhang, Shalini Ghosh, Dawei Li 0006, Heming Zhang 0003, Yalin Wang 0001 |
WACV | 7 |
| 2020 | Hyperbolic Wasserstein Distance for Shape IndexingabstractShape space is an active research topic in computer vision and medical imaging fields. The distance defined in a shape space may provide a simple and refined index to represent a unique shape. This work studies the Wasserstein space and proposes a novel framework to compute the Wasserstein distance between general topological surfaces by integrating hyperbolic Ricci flow, hyperbolic harmonic map, and hyperbolic power Voronoi diagram algorithms. The resulting hyperbolic Wasserstein distance can intrinsically measure the similarity between general topological surfaces. Our proposed algorithms are theoretically rigorous and practically efficient. It has the potential to be a powerful tool for 3D shape indexing research. We tested our algorithm with human face classification and Alzheimer's disease (AD) progression tracking studies. Experimental results demonstrated that our work may provide a succinct and effective shape index. Jie Shi 0001, Yalin Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Continually Modeling Alzheimer's Disease Progression via Deep Multi-order Preserving Weight Consolidation
Jie Zhang 0026, Yalin Wang 0001 |
MICCAI (2) | 2 |
| 2018 | Variational Wasserstein Clustering
Liang Mi, Wen Zhang 0010, Xianfeng Gu, Yalin Wang 0001 |
ECCV (15) | 4 |
| 2018 | Dynamically Hierarchy Revolution: DirNet for Compressing Recurrent Neural Network on Mobile DevicesabstractRecurrent neural networks (RNNs) achieve cutting-edge performance on a variety of problems. However, due to their high computational and memory demands, deploying RNNs on resource constrained mobile devices is a challenging task. To guarantee minimum accuracy loss with higher compression rate and driven by the mobile resource requirement, we introduce a novel model compression approach DirNet based on an optimized fast dictionary learning algorithm, which 1) dynamically mines the dictionary atoms of the projection dictionary matrix within layer to adjust the compression rate 2) adaptively changes the sparsity of sparse codes cross the hierarchical layers. Experimental results on language model and an ASR model trained with a 1000h speech dataset demonstrate that our method significantly outperforms prior approaches. Evaluated on off-the-shelf mobile devices, we are able to reduce the size of original model by eight times with real-time model inference and negligible accuracy loss. Jie Zhang 0026, Xiaolong Wang 0006, Dawei Li 0006, Yalin Wang 0001 |
IJCAI | 4 |
| 2018 | A Tetrahedron-Based Heat Flux Signature for Cortical Thickness Morphometry Analysis
Yonghui Fan, Gang Wang 0029, Natasha Leporé, Yalin Wang 0001 |
MICCAI (3) | 4 |
| 2018 | Multimodal Fusion of Brain Networks with Longitudinal Couplings
Wen Zhang 0010, Kai Shu, Suhang Wang, Huan Liu 0001, Yalin Wang 0001 |
MICCAI (3) | 5 |
| 2017 | An Optimal Transportation Based Univariate Neuroimaging IndexabstractThe alterations of brain structures and functions have been considered closely correlated to the change of cognitive performance due to neurodegenerative diseases such as Alzheimer's disease. In this paper, we introduce a variational framework to compute the optimal transformation (OT) in 3D space and propose a univariate neuroimaging index based on OT to measure such alterations. We compute the OT from each image to a template and measure the Wasserstein distance between them. By comparing the distances from all the images to the common template, we obtain a concise and informative index for each image. Our framework makes use of the Newton's method, which reduces the computational cost and enables itself to be applicable to large-scale datasets. The proposed work is a generic approach and thus may be applicable to various volumetric brain images, including structural magnetic resonance (sMR) and fluorodeoxyglucose positron emission tomography (FDG-PET) images. In the classification between Alzheimer's disease patients and healthy controls, our method achieves an accuracy of 82:30% on the Alzheimers Disease Neuroimaging Initiative (ADNI) baseline sMRI dataset and outperforms several other indices. On FDG-PET dataset, we boost the accuracy to 88:37% by leveraging pairwise Wasserstein distances. In a longitudinal study, we obtain a 5% significance with p-value = 1:13 ×105 in a t-test on FDG-PET. The results demonstrate a great potential of the proposed index for neuroimage analysis and the precision medicine research. Liang Mi, Wen Zhang 0010, Junwei Zhang 0010, Yonghui Fan, Dhruman Goradia, Kewei Chen 0001, Eric Reiman, Xianfeng Gu, Yalin Wang 0001 |
ICCV | 9 |
| 2017 | Intrinsic 3D Dynamic Surface Tracking based on Dynamic Ricci Flow and Teichmüller Mapabstract3D dynamic surface tracking is an important research problem and plays a vital role in many computer vision and medical imaging applications. However, it is still challenging to efficiently register surface sequences which has large deformations and strong noise. In this paper, we propose a novel automatic method for non-rigid 3D dynamic surface tracking with surface Ricci flow and Teichmüller map methods. According to quasi-conformal Teichmüller theory, the Techmüller map minimizes the maximal dilation so that our method is able to automatically register surfaces with large deformations. Besides, the adoption of Delaunay triangulation and quadrilateral meshes makes our method applicable to low quality meshes. In our work, the 3D dynamic surfaces are acquired by a high speed 3D scanner. We first identified sparse surface features using machine learning methods in the texture space. Then we assign landmark features with different curvature settings and the Riemannian metric of the surface is computed by the dynamic Ricci flow method, such that all the curvatures are concentrated on the feature points and the surface is flat everywhere else. The registration among frames is computed by the Teichmüller mappings, which aligns the feature points with least angle distortions. We apply our new method to multiple sequences of 3D facial surfaces with large expression deformations and compare them with two other state-of-the-art tracking methods. The effectiveness of our method is demonstrated by the clearly improved accuracy and efficiency. Xiaokang Yu, Na Lei, Yalin Wang 0001, Xianfeng Gu |
ICCV | 3 |
| 2017 | Conformal invariants for multiply connected surfaces: Application to landmark curve-based brain morphometry analysis
Jie Shi 0001, Wen Zhang 0010, Richard J. Caselli, Yalin Wang 0001 |
Medical Image Anal. | 5 |
| 2017 | Hyperbolic Harmonic Mapping for Surface RegistrationabstractAutomatic computation of surface correspondence via harmonic map is an active research field in computer vision, computer graphics and computational geometry. It may help document and understand physical and biological phenomena and also has broad applications in biometrics, medical imaging and motion capture industries. Although numerous studies have been devoted to harmonic map research, limited progress has been made to compute a diffeomorphic harmonic map on general topology surfaces with landmark constraints. This work conquers this problem by changing the Riemannian metric on the target surface to a hyperbolic metric so that the harmonic mapping is guaranteed to be a diffeomorphism under landmark constraints. The computational algorithms are based on Ricci flow and nonlinear heat diffusion methods. The approach is general and robust. We employ our algorithm to study the constrained surface registration problem which applies to both computer vision and medical imaging applications. Experimental results demonstrate that, by changing the Riemannian metric, the registrations are always diffeomorphic and achieve relatively high performance when evaluated with some popular surface registration evaluation standards. Wei Zeng 0002, Zhengyu Su, Hanna Damasio, Zhonglin Lu, Yalin Wang 0001, Shing-Tung Yau, Xianfeng Gu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2016 | Shape Analysis with Hyperbolic Wasserstein DistanceabstractShape space is an active research field in computer vision study. The shape distance defined in a shape space may provide a simple and refined index to represent a unique shape. Wasserstein distance defines a Riemannian metric for the Wasserstein space. It intrinsically measures the similarities between shapes and is robust to image noise. Thus it has the potential for the 3D shape indexing and classification research. While the algorithms for computing Wasserstein distance have been extensively studied, most of them only work for genus-0 surfaces. This paper proposes a novel framework to compute Wasserstein distance between general topological surfaces with hyperbolic metric. The computational algorithms are based on Ricci flow, hyperbolic harmonic map, and hyperbolic power Voronoi diagram and the method is general and robust. We apply our method to study human facial expression, longitudinal brain cortical morphometry with normal aging, and cortical shape classification in Alzheimer's disease (AD). Experimental results demonstrate that our method may be used as an effective shape index, which outperforms some other standard shape measures in our AD versus healthy control classification study. Jie Shi 0001, Wen Zhang 0010, Yalin Wang 0001 |
CVPR | 3 |
| 2016 | Hyperbolic Space Sparse Coding with Its Application on Prediction of Alzheimer's Disease in Mild Cognitive Impairment
Jie Zhang 0026, Jie Shi 0001, Cynthia M. Stonnington, Qingyang Li 0001, Boris Gutman, Kewei Chen 0001, Eric Reiman, Richard J. Caselli, Paul M. Thompson, Jieping Ye, Yalin Wang 0001 |
MICCAI (1) | 11 |
| 2016 | Large-Scale Collaborative Imaging Genetics Studies of Risk Genetic Factors for Alzheimer's Disease Across Multiple Institutions
Qingyang Li 0001, Tao Yang 0016, Liang Zhan, Derrek P. Hibar, Neda Jahanshad, Yalin Wang 0001, Jieping Ye, Paul M. Thompson, Jie Wang 0005 |
MICCAI (1) | 6 |
| 2015 | Multi-scale Heat Kernel Based Volumetric Morphology Signature
Gang Wang 0029, Yalin Wang 0001 |
MICCAI (3) | 2 |
| 2015 | A novel cortical thickness estimation method based on volumetric Laplace-Beltrami operator and heat kernel
Gang Wang 0029, Xiaofeng Zhang 0003, Qingtang Su, Jie Shi 0001, Richard J. Caselli, Yalin Wang 0001 |
Medical Image Anal. | 6 |
| 2015 | Optimal Mass Transport for Shape Matching and ComparisonabstractSurface based 3D shape analysis plays a fundamental role in computer vision and medical imaging. This work proposes to use optimal mass transport map for shape matching and comparison, focusing on two important applications including surface registration and shape space. The computation of the optimal mass transport map is based on Monge-Brenier theory, in comparison to the conventional method based on Monge-Kantorovich theory, this method significantly improves the efficiency by reducing computational complexity from O(n(2)) to O(n) . For surface registration problem, one commonly used approach is to use conformal map to convert the shapes into some canonical space. Although conformal mappings have small angle distortions, they may introduce large area distortions which are likely to cause numerical instability thus resulting failures of shape analysis. This work proposes to compose the conformal map with the optimal mass transport map to get the unique area-preserving map, which is intrinsic to the Riemannian metric, unique, and diffeomorphic. For shape space study, this work introduces a novel Riemannian framework, Conformal Wasserstein Shape Space, by combing conformal geometry and optimal mass transport theory. In our work, all metric surfaces with the disk topology are mapped to the unit planar disk by a conformal mapping, which pushes the area element on the surface to a probability measure on the disk. The optimal mass transport provides a map from the shape space of all topological disks with metrics to the Wasserstein space of the disk and the pullback Wasserstein metric equips the shape space with a Riemannian metric. We validate our work by numerous experiments and comparisons with prior approaches and the experimental results demonstrate the efficiency and efficacy of our proposed approach. Zhengyu Su, Yalin Wang 0001, Wei Zeng 0002, Jian Sun 0002, Feng Luo 0002, Xianfeng Gu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Hyperbolic Harmonic Mapping for Constrained Brain Surface RegistrationabstractAutomatic computation of surface correspondence via harmonic map is an active research field in computer vision, computer graphics and computational geometry. It may help document and understand physical and biological phenomena and also has broad applications in biometrics, medical imaging and motion capture. Although numerous studies have been devoted to harmonic map research, limited progress has been made to compute a diffeomorphic harmonic map on general topology surfaces with landmark constraints. This work conquer this problem by changing the Riemannian metric on the target surface to a hyperbolic metric, so that the harmonic mapping is guaranteed to be a diffeomorphism under landmark constraints. The computational algorithms are based on the Ricci flow method and the method is general and robust. We apply our algorithm to study constrained human brain surface registration problem. Experimental results demonstrate that, by changing the Riemannian metric, the registrations are always diffeomorphic, and achieve relative high performance when evaluated with some popular cortical surface registration evaluation standards. Wei Zeng 0002, Zhengyu Su, Hanna Damasio, Zhonglin Lu, Yalin Wang 0001, Shing-Tung Yau, Xianfeng Gu |
CVPR | 6 |
| 2013 | Area Preserving Brain MappingabstractBrain mapping transforms the brain cortical surface to canonical planar domains, which plays a fundamental role in morphological study. Most existing brain mapping methods are based on angle preserving maps, which may introduce large area distortions. This work proposes an area preserving brain mapping method based on Monge-Brenier theory. The brain mapping is intrinsic to the Riemannian metric, unique, and diffeomorphic. The computation is equivalent to convex energy minimization and power Voronoi diagram construction. Comparing to the existing approaches based on Monge-Kantorovich theory, the proposed one greatly reduces the complexity (from n2unknowns to n ), and improves the simplicity and efficiency. Experimental results on caudate nucleus surface mapping and cortical surface mapping demonstrate the efficacy and efficiency of the proposed method. Conventional methods for caudate nucleus surface mapping may suffer from numerical instability, in contrast, current method produces diffeomorpic mappings stably. In the study of cortical surface classification for recognition of Alzheimer's Disease, the proposed method outperforms some other morphometry features. Zhengyu Su, Wei Zeng 0002, Yalin Wang 0001, Jian Sun 0002, Xianfeng Gu |
CVPR | 4 |
| 2013 | Multi-source learning with block-wise missing data for Alzheimer's disease predictionabstractWith the advances and increasing sophistication in data collection techniques, we are facing with large amounts of data collected from multiple heterogeneous sources in many applications. For example, in the study of Alzheimer's Disease (AD), different types of measurements such as neuroimages, gene/protein expression data, genetic data etc. are often collected and analyzed together for improved predictive power. It is believed that a joint learning of multiple data sources is beneficial as different data sources may contain complementary information, and feature-pruning and data source selection are critical for learning interpretable models from high-dimensional data. Very often the collected data comes with block-wise missing entries; for example, a patient without the MRI scan will have no information in the MRI data block, making his/her overall record incomplete. There has been a growing interest in the data mining community on expanding traditional techniques for single-source complete data analysis to the study of multi-source incomplete data. The key challenge is how to effectively integrate information from multiple heterogeneous sources in the presence of block-wise missing data. In this paper we first investigate the situation of complete data and present a unified ``bi-level" learning model for multi-source data. Then we give a natural extension of this model to the more challenging case with incomplete data. Our major contributions are threefold: (1) the proposed models handle both feature-level and source-level analysis in a unified formulation and include several existing feature learning approaches as special cases; (2) the model for incomplete data avoids direct imputation of the missing elements and thus provides superior performances. Moreover, it can be easily generalized to other applications with block-wise missing data sources; (3) efficient optimization algorithms are presented for both the complete and incomplete models. We have performed comprehensive evaluations of the proposed models on the application of AD diagnosis. Our proposed models compare favorably against existing approaches. Shuo Xiang, Lei Yuan 0001, Wei Fan 0001, Yalin Wang 0001, Paul M. Thompson, Jieping Ye |
KDD | 4 |
| 2013 | Teichmüller Shape Descriptor and Its Application to Alzheimer's Disease Study
Wei Zeng 0002, Yalin Wang 0001, Shing-Tung Yau, Xianfeng Gu |
Int. J. Comput. Vis. | 3 |
| 2012 | Multi-source learning for joint analysis of incomplete multi-modality neuroimaging dataabstractIncomplete data present serious problems when integrating largescale brain imaging data sets from different imaging modalities. In the Alzheimer's Disease Neuroimaging Initiative (ADNI), for example, over half of the subjects lack cerebrospinal fluid (CSF) measurements; an independent half of the subjects do not have fluorodeoxyglucose positron emission tomography (FDG-PET) scans; many lack proteomics measurements. Traditionally, subjects with missing measures are discarded, resulting in a severe loss of available information. We address this problem by proposing two novel learning methods where all the samples (with at least one available data source) can be used. In the first method, we divide our samples according to the availability of data sources, and we learn shared sets of features with state-of-the-art sparse learning methods. Our second method learns a base classifier for each data source independently, based on which we represent each source using a single column of prediction scores; we then estimate the missing prediction scores, which, combined with the existing prediction scores, are used to build a multi-source fusion model. To illustrate the proposed approaches, we classify patients from the ADNI study into groups with Alzheimer's disease (AD), mild cognitive impairment (MCI) and normal controls, based on the multi-modality data. At baseline, ADNI's 780 participants (172 AD, 397 MCI, 211 Normal), have at least one of four data types: magnetic resonance imaging (MRI), FDG-PET, CSF and proteomics. These data are used to test our algorithms. Comprehensive experiments show that our proposed methods yield stable and promising results. Lei Yuan 0001, Yalin Wang 0001, Paul M. Thompson, Vaibhav A. Narayan, Jieping Ye |
KDD | 2 |
| 2012 | Modeling Dynamic Cellular Morphology in Images
Xing An, Yonggang Shi, Yalin Wang 0001, Shantanu H. Joshi |
MICCAI (1) | 5 |
| 2012 | Brain Surface Conformal Parameterization With the Ricci FlowabstractIn brain mapping research, parameterized 3-D surface models are of great interest for statistical comparisons of anatomy, surface-based registration, and signal processing. Here, we introduce the theories of continuous and discrete surface Ricci flow, which can create Riemannian metrics on surfaces with arbitrary topologies with user-defined Gaussian curvatures. The resulting conformal parameterizations have no singularities and they are intrinsic and stable. First, we convert a cortical surface model into a multiple boundary surface by cutting along selected anatomical landmark curves. Secondly, we conformally parameterize each cortical surface to a parameter domain with a user-designed Gaussian curvature arrangement. In the parameter domain, a shape index based on conformal invariants is computed, and inter-subject cortical surface matching is performed by solving a constrained harmonic map. We illustrate various target curvature arrangements and demonstrate the stability of the method using longitudinal data. To map statistical differences in cortical morphometry, we studied brain asymmetry in 14 healthy control subjects. We used a manifold version of Hotelling's T(2) test, applied to the Jacobian matrices of the surface parameterizations. A permutation test, along with the cumulative distribution of p-values, were used to estimate the overall statistical significance of differences. The results show our algorithm's power to detect subtle group differences in cortical surfaces. Yalin Wang 0001, Jie Shi 0001, Xiaotian Yin, Xianfeng Gu, Tony F. Chan, Shing-Tung Yau, Arthur W. Toga, Paul M. Thompson |
IEEE Trans. Medical Imaging | 1 |
| 2010 | Optimized Conformal Surface Registration with Shape-based Landmark MatchingabstractSurface registration, which transforms different sets of surface data into one common reference space, is an important process which allows us to compare or integrate the surface data effectively. If a nonrigid transformation is required, surface registration is commonly done by parameterizing the surfaces onto a simple parameter domain, such as the unit square or sphere. In this work, we are interested in looking for meaningful registrations between surfaces through parameterizations, using prior features in the form of landmark curves on the surfaces. In particular, we generate optimized conformal parameterizations which match landmark curves exactly with shape-based correspondences between them. We propose a variational method to minimize a compound energy functional that measures the harmonic energy of the parameterization maps and the shape dissimilarity between mapped points on the landmark curves. The novelty is that the computed maps are guaranteed to align the landmark features consistently and give a shape-based diffeomorphism between the landmark curves. We achieve this by intrinsically modeling our search space of maps as flows of smooth vector fields that do not flow across the landmark curves. By using the local surface geometry on the curves to define a shape measure, we compute registrations that ensure consistent correspondences between anatomical features. We test our algorithm on synthetic surface data. An application of our model to medical imaging research is shown, using experiments on brain cortical surfaces, with anatomical (sulcal) landmarks delineated, which show that our computed maps give a shape-based alignment of the sulcal curves without significantly impairing conformality. This ensures correct averaging and comparison of data across subjects. Lok Ming Lui, Sheshadri R. Thiruvenkadam, Yalin Wang 0001, Paul M. Thompson, Tony F. Chan |
SIAM J. Imaging Sci. | 3 |
| 2009 | Shape analysis with conformal invariants for multiply connected domains and its application to analyzing brain morphologyabstractAll surfaces can be classified by the conformal equivalence relation. Conformal invariants, which are shape indices that can be defined intrinsically on a surface, may be used to identify which surfaces are conformally equivalent, and they can also be used to measure surface deformation. Here we propose to compute a conformal invariant, or shape index, that is associated with the perimeter of the inner concentric circle in the hyperbolic parameter plane. With the surface Ricci flow method, we can conformally map a multiply connected domain to a multi-hole disk and this conformal map can preserve the values of the conformal invariant. Our algorithm provides a stable method to map the values of this shape index in the 2D (hyperbolic space) parameter domain. We also applied this new shape index for analyzing abnormalities in brain morphology in Alzheimer's disease (AD) and Williams syndrome (WS). After cutting along various landmark curves on surface models of the cerebral cortex or hippocampus, we obtained multiple connected domains. We conformally projected the surfaces to hyperbolic plane with surface Ricci flow method, accurately computed the proposed conformal invariant for each selected landmark curve, and assembled these into a feature vector.We also detected group differences in brain structure based on multivariate analysis of the surface deformation tensors induced by these Ricci flow mappings. Experimental results with 3D MRI data from 80 subjects demonstrate that our method powerfully detects brain surface abnormalities when combined with a constrained harmonic map based surface registration method. Yalin Wang 0001, Xianfeng Gu, Tony F. Chan, Paul M. Thompson |
CVPR | 1 |
| 2009 | Shape analysis with multivariate tensor-based morphometry and holomorphic differentialsabstractIn this paper, we propose multivariate tensor-based surface morphometry, a new method for surface analysis, using holomorphic differentials; we also apply it to study brain anatomy. Differential forms provide a natural way to parameterize 3D surfaces, but the multivariate statistics of the resulting surface metrics have not previously been investigated. We computed new statistics from the Riemannian metric tensors that retain the full information in the deformation tensor fields. We present the canonical holomorphic one-forms with improved numerical accuracy and computational efficiency. We applied this framework to 3D MRI data to analyze hippocampal surface morphometry in Alzheimer's Disease (AD; 12 subjects), lateral ventricular surface morphometry in HIV/AIDS (11 subjects) and biomarkers in lateral ventricles in HIV/AIDS (11 subjects). Experimental results demonstrated that our method powerfully detected brain surface abnormalities. Multivariate statistics on the local tensors outperformed other TBM methods including analysis of the Jacobian determinant, the largest eigenvalue, or the pair of eigenvalues, of the surface Jacobian matrix. Yalin Wang 0001, Tony F. Chan, Arthur W. Toga, Paul M. Thompson |
ICCV | 1 |
| 2009 | Studying brain morphometry using conformal equivalence classabstractTwo surfaces are conformally equivalent if there exists a bijective angle-preserving map between them. The Teichmüller space for surfaces with the same topology is a finite-dimensional manifold, where each point represents a conformal equivalence class, and the conformal map is homotopic to the identity map. In this paper, we propose a novel method to apply conformal equivalence based shape index to study brain morphometry. The shape index is defined based on Teichmüller space coordinates. It is intrinsic, and invariant under conformal transformations, rigid motions and scaling. It is also simple to compute; no registration of surfaces is needed. Using the Yamabe flow method, we can conformally map a genus-zero open boundary surface to the Poincaré disk. The shape indices that we compute are the lengths of a special set of geodesics under hyperbolic metric. By computing and studying this shape index and its statistical behavior, we can analyze differences in anatomical morphometry due to disease or development. Study on twin lateral ventricular surface data shows it may help detect generic influence on lateral ventricular shapes. In leave-one-out validation tests, we achieved 100% accurate classification (versus only 68% accuracy for volume measures) in distinguishing 11 HIV/AIDS individuals from 8 healthy control subjects, based on Teichmüller coordinates for lateral ventricular surfaces extracted from their 3D MRI scans.Our conformal invariants, the Teichmüller coordinates, successfully classified all lateral ventricular surfaces, showing their promise for analyzing anatomical surface morphometry. Yalin Wang 0001, Yi-Yu Chou, Xianfeng Gu, Tony F. Chan, Arthur W. Toga, Paul M. Thompson |
ICCV | 1 |
| 2009 | Multivariate Tensor-Based Brain Anatomical Surface Morphometry via Holomorphic One-Forms
Yalin Wang 0001, Tony F. Chan, Arthur W. Toga, Paul M. Thompson |
MICCAI (1) | 1 |
| 2009 | Teichmüller Shape Space Theory and Its Application to Brain Morphometry
Yalin Wang 0001, Xianfeng Gu, Tony F. Chan, Shing-Tung Yau, Arthur W. Toga, Paul M. Thompson |
MICCAI (1) | 1 |
| 2008 | Optimized Conformal Parameterization of Cortical Surfaces Using Shape Based Matching of Landmark Curves
Lok Ming Lui, Sheshadri R. Thiruvenkadam, Yalin Wang 0001, Tony F. Chan, Paul M. Thompson |
MICCAI (1) | 3 |
| 2008 | Conformal Slit Mapping and Its Applications to Brain Surface Parameterization
Yalin Wang 0001, Xianfeng Gu, Tony F. Chan, Paul M. Thompson, Shing-Tung Yau |
MICCAI (1) | 1 |
| 2007 | Brain Surface Conformal Parameterization Using Riemann Surface StructureabstractIn medical imaging, parameterized 3-D surface models are useful for anatomical modeling and visualization, statistical comparisons of anatomy, and surface-based registration and signal processing. Here we introduce a parameterization method based on Riemann surface structure, which uses a special curvilinear net structure (conformal net) to partition the surface into a set of patches that can each be conformally mapped to a parallelogram. The resulting surface subdivision and the parameterizations of the components are intrinsic and stable (their solutions tend to be smooth functions and the boundary conditions of the Dirichlet problem can be enforced). Conformal parameterization also helps transform partial differential equations (PDEs) that may be defined on 3-D brain surface manifolds to modified PDEs on a two-dimensional parameter domain. Since the Jacobian matrix of a conformal parameterization is diagonal, the modified PDE on the parameter domain is readily solved. To illustrate our techniques, we computed parameterizations for several types of anatomical surfaces in 3-D magnetic resonance imaging scans of the brain, including the cerebral cortex, hippocampi, and lateral ventricles. For surfaces that are topologically homeomorphic to each other and have similar geometrical structures, we show that the parameterization results are consistent and the subdivided surfaces can be matched to each other. Finally, we present an automatic sulcal landmark location algorithm by solving PDEs on cortical surfaces. The landmark detection results are used as constraints for building conformal maps between surfaces that also match explicitly defined landmarks. Yalin Wang 0001, Lok Ming Lui, Xianfeng Gu, Kiralee M. Hayashi, Tony F. Chan, Arthur W. Toga, Paul M. Thompson, Shing-Tung Yau |
IEEE Trans. Medical Imaging | 1 |
| 2006 | Automatic Landmark Tracking and its Application to the Optimization of Brain Conformal MappingabstractAnatomical features on cortical surfaces are usually represented by landmark curves, called sulci/gyri curves. These landmark curves are important information for neuroscientists to study brain diseases and to match different cortical surfaces. Manual labelling of these landmark curves is time-consuming, especially when there is a large set of data. In this paper, we proposed to trace the landmark curves on cortical surfaces automatically based on the principal directions. Suppose we are given the global conformal parametrization of a cortical surface, By fixing two endpoints, the anchor points, we propose to trace the landmark curves iteratively on the spherical/rectangular parameter domain along the principal direction. Consequently, the landmark curves can be mapped onto the cortical surface. To speed up the iterative scheme, a good initial guess of the landmark curve is necessary. We proposed a method to get a good initialization by extracting the high curvature region on the cortical surface using the Chan-Vese segmentation. This involves solving a PDE on the manifold using our global conformal parametrization technique. Experimental results show that the landmark curves detected by our algorithm closely resemble to those manually labelled curves. As an application, we used these automatically labelled landmark curves to build average cortical surfaces with an optimized brain conformal mapping method. Experimental results show our method can help automatically matching brain cortical surfaces. Lok Ming Lui, Yalin Wang 0001, Tony F. Chan, Paul M. Thompson |
CVPR (2) | 2 |
| 2006 | A Landmark-Based Brain Conformal Parametrization with Automatic Landmark Tracking Technique
Lok Ming Lui, Yalin Wang 0001, Tony F. Chan, Paul M. Thompson |
MICCAI (2) | 2 |
| 2006 | Brain Surface Conformal Parameterization with Algebraic Functions
Yalin Wang 0001, Xianfeng Gu, Tony F. Chan, Paul M. Thompson, Shing-Tung Yau |
MICCAI (2) | 1 |
| 2006 | Document zone content classification and its performance evaluation
Yalin Wang 0001, Ihsin T. Phillips, Robert M. Haralick |
Pattern Recognit. | 1 |
| 2005 | Mutual Information-Based 3D Surface Matching with Applications to Face Recognition and Brain MappingabstractFace recognition and many medical imaging applications require the computation of dense correspondence vector fields that match one surface with another. In brain imaging, surface-based registration is useful for tracking brain change, and for creating statistical shape models of anatomy. Based on surface correspondences, metrics can also be designed to measure differences in facial geometry and expressions. To avoid the need for a large set of manually-defined landmarks to constrain these surface correspondences, we developed an algorithm to automate the matching of surface features. It extends the mutual information method to automatically match general 3D surfaces (including surfaces with a branching topology). We use diffeomorphic flows to optimally align the Riemann surface structures of two surfaces. First, we use holomorphic I-forms to induce consistent conformal grids on both surfaces. High genus surfaces are mapped to a set of rectangles in the Euclidean plane and closed genus-zero surfaces are mapped to the sphere. Next, we compute stable geometric features (mean curvature and conformal factor) and pull them back as scalar fields onto the 2D parameter domains. Mutual information is used as a cost functional to drive a fluid flow in the parameter domain that optimally aligns these surface features. A diffeomorphic surface-to-surface mapping is then recovered that matches surfaces in 3D. Lastly, we present a spectral method that ensures that the grids induced on the target surface remain conformal when pulled through the correspondence field. Using the chain rule, we express the gradient of the mutual information between surfaces in the conformal basis of the source surface. This finite-dimensional linear space generates all conformal reparameterizations of the surface. Illustrative experiments apply the method to face recognition and to the registration of brain structures, such as the hippocampus in 3D MRI scans, a key step in understanding brain shape alterations in Alzheimer's disease and schizophrenia. Yalin Wang 0001, Ming-Chang Chiang, Paul M. Thompson |
ICCV | 1 |
| 2005 | Surface Parameterization Using Riemann Surface StructureabstractWe propose a general method that parameterizes general surfaces with complex (possible branching) topology using Riemann surface structure. Rather than evolve the surface geometry to a plane or sphere, we instead use the fact that all orientable surfaces are Riemann surfaces and admit conformal structures, which induce special curvilinear coordinate systems on the surfaces. We can then automatically partition the surface using a critical graph that connects zero points in the global conformal structure on the surface. The trajectories of iso-parametric curves canonically partition a surface into patches. Each of these patches is either a topological disk or a cylinder and can be conformally mapped to a parallelogram by integrating a holomorphic I-form defined on the surface. The resulting surface subdivision and the parameterizations of the components are intrinsic and stable. For surfaces with similar topology and geometry, we show that the parameterization results are consistent and the subdivided surfaces can be matched to each other using constrained harmonic maps. The surface similarity can be measured by direct computation of distance between each pair of corresponding points on two surfaces. To illustrate the technique, we computed conformal structures for anatomical surfaces in MRI scans of the brain and human face surfaces. We found that the resulting parameterizations were consistent across subjects, even for branching structures such as the ventricles, which are otherwise difficult to parameterize. Our method provides a surface-based framework for statistical comparison of surfaces and for generating grids on surfaces for PDE-based signal processing. Yalin Wang 0001, Xianfeng Gu, Kiralee M. Hayashi, Tony F. Chan, Paul M. Thompson, Shing-Tung Yau |
ICCV | 1 |
| 2005 | Automated Surface Matching Using Mutual Information Applied to Riemann Surface Structures
Yalin Wang 0001, Ming-Chang Chiang, Paul M. Thompson |
MICCAI (2) | 1 |
| 2005 | Brain Surface Parameterization Using Riemann Surface Structure
Yalin Wang 0001, Xianfeng Gu, Kiralee M. Hayashi, Tony F. Chan, Paul M. Thompson, Shing-Tung Yau |
MICCAI (2) | 1 |
| 2005 | Optimization of Brain Conformal Mapping with Landmarks
Yalin Wang 0001, Lok Ming Lui, Tony F. Chan, Paul M. Thompson |
MICCAI (2) | 1 |
| 2004 | Optimal Global Conformal Surface ParameterizationabstractAll orientable metric surfaces are Riemann surfaces and admit global conformal parameterizations. Riemann surface structure is a fundamental structure and governs many natural physical phenomena, such as heat diffusion and electro-magnetic fields on the surface. A good parameterization is crucial for simulation and visualization. This paper provides an explicit method for finding optimal global conformal parameterizations of arbitrary surfaces. It relies on certain holomorphic differential forms and conformal mappings from differential geometry and Riemann surface theories. Algorithms are developed to modify topology, locate zero points, and determine cohomology types of differential forms. The implementation is based on a finite dimensional optimization method. The optimal parameterization is intrinsic to the geometry, preserves angular structure, and can play an important role in various applications including texture mapping, remeshing, morphing and simulation. The method is demonstrated by visualizing the Riemann surface structure of real surfaces represented as triangle meshes. Miao Jin, Yalin Wang 0001, Shing-Tung Yau, Xianfeng Gu |
IEEE Visualization | 2 |
| 2004 | Table structure understanding and its performance evaluation
Yalin Wang 0001, Ihsin T. Phillips, Robert M. Haralick |
Pattern Recognit. | 1 |
| 2004 | Genus zero surface conformal mapping and its application to brain surface mappingabstractWe developed a general method for global conformal parameterizations based on the structure of the cohomology group of holomorphic one-forms for surfaces with or without boundaries (Gu and Yau, 2002), (Gu and Yau, 2003). For genus zero surfaces, our algorithm can find a unique mapping between any two genus zero manifolds by minimizing the harmonic energy of the map. In this paper, we apply the algorithm to the cortical surface matching problem. We use a mesh structure to represent the brain surface. Further constraints are added to ensure that the conformal map is unique. Empirical tests on magnetic resonance imaging (MRI) data show that the mappings preserve angular relationships, are stable in MRIs acquired at different times, and are robust to differences in data triangulation, and resolution. Compared with other brain surface conformal mapping algorithms, our algorithm is more stable and has good extensibility. Xianfeng Gu, Yalin Wang 0001, Tony F. Chan, Paul M. Thompson, Shing-Tung Yau |
IEEE Trans. Medical Imaging | 2 |
| 2002 | Detecting Tables in HTML Documents
Yalin Wang 0001, Jianying Hu |
Document Analysis Systems | 1 |
| 2002 | A Study on the Document Zone Content Classification Problem
Yalin Wang 0001, Ihsin T. Phillips, Robert M. Haralick |
Document Analysis Systems | 1 |
| 2002 | Table Detection via Probability Optimization
Yalin Wang 0001, Ihsin T. Phillips, Robert M. Haralick |
Document Analysis Systems | 1 |
| 2002 | A machine learning based approach for table detection on the webabstractTable is a commonly used presentation scheme, especially for describing relational information. However, table understanding remains an open problem. In this paper, we consider the problem of table detection in web documents. Its potential applications include web mining, knowledge management, and web content summarization and delivery to narrow-bandwidth devices. We describe a machine learning based approach to classify each given table entity as either genuine or non-genuine. Various features reflecting the layout as well as content characteristics of tables are studied.In order to facilitate the training and evaluation of our table classifier, we designed a novel web document table ground truthing protocol and used it to build a large table ground truth database. The database consists of 1,393 HTML files collected from hundreds of different web sites and contains 11,477 leaf TABLE elements, out of which 1,740 are genuine tables. Experiments were conducted using the cross validation method and an F-measure of 95.89% was achieved. Yalin Wang 0001, Jianying Hu |
WWW | 1 |
| 2001 | Automatic Table Ground Truth Generation and a Background-Analysis-Based Table Structure Extraction MethodabstractWe first describe an automatic table ground truth generation system which can efficiently generate a large amount of accurate table ground truth suitable for the development of table detection algorithms. Then a novel background analysis-based, coarse-to-fine table identification algorithm and an X-Y cut table decomposition algorithm are described. We discuss an experimental protocol to evaluate the table detection algorithms. For a total of 1,125 document pages having 518 table entities and a total of 10,941 cell entities, our table detection algorithm takes line, word segmentation results as input and obtains around 90% cell correct detection rates. Yalin Wang 0001, Robert M. Haralick, Ihsin T. Phillips |
ICDAR | 1 |
| 2001 | Zone Content Classification and its Performance EvaluationabstractWe present an improved zone content classification method and its performance evaluation. We added two new features to the feature vector from one previously published method (Sivaramakrishnan et al., 1995). We assumed different independence relationships in two zone sets. We used an optimized binary decision tree to estimate the maximum zone content class probability in one set while using the Viterbi algorithm to find the optimal solution for a zone sequence in the other set. The training, pruning and testing data set for the algorithm include 1,600 images drawn from the UWCDROM III document image database. The classifier is able to classify each given scientific and technical document zone into one of the nine classes, 2 text classes (of font size 4 - 18pt and font size 19 - 32 pt), math, table, halftone, map/drawing, ruling, logo, and others. Compared with our previous work (Wang et al., 2000), it raised the accuracy rate to 98.52% from 97.53% and reduced the mean false alarm rate to 0.53% from 1.26%. Yalin Wang 0001, Robert M. Haralick, Ihsin T. Phillips |
ICDAR | 1 |
| 2000 | Algorithm Performance ContestabstractThis contest involved the running and evaluation of computer vision and pattern recognition techniques on different data sets with known groundwidth. The contest included three areas; binary shape recognition, symbol recognition and image flow estimation. A package was made available for each area. Each package contained either real images with manual groundtruth or programs to generate data sets of ideal as well as noisy images with known groundtruth. They also contained programs to evaluate the results of an algorithm according to the given groundtruth. These evaluation criteria included the generation of confusion matrices, computation of the misdetection and false alarm rates and other performance measures suitable for the problems. The paper summarizes the data generation for each area and experimental results for a total of six participating algorithms. Selim Aksoy, Michael L. Schauf, Mingzhou Song 0001, Yalin Wang 0001, Robert M. Haralick, Jim R. Parker, Juraj Pivovarov, Dominik Royko, Changming Sun, Gunnar Farnebäck |
ICPR | 5 |
| 2000 | Statistical-Based Approach to Word SegmentationabstractThis paper presents a text word extraction algorithm that takes a set of bounding boxes of glyphs and their associated text lines of a given document and partitions the glyphs into a set of text words, using only the geometric information of the input glyphs. The algorithm is probability based. An iterative, relaxation-like method is used to find the partitioning solution that maximizes the joint probability. To evaluate the performance of our test word extraction algorithm, we used a 3-fold validation method and developed a quantitative performance measure. The algorithm was evaluated on the UW-III database of some 1600 scanned document image pages. An area-overlap measure was used to find the correspondence between the detected entities and the ground-truth. For a total of 827, 433 ground truth words, the algorithm identified and segmented 800, 149 words correctly, an accuracy of 97.43%. Yalin Wang 0001, Robert M. Haralick, Ihsin T. Phillips |
ICPR | 1 |