Longin Jan Latecki

dblp:44/1633 · DBLP profile ↗
← Back
179ranked-venue papers
35as first author
29since 2021 · last 2026
0000-0002-5102-8244ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 127 · 26 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 85 · 12 first-author · 19 since 2021Databases, data management, data science and information retrieval · 15 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Direct Visual Grounding by Directing Attention of Visual Tokens
abstract
Vision Language Models (VLMs) mix visual tokens and text tokens. A puzzling issue is the fact that visual tokens most related to the query receive little to no attention in the final layers of the LLM module of VLMs from the answer tokens, where all tokens are treated equally, in particular, visual and language tokens in the LLM attention layers. This fact may result in wrong answers to visual questions, as our experimental results confirm. It appears that the standard next-token prediction (NTP) loss provides an insufficient signal for directing attention to visual tokens. We hypothesize that a more direct supervision of the attention of visual tokens to corresponding language tokens in the LLM module of VLMs will lead to improved performance on visual tasks. To demonstrate that this is indeed the case, we propose a novel loss function that directly supervises the attention of visual tokens. It directly grounds the answer language tokens in images by directing their attention to the relevant visual tokens. This is achieved by aligning the attention distribution of visual tokens to ground truth attention maps with KL divergence. The ground truth attention maps are obtained from task geometry in synthetic cases or from standard grounding annotations (e.g., bounding boxes or point annotations) in real images, and are used inside the LLM for attention supervision without requiring new labels. The obtained KL attention loss (KLAL) when combined with NTP encourages VLMs to attend to relevant visual tokens while generating answer tokens. This results in notable improvements across geometric tasks, pointing, and referring expression comprehension on both synthetic and real-world data, as demonstrated by our experiments. We also introduce a new dataset to evaluate the line tracing abilities of VLMs. Surprisingly, even commercial VLMs do not perform well on this task.
Parsa Esmaeilkhani, Longin Jan Latecki
WACV2
2025 Generalizable Detection of Student Engagement in Online Learning Environments
Lu Pang 0003, Tony Siu, Anis Alazzawe, Krishna Kant 0001, Longin Jan Latecki
CAIP (2)5
2025 VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose Estimation
abstract
Localizing predefined 3D keypoints in a 2D image is an effective way to establish 3D-2D correspondences for instance-level 6DoF object pose estimation. However, unreliable localization results of invisible keypoints degrade the quality of correspondences. In this paper, we address this issue by localizing the important keypoints in terms of visibility. Since keypoint visibility information is currently missing in the dataset collection process, we propose an efficient way to generate binary visibility labels from available object-level annotations, for keypoints of both asymmetric objects and symmetric objects. We further derive real-valued visibility-aware importance from binary labels based on the PageRank algorithm. Taking advantage of the flexibility of our visibility-aware importance, we construct VAPO (Visibility-Aware POse estimator) by integrating the visibility-aware importance with a state-of-the-art pose estimation algorithm, along with additional positional encoding. VAPO can work in both CAD-based and CAD-free settings. Extensive experiments are conducted on popular pose estimation benchmarks including Linemod, Linemod-Occlusion, and YCB-V, demonstrating that VAPO clearly achieves state-of-the-art performances. Project page: https://github.com/RuyiLian/VAPO.
Ruyi Lian, Yuewei Lin, Longin Jan Latecki, Haibin Ling
IROS3
2025 DPNet: Dual-Path Network for Real-Time Object Detection With Lightweight Attention
abstract
The recent advances in compressing high-accuracy convolutional neural networks (CNNs) have witnessed remarkable progress in real-time object detection. To accelerate detection speed, lightweight detectors always have few convolution layers using a single-path backbone. Single-path architecture, however, involves continuous pooling and downsampling operations, always resulting in coarse and inaccurate feature maps that are disadvantageous to locate objects. On the other hand, due to limited network capacity, recent lightweight networks are often weak in representing large-scale visual data. To address these problems, we present a dual-path network, named DPNet, with a lightweight attention scheme for real-time object detection. The dual-path architecture enables us to extract in parallel high-level semantic features and low-level object details. Although DPNet has a nearly duplicated shape with respect to single-path detectors, the computational costs and model size are not significantly increased. To enhance representation capability, a lightweight self-correlation module (LSCM) is designed to capture global interactions, with only a few computational overheads and network parameters. In the neck, LSCM is extended into a lightweight cross correlation module (LCCM), capturing mutual dependencies among neighboring scale features. We have conducted exhaustive experiments on MS COCO, Pascal VOC 2007, and ImageNet datasets. The experimental results demonstrate that DPNet achieves a state-of-the-art trade off between detection accuracy and implementation efficiency. More specifically, DPNet achieves 31.3% AP on MS COCO test-dev, 82.7% mAP on Pascal VOC 2007 test set, and 41.6% mAP on ImageNet validation set, together with nearly 2.5M model size, 1.04 GFLOPs, and 164 and 196 frames/s (FPS) FPS for input images of three datasets.
Quan Zhou 0004, Huimin Shi, Weikang Xiang, Bin Kang, Longin Jan Latecki
IEEE Trans. Neural Networks Learn. Syst.5
2024 SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions
abstract
We present SciDMT, an enhanced and expanded corpus for scientific mention detection, offering a significant advancement over existing related resources. SciDMT contains annotated scientific documents for datasets (D), methods (M), and tasks (T). The corpus consists of two components: 1) the SciDMT main corpus, which includes 48 thousand scientific articles with over 1.8 million weakly annotated mention annotations in the format of in-text span, and 2) an evaluation set, which comprises 100 scientific articles manually annotated for evaluation purposes. To the best of our knowledge, SciDMT is the largest corpus for scientific entity mention detection. The corpus’s scale and diversity are instrumental in developing and refining models for tasks such as indexing scientific papers, enhancing information retrieval, and improving the accessibility of scientific knowledge. We demonstrate the corpus’s utility through experiments with advanced deep learning architectures like SciBERT and GPT-3.5. Our findings establish performance baselines and highlight unresolved challenges in scientific mention detection. SciDMT serves as a robust benchmark for the research community, encouraging the development of innovative models to further the field of scientific information extraction.
Huitong Pan, Cornelia Caragea, Eduard C. Dragut, Longin Jan Latecki
LREC/COLING5
2024 FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding
abstract
Flowcharts are graphical tools for representing complex concepts in concise visual representations. This paper introduces the FlowLearn dataset, a resource tailored to enhance the understanding of flowcharts. FlowLearn contains complex scientific flowcharts and simulated flowcharts. The scientific subset contains 3,858 flowcharts sourced from scientific literature and the simulated subset contains 10,000 flowcharts created using a customizable script. The dataset is enriched with annotations for visual components, OCR, Mermaid code representation, and VQA question-answer pairs. Despite the proven capabilities of Large Vision-Language Models (LVLMs) in various visual understanding tasks, their effectiveness in decoding flowcharts—a crucial element of scientific communication—has yet to be thoroughly investigated. The FlowLearn test set is crafted to assess the performance of LVLMs in flowchart comprehension. Our study thoroughly evaluates state-of-the-art LVLMs, identifying existing limitations and establishing a foundation for future enhancements in this relatively underexplored domain. For instance, in tasks involving simulated flowcharts, GPT-4V achieved the highest accuracy (58%) in counting the number of nodes, while Claude recorded the highest accuracy (83%) in OCR tasks. Notably, no single model excels in all tasks within the FlowLearn framework, highlighting significant opportunities for further development.
Huitong Pan, Cornelia Caragea, Eduard C. Dragut, Longin Jan Latecki
ECAI5
2024 SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents
abstract
Scientific information extraction (SciIE) is critical for converting unstructured knowledge from scholarly articles into structured data (entities and relations).Several datasets have been proposed for training and validating SciIE models.However, due to the high complexity and cost of annotating scientific texts, those datasets restrict their annotations to specific parts of paper, such as abstracts, resulting in the loss of diverse entity mentions and relations in context.In this paper, we release a new entity and relation extraction dataset for entities related to datasets, methods, and tasks in scientific articles.Our dataset contains 106 manually annotated full-text scientific publications with over 24k entities and 12k relations.To capture the intricate use and interactions among entities in full texts, our dataset contains a finegrained tag set for relations.Additionally, we provide an out-of-distribution test set to offer a more realistic evaluation.We conduct comprehensive experiments, including state-of-the-art supervised models and our proposed LLM baselines, and highlight the challenges presented by our dataset, encouraging the development of innovative models to further the field of SciIE.1
Zhijia Chen, Huitong Pan, Cornelia Caragea, Longin Jan Latecki, Eduard C. Dragut
EMNLP5
2024 Depth-Guided Dominant Plane Perception for Unsupervised Homography Estimation
abstract
Homography describes the mapping relations of the same plane across views. In scenarios with multiple planes, single homography estimation aims to obtain the optimal solution generated by the largest consistent plane to obey the coplanar constraints. However, existing methods typically consider all planes equally, neglecting the negative impact of regions that differ significantly from the largest approximate planar areas (dominant plane). In this work, we propose a depth-guided dominant plane perception network to achieve unsupervised homography estimation with additional attention on the dominant plane. Specifically, we leverage the depth-wise prior to adaptively detecting the approximate dominant plane, invoking essential scene structures for unsupervised homography estimation. Then, we enhance the corresponding features of the dominant plane and explore their correlations through a specially designed perceptual module. Finally, we employ dominant plane perception on multi-scale features progressively to estimate the homography in a coarse-to-fine manner. Extensive experiments on a large parallax dataset demonstrate that our method improves the alignment performance by 10.29%, yielding more accurate alignment than previous competitive methods.
Xiaomei Feng, Qi Jia 0001, Yu Liu 0012, Xin Fan 0001, Longin Jan Latecki
ICASSP5
2024 Fuzzy Boundary-Guided Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is a challenging task that identifies camouflaged objects from highly similar backgrounds. Existing methods typically treat the whole object equally while neglecting the indistinguishable regions that require more attention than other regions. In this paper, we propose a Fuzzy Boundary-Guided Network (FBG-Net) for camouflaged object detection, which mimics the human behavior that pays more attention to these low-confidence regions when observing objects. Specifically, we devise two main building blocks: (1) Mixed Semantics Aggregation Module (MSAM) to integrate boundary and texture features cumulatively in the high-to-low scales, and (2) Fuzzy Boundary-Guided Module (FBGM) to locate and enhance the low-confidence regions under the guidance of fuzzy boundary. Extensive experiments demonstrate the effectiveness of FBG-Net with superior performance to existing state-of-the-art methods. Code is available at https://github.com/YAOSL98/FBG-Net.
Qi Jia 0001, Shuilian Yao, Youcan Xu, Yu Liu 0012, Dehao Kong, Longin Jan Latecki
ICME6
2024 Learning Object Focused Attention
Vivek Trivedy, Amani Almalki, Longin Jan Latecki
ICPR (9)3
2024 Image Retrieval with Self-Supervised Divergence Minimization and Cross-Attention Classification
Vivek Trivedy, Longin Jan Latecki
IJCAI2
2024 Self-Supervised Learning with Masked Autoencoders for Teeth Segmentation from Intra-oral 3D Scans
abstract
In modern dentistry, teeth localization, segmentation, and labeling from intra-oral 3D scans are crucial for improving dental diagnostics, treatment planning, and population-based studies on oral health. However, creating automated algorithms for teeth analysis is a challenging task due to the limited availability of accessible data for training, particularly from the point of view of deep learning. This study extends the self-supervised learning framework of the mesh masked autoencoder (MeshMAE) transformer. While the MeshMAE loss measures the quality of reconstructed masked mesh triangles, the loss of the proposed DentalMAE evaluates the predicted deep embeddings of masked mesh triangles. This yields a better generalization ability on a very limited number of 3D dental scans, as documented by our results on teeth segmentation of intra-oral scans. Our results show that masking-based unsupervised learning methods may, for the first time, provide convincing transfer learning improvements on 3D intra-oral scans, increasing the overall accuracy over both MeshMAE and prior self-supervised pre-training.
Amani Almalki, Longin Jan Latecki
WACV2
2024 Rank-based Hashing for Effective and Efficient Nearest Neighbor Search for Image Retrieval
abstract
The large and growing amount of digital data creates a pressing need for approaches capable of indexing and retrieving multimedia content. A traditional and fundamental challenge consists of effectively and efficiently performing nearest-neighbor searches. After decades of research, several different methods are available, including trees, hashing, and graph-based approaches. Most of the current methods exploit learning to hash approaches based on deep learning. In spite of effective results and compact codes obtained, such methods often require a significant amount of labeled data for training. Unsupervised approaches also rely on expensive training procedures usually based on a huge amount of data. In this work, we propose an unsupervised data-independent approach for nearest neighbor searches, which can be used with different features, including deep features trained by transfer learning. The method uses a rank-based formulation and exploits a hashing approach for efficient ranked list computation at query time. A comprehensive experimental evaluation was conducted on seven public datasets, considering deep features based on CNNs and Transformers. Both effectiveness and efficiency aspects were evaluated. The proposed approach achieves remarkable results in comparison to traditional and state-of-the-art methods. Hence, it is an attractive and innovative solution, especially when costly training procedures need to be avoided.
Vinicius Atsushi Sato Kawai, Lucas Pascotti Valem, Alexandro Baldassin, Edson Borin, Daniel C. G. Pedronette, Longin Jan Latecki
ACM Trans. Multim. Comput. Commun. Appl.6
2024 A rotation robust shape transformer for cartoon character recognition
Qi Jia 0001, Yi Wang 0037, Xin Fan 0001, Haibin Ling, Longin Jan Latecki
Vis. Comput.6
2023 Strokes Trajectory Recovery for Unconstrained Handwritten Documents with Automatic Evaluation
Sidra Hanif, Longin Jan Latecki
ICPRAM2
2023 Learning Pixel-wise Alignment for Unsupervised Image Stitching
abstract
Image stitching aims to align a pair of images in the same view. Generating precise alignment with natural structures is challenging for image stitching, as there is no wider field-of-view image as a reference, especially in non-coplanar practical scenarios. In this paper, we propose an unsupervised image stitching framework, breaking through the coplanar constraints in homography estimation, yielding accurate pixel-wise alignment under limited overlapping regions. First, we generate a global transformation by an iterative dense feature matching combined with an error control strategy to alleviate the difference introduced by large parallax. Second, we propose a pixel-wise warping network embedded within a large-scale feature extractor and a correlative feature enhancement module to explicitly learn correspondences between the inputs, and generate accurate pixel-level offsets upon novel constraints on both overlapping and non-overlapping regions. Notably, we leverage the pixel-level offsets in the overlapping area to guide the adjustment in the non-overlapping area upon content and structure consistency constraints, rendering a natural transition between two regions and distortions suppression over the entire stitched image. The proposed method achieves state-of-the-art performance that surpasses both traditional and deep learning approaches by a large margin. It also achieves the shortest execution time and has the best generalization ability on the traditional dataset.
Qi Jia 0001, Xiaomei Feng, Yu Liu 0012, Xin Fan 0001, Longin Jan Latecki
ACM Multimedia5
2023 Self-Supervised Learning with Masked Image Modeling for Teeth Numbering, Detection of Dental Restorations, and Instance Segmentation in Dental Panoramic Radiographs
abstract
The computer-assisted radiologic informative report is currently emerging in dental practice to facilitate dental care and reduce time consumption in manual panoramic radiographic interpretation. However, the amount of dental radiographs for training is very limited, particularly from the point of view of deep learning. This study aims to utilize recent self-supervised learning methods like SimMIM and UM-MAE to increase the model efficiency and understanding of the limited number of dental radiographs. We use the Swin Transformer for teeth numbering, detection of dental restorations, and instance segmentation tasks. To the best of our knowledge, this is the first study that applied self-supervised learning methods to Swin Transformer on dental panoramic radiographs. Our results show that the SimMIM method obtained the highest performance of 90.4% and 88.9% on detecting teeth and dental restorations and instance segmentation, respectively, increasing the average precision by 13.4 and 12.8 over the random initialization baseline. Moreover, we augment and correct the existing dataset of panoramic radiographs. The code and the dataset are available at https://github.com/AmaniHAlmalki/DentalMIM.
Amani Almalki, Longin Jan Latecki
WACV2
2023 CNN2Graph: Building Graphs for Image Classification
abstract
Neural Network classifiers generally operate via the i.i.d. assumption where examples are passed through independently during training. We propose CNN2GNN and CNN2Transformer which instead leverage inter-example information for classification. We use Graph Neural Networks (GNNs) to build a latent space bipartite graph and compute cross-attention scores between input images and a proxy set. Our approach addresses several challenges of existing methods. Firstly, it is end-to-end differentiable despite the generally discrete nature of graph construction. Secondly, it allows inductive inference at no extra cost. Thirdly, it presents a simple method to construct graphs from arbitrary datasets that captures both example level and class level information. Finally, it addresses the proxy collapse problem by combining contrastive and cross-entropy losses rather than separate clustering algorithms. Our results increase classification performance over baseline experiments and outperform other methods. We also conduct an empirical investigation showing that Transformer style attention scales better than GAT attention with dataset size.
Vivek Trivedy, Longin Jan Latecki
WACV2
2023 Graph Convolutional Networks based on manifold learning for semi-supervised image classification
Lucas Pascotti Valem, Daniel C. G. Pedronette, Longin Jan Latecki
Comput. Vis. Image Underst.3
2023 DMDD: A Large-Scale Dataset for Dataset Mentions Detection
abstract
Abstract The recognition of dataset names is a critical task for automatic information extraction in scientific literature, enabling researchers to understand and identify research opportunities. However, existing corpora for dataset mention detection are limited in size and naming diversity. In this paper, we introduce the Dataset Mentions Detection Dataset (DMDD), the largest publicly available corpus for this task. DMDD consists of the DMDD main corpus, comprising 31,219 scientific articles with over 449,000 dataset mentions weakly annotated in the format of in-text spans, and an evaluation set, which comprises 450 scientific articles manually annotated for evaluation purposes. We use DMDD to establish baseline performance for dataset mention detection and linking. By analyzing the performance of various models on DMDD, we are able to identify open problems in dataset mention detection. We invite the community to use our dataset as a challenge to develop novel dataset mention detection models.
Huitong Pan, Eduard C. Dragut, Cornelia Caragea, Longin Jan Latecki
Trans. Assoc. Comput. Linguistics5
2023 Characteristic Mapping for Ellipse Detection Acceleration
abstract
It is challenging to characterize the intrinsic geometry of high-degree algebraic curves with lower-degree algebraic curves. The reduction in the curve's degree implies lower computation costs, which is crucial for various practical computer vision systems. In this paper, we develop a characteristic mapping (CM) to recursively degenerate 3n points on a planar curve of n th order to 3(n-1) points on a curve of (n-1) th order. The proposed characteristic mapping enables curve grouping on a line, a curve of the lowest order, that preserves the intrinsic geometric properties of a higher-order curve (ellipse). We prove a necessary condition and derive an efficient arc grouping module that finds valid elliptical arc segments by determining whether the mapped three points are colinear, invoking minimal computation. We embed the module into two latest arc-based ellipse detection methods, which reduces their running time by 25% and 50% on average over five widely used data sets. This yields faster detection than the state-of-the-art algorithms while keeping their precision comparable or even higher. Two CM embedded methods also significantly surpass a deep learning method on all evaluation metrics.
Qi Jia 0001, Xin Fan 0001, Yang Yang 0120, Xuxu Liu, Zhongxuan Luo, Xinchen Zhou, Longin Jan Latecki
IEEE Trans. Image Process.8
2023 Rank Flow Embedding for Unsupervised and Semi-Supervised Manifold Learning
abstract
Impressive advances in acquisition and sharing technologies have made the growth of multimedia collections and their applications almost unlimited. However, the opposite is true for the availability of labeled data, which is needed for supervised training, since such data is often expensive and time-consuming to obtain. While there is a pressing need for the development of effective retrieval and classification methods, the difficulties faced by supervised approaches highlight the relevance of methods capable of operating with few or no labeled data. In this work, we propose a novel manifold learning algorithm named Rank Flow Embedding (RFE) for unsupervised and semi-supervised scenarios. The proposed method is based on ideas recently exploited by manifold learning approaches, which include hypergraphs, Cartesian products, and connected components. The algorithm computes context-sensitive embeddings, which are refined following a rank-based processing flow, while complementary contextual information is incorporated. The generated embeddings can be exploited for more effective unsupervised retrieval or semi-supervised classification based on Graph Convolutional Networks. Experimental results were conducted on 10 different collections. Various features were considered, including the ones obtained with recent Convolutional Neural Networks (CNN) and Vision Transformer (ViT) models. High effective results demonstrate the effectiveness of the proposed method on different tasks: unsupervised image retrieval, semi-supervised classification, and person Re-ID. The results demonstrate that RFE is competitive or superior to the state-of-the-art in diverse evaluated scenarios.
Lucas Pascotti Valem, Daniel C. G. Pedronette, Longin Jan Latecki
IEEE Trans. Image Process.3
2023 Entropy Minimization Versus Diversity Maximization for Domain Adaptation
abstract
Entropy minimization has been widely used in unsupervised domain adaptation (UDA). However, existing works reveal that the use of entropy-minimization-only may lead to collapsed trivial solutions for UDA. In this article, we try to seek possible close-to-ideal UDA solutions by focusing on some intuitive properties of the ideal domain adaptation solution. In particular, we propose to introduce diversity maximization for further regulating entropy minimization. In order to achieve the possible minimum target risk for UDA, we show that diversity maximization should be elaborately balanced with entropy minimization, the degree of which can be finely controlled with the use of deep embedded validation in an unsupervised manner. The proposed minimal-entropy diversity maximization (MEDM) can be directly implemented by stochastic gradient descent without the use of adversarial learning. Empirical evidence demonstrates that MEDM outperforms the state-of-the-art methods on four popular domain adaptation datasets.
Xiaofu Wu, Suofei Zhang, Quan Zhou 0004, Zhen Yang 0001, Chunming Zhao 0001, Longin Jan Latecki
IEEE Trans. Neural Networks Learn. Syst.6
2022 DPNET: Dual-Path Network for Efficient Object Detection with Lightweight Self-Attention
abstract
Object detection often costs a considerable amount of computation to get satisfied performance, which is unfriendly to be deployed in edge devices. To address the trade-off be-tween computational cost and detection accuracy, this paper presents a dual path network, named DPNet, for efficient object detection with lightweight self-attention. In backbone, a single input/output lightweight self-attention module (LSAM) is designed to encode global interactions between different positions. LSAM is also extended into a multiple-inputs version in feature pyramid network (FPN), which is employed to capture cross-resolution dependencies in two paths. Extensive experiments on the COCO dataset demonstrate that our method achieves promising detection results. More specifically, DPNet obtains 29.0% AP on COCO test-dev, with only 1.14 GFLOPs and 2.27M model size for a 320 × 320 image.
Huimin Shi, Quan Zhou 0004, Yinghao Ni, Xiaofu Wu, Longin Jan Latecki
ICIP5
2022 DRBANET: A Lightweight Dual-Resolution Network for Semantic Segmentation with Boundary Auxiliary
abstract
Due to the powerful ability to encode image details and semantics, many lightweight dual-resolution networks have been proposed in recent years. However, most of them ignore the benefit of boundary information. This paper introduces a lightweight dual-resolution network, called DRBANet, aiming to refine semantic segmentation results with the aid of boundary information. DRBANet also adopts dual parallel architecture, including: high resolution branch (HRB) and low resolution branch (LRB). Specifically, HRB mainly consists of a set of Efficient Inverted Bottleneck Modules (EIBMs), which learn feature representations with larger receptive fields. LRB is composed of a series of EIBMs and an Extremely Lightweight Pyramid Pooling Module (ELPPM), where ELPPM is utilized to capture multi-scale context through hierarchical residual connections. Finally, a boundary supervision head is designed to capture object boundaries in HRB. Extensive experiments on Cityscapes and CamVid datasets demonstrate that our method achieves promising trade-off between segmentation accuracy and running efficiency.
Quan Zhou 0004, Chenfeng Jiang, Xiaofu Wu, Longin Jan Latecki
ICIP5
2022 Contextual ensemble network for semantic segmentation
Quan Zhou 0004, Xiaofu Wu, Suofei Zhang, Bin Kang, ZongYuan Ge, Longin Jan Latecki
Pattern Recognit.6
2022 BANet: Boundary-Assistant Encoder-Decoder Network for Semantic Segmentation
abstract
Recently, boundary information has gained great attraction for semantic segmentation. This paper presents a novel encoder-decoder network, called BANet, for accurate semantic segmentation, where boundary information is employed as an additional assistance for producing more consistent segmentation outputs. BANet is composed of three components: the pre-trained backbone using dilated-ResNet101, semantic flow branch (SFB) and boundary flow branch (BFB) for semantic segmentation and boundary detection, respectively. More specifically, to delineate more accurate object shapes and boundaries, a global attention block (GAB) is designed in SFB as global guidance for high-level feature. On the other hand, BFB directly extracts features on boundaries, avoiding the unexpected interference from the non-boundary parts. Finally, we adopt a joint loss function to further optimize the segmentation results and boundary outputs synchronously. Moreover, compared with previous state-of-the-art methods, e.g., non-local block and ASPP module, our BFB leverages detection accuracy and computational efficiency in a lightweight fashion. To evaluate BANet, we have conducted extensive experiments on several semantic segmentation datasets: Cityscapes, PASCAL Context, and ADE20K. The experimental results show that, with the aid of boundary information, BANet is able to produce more consistent segmentation predictions with accurately delineated object shapes and boundaries, leading to the state-of-the-art performance on Cityscapes, and competitive results on PASCAL Context and ADE20K with respect to recent semantic segmentation networks.
Quan Zhou 0004, Yong Qiang, Yuwei Mo, Xiaofu Wu, Longin Jan Latecki
IEEE Trans. Intell. Transp. Syst.5
2021 Leveraging Line-Point Consistence To Preserve Structures for Wide Parallax Image Stitching
abstract
Generating high-quality stitched images with natural structures is a challenging task in computer vision. In this paper, we succeed in preserving both local and global geometric structures for wide parallax images, while reducing artifacts and distortions. A projective invariant, Characteristic Number, is used to match co-planar local sub-regions for input images. The homography between these well-matched sub-regions produces consistent line and point pairs, suppressing artifacts in overlapping areas. We explore and introduce global collinear structures into an objective function to specify and balance the desired characters for image warping, which can preserve both local and global structures while alleviating distortions. We also develop comprehensive measures for stitching quality to quantify the collinearity of points and the discrepancy of matched line pairs by considering the sensitivity to linear structures for human vision. Extensive experiments demonstrate the superior performance of the proposed method over the state-of-the-art by presenting sharp textures and preserving prominent natural structures in stitched images. Especially, our method not only exhibits lower errors but also the least divergence across all test images. Code is available at https://github.com/dut-media-lab/Image-Stitching.
Qi Jia 0001, Zhengjun Li, Xin Fan 0001, Shiyu Teng, Xinchen Ye, Longin Jan Latecki
CVPR7
2021 Rank-based self-training for graph convolutional networks
Daniel C. G. Pedronette, Longin Jan Latecki
Inf. Process. Manag.2
2020 Combination Of Handcrafted And Deep Learning-Based Features For 3d Mesh Quality Assessment
abstract
We propose in this paper a novel objective method to evaluate the perceived visual quality of 3D meshes. The proposed method in no-reference, it relies only on the distorted mesh for the quality estimation. It is based on a pre-trained convolutional neural network (i.e VGG to extract features from the distorted mesh) and handcrafted features extracted directly from the 3D mesh (i.e curvature and dihedral angle). A General Regression Neural Network (GRNN) is used to learn the statistical parameters of the feature vectors and estimate the quality score. Experimental results from for subjective databases (LIRIS masking, LIRIS/EPFL generalpurpose, UWB compression and LEETA simplification) and comparisons with objective metrics cited in the state-of-the-art demonstrate the efficacy of the proposed metric in terms of the correlation to the mean opinion scores across these databases.
Ilyass Abouelaziz, Aladine Chetouani, Mohammed El Hassouni, Longin Jan Latecki, Hocine Cherifi
ICIP4
2020 DCM: A Dense-Attention Context Module For Semantic Segmentation
abstract
For image semantic segmentation, a fully convolutional network is usually employed as the encoder to abstract visual features of the input image. A meticulously designed decoder is used to decoding the final feature map of the backbone. The output resolution of backbones which are designed for image classification task is too low to match segmentation task. Most existing methods for obtaining the final high-resolution feature map can not fully utilize the information of different layers of the backbone. To adequately extract the information of a single layer, the multi-scale context information of different layers, and the global information of backbone, we present a new attention-augmented module named Dense-attention Context Module (DCM), which is used to connect the common backbones and the other decoding heads. The experiments show the promising results of our method on Cityscapes dataset.
Shenghua Li, Quan Zhou 0004, Jie Wang 0024, Yawen Fan, Xiaofu Wu, Longin Jan Latecki
ICIP7
2020 Coupling Deep Textural and Shape Features for Sketch Recognition
abstract
Recognizing freehand sketches with high arbitrariness is such a great challenge that the automatic recognition rate has reached a ceiling in recent years. In this paper, we explicitly explore the shape properties of sketches, which has almost been neglected before in the context of deep learning, and propose a sequential dual learning strategy that combines both shape and texture features. We devise a two-stage recurrent neural network to balance these two types of features. Our architecture also considers stroke orders of sketches to reduce the intra-class variations of input features. Extensive experiments on the TU-Berlin benchmark set show that our method achieves over 90% recognition rate for the first time on this task, outperforming both humans and state-of-the-art algorithms by over 19 and 7.5 percentage points, respectively. Especially, our approach can distinguish the sketches with similar textures but different shapes more effectively than recent deep networks. Based on the proposed method, we develop an on-line sketch retrieval and imitation application to teach children or adults to draw. The application is available as Sketch.Draw.
Qi Jia 0001, Xin Fan 0001, Meiyu Yu, Yuqing Liu 0001, Dingrong Wang, Longin Jan Latecki
ACM Multimedia6
2020 Learning adaptive contrast combinations for visual saliency detection
Quan Zhou 0004, Huimin Lu 0001, Yawen Fan, Suofei Zhang, Xiaofu Wu, Baoyu Zheng, Weihua Ou, Longin Jan Latecki
Multim. Tools Appl.9
2020 3D visual saliency and convolutional neural network for blind mesh quality assessment
Ilyass Abouelaziz, Aladine Chetouani, Mohammed El Hassouni, Longin Jan Latecki, Hocine Cherifi
Neural Comput. Appl.4
2020 No-reference mesh visual quality assessment via ensemble of convolutional neural networks and compact multi-linear pooling
Ilyass Abouelaziz, Aladine Chetouani, Mohammed El Hassouni, Longin Jan Latecki, Hocine Cherifi
Pattern Recognit.4
2019 Re-Ranking via Metric Fusion for Object Retrieval and Person Re-Identification
abstract
This work studies the unsupervised re-ranking procedure for object retrieval and person re-identification with a specific concentration on an ensemble of multiple metrics (or similarities). While the re-ranking step is involved by running a diffusion process on the underlying data manifolds, the fusion step can leverage the complementarity of multiple metrics. We give a comprehensive summary of existing fusion with diffusion strategies, and systematically analyze their pros and cons. Based on the analysis, we propose a unified yet robust algorithm which inherits their advantages and discards their disadvantages. Hence, we call it Unified Ensemble Diffusion (UED). More interestingly, we derive that the inherited properties indeed stem from a theoretical framework, where the relevant works can be elegantly summarized as special cases of UED by imposing additional constraints on the objective function and varying the solver of similarity propagation. Extensive experiments with 3D shape retrieval, image retrieval and person re-identification demonstrate that the proposed framework outperforms the state of the arts, and at the same time suggest that re-ranking via metric fusion is a promising tool to further improve the retrieval performance of existing algorithms.
Song Bai 0001, Peng Tang 0005, Philip Torr 0001, Longin Jan Latecki
CVPR4
2019 Efficient Local Search for Minimum Dominating Sets in Large Graphs
Yi Fan 0001, Yongxuan Lai, Chengqian Li, Nan Li 0021, Zongjie Ma, Jun Zhou 0001, Longin Jan Latecki, Kaile Su
DASFAA (2)7
2019 Lednet: A Lightweight Encoder-Decoder Network for Real-Time Semantic Segmentation
abstract
The extensive computational burden limits the usage of CNNs in mobile devices for dense estimation tasks. In this paper, we present a lightweight network to address this problem, namely LEDNet, which employs an asymmetric encoder-decoder architecture for the task of real-time semantic segmentation. More specifically, the encoder adopts a ResNet as backbone network, where two new operations, channel split and shuffle, are utilized in each residual block to greatly reduce computation cost while maintaining higher segmentation accuracy. On the other hand, an attention pyramid network (APN) is employed in the decoder to further lighten the entire network complexity. Our model has less than 1M parameters, and is able to run at over 71 FPS in a single GTX 1080Ti GPU. The comprehensive experiments demonstrate that our approach achieves state-of-the-art results in terms of speed and accuracy trade-off on CityScapes dataset.
Yu Wang 0109, Quan Zhou 0004, Jian Xiong 0005, Guangwei Gao, Xiaofu Wu, Longin Jan Latecki
ICIP7
2019 Image Retrieval with Similar Object Detection and Local Similarity to Detected Objects
Sidra Hanif, Chao Li 0007, Anis Alazzawe, Longin Jan Latecki
PRICAI (3)4
2019 Scene Parsing Via Dense Recurrent Neural Networks With Attentional Selection
abstract
Recurrent neural networks (RNNs) have shown the ability to improve scene parsing through capturing long-range dependencies among image units. In this paper, we propose dense RNNs for scene labeling by exploring various long-range semantic dependencies among image units. Different from existing RNN based approaches, our dense RNNs are able to capture richer contextual dependencies for each image unit by enabling immediate connections between each pair of image units, which significantly enhances their discriminative power. Besides, to select relevant dependencies and meanwhile to restrain irrelevant ones for each unit from dense connections, we introduce an attention model into dense RNNs. The attention model allows automatically assigning more importance to helpful dependencies while less weight to unconcerned dependencies. Integrating with convolutional neural networks (CNNs), we develop an end-to-end scene labeling system. Extensive experiments on three large-scale benchmarks demonstrate that the proposed approach can improve the baselines by large margins and outperform other state-of-the-art algorithms.
Heng Fan 0001, Peng Chu, Longin Jan Latecki, Haibin Ling
WACV3
2019 An open-source project for real-time image semantic segmentation
Quan Zhou 0004, Yu Wang 0109, Xin Jin 0015, Longin Jan Latecki
Sci. China Inf. Sci.5
2019 Weakly supervised mitosis detection in breast histopathology images using concentric loss
Chao Li 0007, Xinggang Wang, Wenyu Liu 0001, Longin Jan Latecki, Bo Wang 0044, Junzhou Huang
Medical Image Anal.4
2019 Regularized Diffusion Process on Bidirectional Context for Object Retrieval
abstract
Diffusion process has advanced object retrieval greatly as it can capture the underlying manifold structure. Recent studies have experimentally demonstrated that tensor product diffusion can better reveal the intrinsic relationship between objects than other variants. However, the principle remains unclear, i.e., what kind of manifold structure is captured. In this paper, we propose a new affinity learning algorithm called Regularized Diffusion Process (RDP). By deeply exploring the properties of RDP, our first yet basic contribution is providing a manifold-based explanation for tensor product diffusion. A novel criterion measuring the smoothness of the manifold is defined, which simultaneously regularizes four vertices in the affinity graph. Inspired by this observation, we further contribute two variants towards two specific goals. While ARDP can learn similarities across heterogeneous domains, HRDP performs affinity learning on tensor product hypergraph, considering the relationships between objects are generally more complex than pairwise. Consequently, RDP, ARDP and HRDP constitute a generic tool for object retrieval in most commonly-used settings, no matter the input relationships between objects are derived from the same domain or not, and in pairwise formulation or not. Comprehensive experiments on 10 retrieval benchmarks, especially on large scale data, validate the effectiveness and generalization of our work.
Song Bai 0001, Xiang Bai, Qi Tian 0001, Longin Jan Latecki
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Training convolutional neural network from multi-domain contour images for 3D shape retrieval
Zongxiao Zhu, Cong Rao, Song Bai 0001, Longin Jan Latecki
Pattern Recognit. Lett.4
2019 Automatic Ensemble Diffusion for 3D Shape and Image Retrieval
abstract
As a post-processing procedure, the diffusion process has demonstrated its ability of substantially improving the performance of various visual retrieval systems. Whereas, great efforts are also devoted to similarity (or metric) fusion, seeing that only one individual type of similarity cannot fully reveal the intrinsic relationship between objects. This stimulates a great research interest of considering similarity fusion in the framework of the diffusion process (i.e., fusion with diffusion) for robust retrieval. In this paper, we first revisit representative methods about fusion with diffusion and provide new insights which are ignored by previous researchers. Then, observing that existing algorithms are susceptible to noisy similarities, the proposed regularized ensemble diffusion (RED) is bundled with an automatic weight learning paradigm, so that the negative impacts of noisy similarities are suppressed. Though formulated as a convex optimization problem, one advantage of RED is that it converts back into the iteration-based solver with the same computational complexity as the conventional diffusion process. At last, we integrate several recently-proposed similarities with the proposed framework. The experimental results suggest that we can achieve new state-of-the-art performances on various retrieval tasks, including 3D shape retrieval on the ModelNet data set, and image retrieval on the Holidays and Ukbench data sets.
Song Bai 0001, Jingdong Wang 0001, Xiang Bai, Longin Jan Latecki, Qi Tian 0001
IEEE Trans. Image Process.5
2019 Multi-scale deep context convolutional neural networks for semantic segmentation
Quan Zhou 0004, Guangwei Gao, Weihua Ou, Huimin Lu 0001, Longin Jan Latecki
World Wide Web7
2018 Convolutional Neural Network for Blind Mesh Visual Quality Assessment Using 3D Visual Saliency
abstract
In this work, we propose a convolutional neural network (CNN) framework to estimate the perceived visual quality of 3D meshes without having access to the reference. The proposed CNN architecture is fed by small patches selected carefully according to their level of saliency. To do so, the visual saliency of the 3D mesh is computed, then we render 2D projections from the 3D mesh and its corresponding 3D saliency map. Afterward, the obtained views are split to obtain 2D small patches that pass through a saliency filter to select the most relevant patches. Experiments are conducted on two MVQ assessment databases, and the results show that the trained CNN achieves good rates in terms of correlation with human judgment.
Ilyass Abouelaziz, Aladine Chetouani, Mohammed El Hassouni, Longin Jan Latecki, Hocine Cherifi
ICIP4
2018 Dense Deconvolutional Network for Semantic Segmentation
abstract
Recently, exploring multiple feature maps from different layers in fully convolutional networks (FCNs) has gained substantial attention to capture context information for semantic segmentation. This paper presents a novel encoder-decoder architecture, called dense deconvolutional network (DDN), for semantic segmentation, where the feature maps of deeper convolutional layers are densely upsampled for the shallow deconvolutional layers. The proposed DDN is trainable end-to-end, and allows us to fully investigate multiple scale context cues embedded in images. The experimental results show that our DDN outperforms previous FCNs and encoder-decoder networks (EDNs) on PASCAL VOC 2012 dataset.
Quan Zhou 0004, Jingnan Lu, Xiaofu Wu, Suofei Zhang, Longin Jan Latecki
ICIP6
2018 DeepMitosis: Mitosis detection via deep detection, verification and segmentation networks
Chao Li 0007, Xinggang Wang, Wenyu Liu 0001, Longin Jan Latecki
Medical Image Anal.4
2018 Face recognition via fast dense correspondence
Quan Zhou 0004, Wenbin Yu 0002, Yawen Fan, Hu Zhu, Xiaofu Wu, Weihua Ou, Wei-Ping Zhu 0001, Longin Jan Latecki
Multim. Tools Appl.9
2017 Multidimensional Scaling on Multiple Input Distance Matrices
abstract
Multidimensional Scaling (MDS) is a classic technique that seeks vectorial representations for data points, given the pairwise distances between them. In recent years, data are usually collected from diverse sources or have multiple heterogeneous representations. However, how to do multidimensional scaling on multiple input distance matrices is still unsolved to our best knowledge. In this paper, we first define this new task formally. Then, we propose a new algorithm called Multi-View Multidimensional Scaling (MVMDS) by considering each input distance matrix as one view. The proposed algorithm can learn the weights of views (i.e., distance matrices) automatically by exploring the consensus information and complementary nature of views. Experimental results on synthetic as well as real datasets demonstrate the effectiveness of MVMDS. We hope that our work encourages a wider consideration in many domains where MDS is needed.
Song Bai 0001, Xiang Bai, Longin Jan Latecki, Qi Tian 0001
AAAI3
2017 Regularized Diffusion Process for Visual Retrieval
abstract
Diffusion process has advanced visual retrieval greatly owing to its capacity in capturing the geometry structure of the underlying manifold. Recent studies (Donoser and Bischof 2013) have experimentally demonstrated that diffusion process on the tensor product graph yields better retrieval performances than that on the original affinity graph. However, the principle behind this kind of diffusion process remains unclear, i.e., what kind of manifold structure is captured and how it is reflected. In this paper, we propose a new variant o diffusion process, which also operates on a tensor product graph. It is defined in three equivalent formulations (regularization framework, iterative framework and limit framework, respectively). Based on our study, three insightful conclusions are drawn which theoretically explain how this kind of diffusion process can better reveal the intrinsic relationship between objects. Besides, extensive experimental results on various retrieval tasks testify the validity of the proposed method.
Song Bai 0001, Xiang Bai, Qi Tian 0001, Longin Jan Latecki
AAAI4
2017 Amodal Detection of 3D Objects: Inferring 3D Bounding Boxes from 2D Ones in RGB-Depth Images
abstract
This paper addresses the problem of amodal perception of 3D object detection. The task is to not only find object localizations in the 3D world, but also estimate their physical sizes and poses, even if only parts of them are visible in the RGB-D image. Recent approaches have attempted to harness point cloud from depth channel to exploit 3D features directly in the 3D space and demonstrated the superiority over traditional 2.5D representation approaches. We revisit the amodal 3D detection problem by sticking to the 2.5D representation framework, and directly relate 2.5D visual appearance to 3D objects. We propose a novel 3D object detection system that simultaneously predicts objects 3D locations, physical sizes, and orientations in indoor scenes. Experiments on the NYUV2 dataset show our algorithm significantly outperforms the state-of-the-art and indicates 2.5D representation is capable of encoding features for 3D amodal object detection. All source code and data is on https://github.com/phoenixnn/Amodal3Det.
Longin Jan Latecki
CVPR2
2017 Ensemble Diffusion for Retrieval
abstract
As a postprocessing procedure, diffusion process has demonstrated its ability of substantially improving the performance of various visual retrieval systems. Whereas, great efforts are also devoted to similarity (or metric) fusion, seeing that only one individual type of similarity cannot fully reveal the intrinsic relationship between objects. This stimulates a great research interest of considering similarity fusion in the framework of diffusion process (i.e., fusion with diffusion) for robust retrieval. In this paper, we firstly revisit representative methods about fusion with diffusion, and provide new insights which are ignored by previous researchers. Then, observing that existing algorithms are susceptible to noisy similarities, the proposed Regularized Ensemble Diffusion (RED) is bundled with an automatic weight learning paradigm, so that the negative impacts of noisy similarities are suppressed. At last, we integrate several recently-proposed similarities with the proposed framework. The experimental results suggest that we can achieve new state-of-the-art performances on various retrieval tasks, including 3D shape retrieval on ModelNet dataset, and image retrieval on Holidays and Ukbench dataset.
Song Bai 0001, Jingdong Wang 0001, Xiang Bai, Longin Jan Latecki, Qi Tian 0001
ICCV5
2017 Efficient Local Search for Maximum Weight Cliques in Large Graphs
abstract
In this paper, we develop a local search algorithm to solve the Maximum Weight Clique (MWC) problem. Firstly we design a novel scoring function to measure the benefits of a local move. Then we develop a Cycle Estimation based ReStart (CERS) strategy to resolve the cycling issue in the local search process. Experimental results show that our solver achieves state-of-the-art performances on the large sparse graphs as well as large dense graphs. Also we present a theorem which shows the necessity of the restart strategies in current state-of-the-art local search algorithms.
Yi Fan 0001, Zongjie Ma, Kaile Su, Chengqian Li, Cong Rao, Ren-Hau Liu, Longin Jan Latecki
ICTAI7
2017 Restart and Random Walk in Local Search for Maximum Vertex Weight Cliques with Evaluations in Clustering Aggregation
abstract
The Maximum Vertex Weight Clique (MVWC) problem is NP-hard and also important in real-world applications. In this paper we propose to use the restart and the random walk strategies to improve local search for MVWC. If a solution is revisited in some particular situation, the search will restart. In addition, when the local search has no other options except dropping vertices, it will use random walk. Experimental results show that our solver outperforms state-of-the-art solvers in DIMACS and finds a new best-known solution. Also it is the unique solver which is comparable with state-of-the-art methods on both BHOSLIB and large crafted graphs. Furthermore we evaluated our solver in clustering aggregation. Experimental results on a number of real data sets demonstrate that our solver outperforms the state-of-the-art for solving the derived MVWC problem and helps improve the final clustering results.
Yi Fan 0001, Nan Li 0021, Chengqian Li, Zongjie Ma, Longin Jan Latecki, Kaile Su
IJCAI5
2017 Affinity Learning for Mixed Data Clustering
abstract
In this paper, we propose a novel affinity learning based framework for mixed data clustering, which includes: how to process data with mixed-type attributes, how to learn affinities between data points, and how to exploit the learned affinities for clustering. In the proposed framework, each original data attribute is represented with several abstract objects defined according to the specific data type and values. Each attribute value is transformed into the initial affinities between the data point and the abstract objects of attribute. We refine these affinities and infer the unknown affinities between data points by taking into account the interconnections among the attribute values of all data points. The inferred affinities between data points can be exploited for clustering. Alternatively, the refined affinities between data points and the abstract objects of attributes can be transformed into new data features for clustering. Experimental results on many real world data sets demonstrate that the proposed framework is effective for mixed data clustering.
Nan Li 0021, Longin Jan Latecki
IJCAI2
2017 Unsupervised object region proposals for RGB-D indoor scenes
Sinisa Todorovic, Longin Jan Latecki
Comput. Vis. Image Underst.3
2017 GIFT: Towards Scalable 3D Shape Retrieval
abstract
Projective analysis is an important solution in three-dimensional (3D) shape retrieval, since human visual perceptions of 3D shapes rely on various 2D observations from different viewpoints. Although multiple informative and discriminative views are utilized, most projection-based retrieval systems suffer from heavy computational cost, and thus cannot satisfy the basic requirement of scalability for search engines. In the past three years, shape retrieval contest (SHREC) pays much attention to the scalability of 3D shape retrieval algorithms, and organizes several large scale tracks accordingly [1]- [3]. However, the experimental results indicate that conventional algorithms cannot be directly applied to large datasets. In this paper, we present a real-time 3D shape search engine based on the projective images of 3D shapes. The real-time property of our search engine results from the following aspects: (1) efficient projection and view feature extraction using GPU acceleration; (2) the first inverted file, called F-IF, is utilized to speed up the procedure of multiview matching; and (3) the second inverted file, which captures a local distribution of 3D shapes in the feature manifold, is adopted for efficient context-based reranking. As a result, for each query the retrieval task can be finished within one second despite the necessary cost of IO overhead. We name the proposed 3D shape search engine, which combines GPU acceleration and inverted file (t wice), as GIFT. Besides its high efficiency, GIFT also outperforms state-of-the-art methods significantly in retrieval accuracy on various shape benchmarks (ModelNet40 dataset, ModelNet10 dataset, PSB dataset, McGill dataset) and competitions (SHREC14LSGTB, ShapeNet Core55, WM-SHREC07).
Song Bai 0001, Xiang Bai, Zhaoxiang Zhang 0001, Qi Tian 0001, Longin Jan Latecki
IEEE Trans. Multim.6
2016 GIFT: A Real-Time and Scalable 3D Shape Search Engine
abstract
Projective analysis is an important solution for 3D shape retrieval, since human visual perceptions of 3D shapes rely on various 2D observations from different view points. Although multiple informative and discriminative views are utilized, most projection-based retrieval systems suffer from heavy computational cost, thus cannot satisfy the basic requirement of scalability for search engines. In this paper, we present a real-time 3D shape search engine based on the projective images of 3D shapes. The real-time property of our search engine results from the following aspects: (1) efficient projection and view feature extraction using GPU acceleration, (2) the first inverted file, referred as F-IF, is utilized to speed up the procedure of multi-view matching, (3) the second inverted file (S-IF), which captures a local distribution of 3D shapes in the feature manifold, is adopted for efficient context-based reranking. As a result, for each query the retrieval task can be finished within one second despite the necessary cost of IO overhead. We name the proposed 3D shape search engine, which combines GPU acceleration and Inverted File (Twice), as GIFT. Besides its high efficiency, GIFT also outperforms the state-of-the-art methods significantly in retrieval accuracy on various shape benchmarks and competitions.
Song Bai 0001, Xiang Bai, Zhaoxiang Zhang 0001, Longin Jan Latecki
CVPR5
2016 Semi-Supervised Learning on an Augmented Graph with Class Labels
abstract
In this paper, we propose a novel graph-based method for semi-supervised learning. Our method runs a diffusion-based affinity learning algorithm on an augmented graph consisting of not only the nodes of labeled and unlabeled data but also artificial nodes representing class labels. The learned affinities between unlabeled data and class labels are used for classification. Our method achieves superior results on many standard data sets.
Nan Li 0021, Longin Jan Latecki
ECAI2
2016 Context-regularized learning of fully convolutional networks for scene labeling
abstract
This paper addresses the problem of pixel-wise semantic labeling of images. To this end, we use a fully convolutional network (FCN) whose input are raw pixels, and output are pixel labels. Our key novelty is that we regularize a supervised learning of FCN, such that FCN correctly predicts pixel labels and additionally does not violate a given set of spatial object relationships of interest. The frequency of occurrence of these object relationships in training images is used to estimate a new loss function for the regularized learning of FCN. The results on the benchmark PASCAL 2011, 2012 and NYU v2 datasets demonstrate that our regularized FCN outperforms a non-regularized FCN and other related state-of-the-art approaches. Importantly, in cases of error in semantic labeling, the regularized FCN does not violate the object relationships of interest, unlike the non-regularized counterparts.
Sinisa Todorovic, Longin Jan Latecki
ICPR3
2016 Location-Aware Image Classification
Xinggang Wang, Xin Yang 0008, Wenyu Liu 0001, Chen Duan, Longin Jan Latecki
MMM (1)5
2016 Enhanced Affinity Inference Based Recommender Systems
abstract
In this paper, we focus on improving the prediction accuracy and scalability of affinity inference based recommender systems, which predict unknown ratings based on the inferred affinities between abstract objects of users and item-rating pairs. Instead of treating each item evenly when inferring affinities, we propose to take into account the variance in ratings of each item. We also propose to reduce the number of item-rating pairs defined for each item, which can significantly reduce the computing time of affinity inference. Experimental results on the standard MovieLens dataset demonstrate the improvements.
Nan Li 0021, Longin Jan Latecki
WI2
2016 Similarity Fusion for Visual Tracking
Yu Zhou 0016, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
Int. J. Comput. Vis.4
2016 Multi-scale context for scene labeling via flexible segmentation graph
Quan Zhou 0004, Baoyu Zheng, Wei-Ping Zhu 0001, Longin Jan Latecki
Pattern Recognit.4
2016 Efficient shape representation, matching, ranking, and its applications
Xiang Bai, Michael Donoser, Hairong Liu, Longin Jan Latecki
Pattern Recognit. Lett.4
2015 Integration of Single-view Graphs with Diffusion of Tensor Product Graphs for Multi-view Spectral Clustering
Le Shu, Longin Jan Latecki
ACML2
2015 Transductive Domain Adaptation with Affinity Learning
abstract
We study the problem of domain adaptation, which aims to adapt the classifiers trained on a labeled source domain to an unlabeled target domain. We propose a novel method to solve domain adaptation task in a transductive setting. The proposed method bridges the distribution gap between source domain and target domain through affinity learning. It exploits the existence of a subset of data points in target domain which distribute similarly to the data points in the source domain. These data points act as the bridge that facilitates the data similarities propagation across domains. We also propose to control the relative importance of intra- and inter-domain similarities to boost the similarity propagation. In our approach, we first construct the similarity matrix which encodes both the intra- and inter-domain similarities. We then learn the true similarities among data points in joint manifold using graph diffusion.
Le Shu, Longin Jan Latecki
CIKM2
2015 Shape Similarity Based on the Qualitative Spatial Reasoning Calculus eOPRAm
Christopher H. Dorr, Longin Jan Latecki, Reinhard Moratz
COSIT2
2015 Salient object detection via background contrast
abstract
This paper addresses the problem of salient object detection. We introduce a novel framework which aims to automatically identify salient regions in natural images based on two key ideas. The first one is to consider the statistical spatial distribution of saliency and non-saliency regions as two complementary processes. The second one is based on the assumption that contrast saliency with respect to background regions outperforms those with respect to entire image. Experimental results demonstrate the effectiveness of our approach over 12 state-of-the-art models.
Quan Zhou 0004, Nianyi Li, Shu Cai, Longin Jan Latecki
ICASSP5
2015 Semantic Segmentation of RGBD Images with Mutex Constraints
abstract
In this paper, we address the problem of semantic scene segmentation of RGB-D images of indoor scenes. We propose a novel image region labeling method which augments CRF formulation with hard mutual exclusion (mutex) constraints. This way our approach can make use of rich and accurate 3D geometric structure coming from Kinect in a principled manner. The final labeling result must satisfy all mutex constraints, which allows us to eliminate configurations that violate common sense physics laws like placing a floor above a night stand. Three classes of mutex constraints are proposed: global object co-occurrence constraint, relative height relationship constraint, and local support relationship constraint. We evaluate our approach on the NYU-Depth V2 dataset, which consists of 1449 cluttered indoor scenes, and also test generalization of our model trained on NYU-Depth V2 dataset directly on a recent SUN3D dataset without any new training. The experimental results show that we significantly outperform the state-of-the-art methods in scene labeling on both datasets.
Sinisa Todorovic, Longin Jan Latecki
ICCV3
2015 Sequential Monte Carlo for Maximum Weight Subgraphs with Application to Solving Image Jigsaw Puzzles
Nagesh Adluru, Xingwei Yang, Longin Jan Latecki
Int. J. Comput. Vis.3
2015 3D Shape Matching via Two Layer Coding
abstract
View-based 3D shape retrieval is a popular branch in 3D shape analysis owing to the high discriminative property of 2D views. However, many previous works do not scale up to large 3D shape databases. We propose a two layer coding (TLC) framework to conduct shape matching much more efficiently. The first layer coding is applied to pairs of views represented as depth images. The spatial relationship of each view pair is captured with so-called eigen-angle, which is the planar angle between the two views measured at the center of the 3D shape. Prior to the second layer coding, the view pairs are divided into subsets according to their eigen-angles. Consequently, view pairs that differ significantly in their eigen-angles are encoded with different codewords, which implies that spatial arrangement of views is preserved in the second layer coding. The final feature vector of a 3D shape is the concatenation of all the encoded features from different subsets, which is used for efficient indexing directly. TLC is not limited to encode the local features from 2D views, but can be also applied to encoding 3D features. Exhaustive experimental results confirm that TLC achieves state-of-the-art performance in both retrieval accuracy and efficiency.
Xiang Bai, Song Bai 0001, Zhuotun Zhu, Longin Jan Latecki
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Dense Subgraph Partition of Positive Hypergraphs
abstract
In this paper, we present a novel partition framework, called dense subgraph partition (DSP), to automatically, precisely and efficiently decompose a positive hypergraph into dense subgraphs. A positive hypergraph is a graph or hypergraph whose edges, except self-loops, have positive weights. We first define the concepts of core subgraph, conditional core subgraph, and disjoint partition of a conditional core subgraph, then define DSP based on them. The result of DSP is an ordered list of dense subgraphs with decreasing densities, which uncovers all underlying clusters, as well as outliers. A divide-and-conquer algorithm, called min-partition evolution, is proposed to efficiently compute the partition. DSP has many appealing properties. First, it is a nonparametric partition and it reveals all meaningful clusters in a bottom-up way. Second, it has an exact and efficient solution, called min-partition evolution algorithm. The min-partition evolution algorithm is a divide-and-conquer algorithm, thus time-efficient and memory-friendly, and suitable for parallel processing. Third, it is a unified partition framework for a broad range of graphs and hypergraphs. We also establish its relationship with the densest k-subgraph problem (DkS), an NP-hard but fundamental problem in graph theory, and prove that DSP gives precise solutions to DkS for all kin a graph-dependent set, called critical k-set. To our best knowledge, this is a strong result which has not been reported before. Moreover, as our experimental results show, for sparse graphs, especially web graphs, the size of critical k-set is close to the number of vertices in the graph. We test the proposed partition framework on various tasks, and the experimental results clearly illustrate its advantages.
Hairong Liu, Longin Jan Latecki, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Locality Preserving Projection for Domain Adaptation with Multi-Objective Learning
abstract
In many practical cases, we need to generalize a model trained in a source domain to a new target domain.However, the distribution of these two domains may differ very significantly, especially sometimes some crucial target features may not have support in the source domain.This paper proposes a novel locality preserving projection method for domain adaptation task,which can find a linear mapping preserving the 'intrinsic structure' for both source and target domains.We first construct two graphs encoding the neighborhood information for source and target domains separately.We then find linear projection coefficients which have the property of locality preserving for each graph.Instead of combing the two objective terms under compatibility assumption and requiring the user to decide the importance of each objective function,we propose a multi-objective formulation for this problem and solve it simultaneously using Pareto optimization.The Pareto frontier captures all possible good linear projection coefficients that are preferred by one or more objectives.The effectiveness of our approach is justified by both theoretical analysis and empirical results on real world data sets.The new feature representation shows better prediction accuracy as our experiments demonstrate.
Le Shu, Tianyang Ma, Longin Jan Latecki
AAAI3
2014 Unsupervised Segmentation of RGB-D Images
Longin Jan Latecki
ACCV (3)2
2014 Human Detection Using Learned Part Alphabet and Pose Dictionary
Cong Yao, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
ECCV (5)4
2014 3D object retrieval by 3D curve matching
abstract
In this paper, we introduce a novel approach to 3D object retrieval by 3D curve matching. First, we project 2D object edges obtained from a depth image into 3D space. Second, we find distinctive feature points on the object. Third, we represent the shortest paths between the features by robust descriptors invariant to rotation, scaling, and translation. Finally, we match two 3D objects using the Maximum Weight Subgraph search. The most important contribution of this paper is the powerful object representation by 3D curves together with the corresponding matching algorithm. Excellent retrieval results achieved with our method show its benefits compared to the state-of-the-art.
Christian Feinen, Joanna Czajkowska, Marcin Grzegorzek, Longin Jan Latecki
ICIP4
2014 Online Multiple targets Detection and Tracking from Mobile robot in Cluttered indoor Environments with Depth Camera
abstract
Indoor environment is a common scene in our everyday life, and detecting and tracking multiple targets in this environment is a key component for many applications. However, this task still remains challenging due to limited space, intrinsic target appearance variation, e.g. full or partial occlusion, large pose deformation, and scale change. In the proposed approach, we give a novel framework for detection and tracking in indoor environments, and extend it to robot navigation. One of the key components of our approach is a virtual top view created from an RGB-D camera, which is named ground plane projection (GPP). The key advantage of using GPP is the fact that the intrinsic target appearance variation and extrinsic noise is far less likely to appear in GPP than in a regular side-view image. Moreover, it is a very simple task to determine free space in GPP without any appearance learning even from a moving camera. Hence GPP is very different from the top-view image obtained from a ceiling mounted camera. We perform both object detection and tracking in GPP. Two kinds of GPP images are utilized: gray GPP, which represents the maximal height of 3D points projecting to each pixel, and binary GPP, which is obtained by thresholding the gray GPP. For detection, a simple connected component labeling is used to detect footprints of targets in binary GPP. For tracking, a novel Pixel Level Association (PLA) strategy is proposed to link the same target in consecutive frames in gray GPP. It utilizes optical flow in gray GPP, which to our best knowledge has never been done before. Then we "back project" the detected and tracked objects in GPP to original, side-view (RGB) images. Hence we are able to detect and track objects in the side-view (RGB) images. Our system is able to robustly detect and track multiple moving targets in real time. The detection process does not rely on any target model, which means we do not need any training process. Moreover, tracking does not require any manual initialization, since all entering objects are robustly detected. We also extend the novel framework to robot navigation by tracking. As our experimental results demonstrate, our approach can achieve near prefect detection and tracking results. The performance gain in comparison to state-of-the-art trackers is most significant in the presence of occlusion and background clutter.
Yu Zhou 0016, Yinfei Yang, Meng Yi, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
Int. J. Pattern Recognit. Artif. Intell.6
2014 Bag of contour fragments for robust shape classification
Xinggang Wang, Bin Feng 0001, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
Pattern Recognit.5
2013 Graph Transduction Learning with Connectivity Constraints with Application to Multiple Foreground Cosegmentation
abstract
The proposed approach is based on standard graph transduction, semi-supervised learning (SSL) framework. Its key novelty is the integration of global connectivity constraints into this framework. Although connectivity leads to higher order constraints and their number is an exponential, finding the most violated connectivity constraint can be done efficiently in polynomial time. Moreover, each such constraint can be represented as a linear inequality. Based on this fact, we design a cutting-plane algorithm to solve the integrated problem. It iterates between solving a convex quadratic problem of label propagation with linear inequality constraints, and finding the most violated constraint. We demonstrate the benefits of the proposed approach on a realistic and very challenging problem of co segmentation of multiple foreground objects in photo collections in which the foreground objects are not present in all photos. The obtained results not only demonstrate performance boost induced by the connectivity constraints, but also show a significant improvement over the state-of-the-art methods.
Tianyang Ma, Longin Jan Latecki
CVPR2
2013 Skeleton pruning as trade-off between skeleton simplicity and reconstruction error
Wei Shen 0002, Xiang Bai, Xingwei Yang, Longin Jan Latecki
Sci. China Inf. Sci.4
2013 Face identification using reference-based features with message passing model
Wei Shen 0002, Bo Wang 0044, Xiang Bai, Longin Jan Latecki
Neurocomputing5
2013 Fast Detection of Dense Subgraphs with Iterative Shrinking and Expansion
abstract
In this paper, we propose an efficient algorithm to detect dense subgraphs of a weighted graph. The proposed algorithm, called the shrinking and expansion algorithm (SEA), iterates between two phases, namely, the expansion phase and the shrink phase, until convergence. For a current subgraph, the expansion phase adds the most related vertices based on the average affinity between each vertex and the subgraph. The shrink phase considers all pairwise relations in the current subgraph and filters out vertices whose average affinities to other vertices are smaller than the average affinity of the result subgraph. In both phases, SEA operates on small subgraphs; thus it is very efficient. Significant dense subgraphs are robustly enumerated by running SEA from each vertex of the graph. We evaluate SEA on two different applications: solving correspondence problems and cluster analysis. Both theoretic analysis and experimental results show that SEA is very efficient and robust, especially when there exists a large amount of noise in edge weights.
Hairong Liu, Longin Jan Latecki, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Affinity Learning with Diffusion on Tensor Product Graph
abstract
In many applications, we are given a finite set of data points sampled from a data manifold and represented as a graph with edge weights determined by pairwise similarities of the samples. Often the pairwise similarities (which are also called affinities) are unreliable due to noise or due to intrinsic difficulties in estimating similarity values of the samples. As observed in several recent approaches, more reliable similarities can be obtained if the original similarities are diffused in the context of other data points, where the context of each point is a set of points most similar to it. Compared to the existing methods, our approach differs in two main aspects. First, instead of diffusing the similarity information on the original graph, we propose to utilize the tensor product graph (TPG) obtained by the tensor product of the original graph with itself. Since TPG takes into account higher order information, it is not a surprise that we obtain more reliable similarities. However, it comes at the price of higher order computational complexity and storage requirement. The key contribution of the proposed approach is that the information propagation on TPG can be computed with the same computational complexity and the same amount of storage as the propagation on the original graph. We prove that a graph diffusion process on TPG is equivalent to a novel iterative algorithm on the original graph, which is guaranteed to converge. After its convergence we obtain new edge weights that can be interpreted as new, learned affinities. We stress that the affinities are learned in an unsupervised setting. We illustrate the benefits of the proposed approach for data manifolds composed of shapes, images, and image patches on two very different tasks of image retrieval and image segmentation. With learned affinities, we achieve the bull's eye retrieval score of 99.99 percent on the MPEG-7 shape dataset, which is much higher than the state-of-the-art algorithms. When the data- points are image patches, the NCut with the learned affinities not only significantly outperforms the NCut with the original affinities, but it also outperforms state-of-the-art image segmentation methods.
Xingwei Yang, Lakshman Prasad, Longin Jan Latecki
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Shape clustering: Common structure discovery
Wei Shen 0002, Yan Wang 0033, Xiang Bai, Longin Jan Latecki
Pattern Recognit.5
2012 Navigation toward Non-static Target Object Using Footprint Detection Based Tracking
Meng Yi, Yinfei Yang, Wenjing Qi, Yu Zhou 0016, Zygmunt Pizlo, Longin Jan Latecki
ACCV (3)7
2012 Maximum weight cliques with mutex constraints for video object segmentation
abstract
In this paper, we address the problem of video object segmentation, which is to automatically identify the primary object and segment the object out in every frame. We propose a novel formulation of selecting object region candidates simultaneously in all frames as finding a maximum weight clique in a weighted region graph. The selected regions are expected to have high objectness score (unary potential) as well as share similar appearance (binary potential). Since both unary and binary potentials are unreliable, we introduce two types of mutex (mutual exclusion) constraints on regions in the same clique: intra-frame and inter-frame constraints. Both types of constraints are expressed in a single quadratic form. We propose a novel algorithm to compute the maximal weight cliques that satisfy the constraints. We apply our method to challenging benchmark videos and obtain very competitive results that outperform state-of-the-art methods.
Tianyang Ma, Longin Jan Latecki
CVPR2
2012 Fan Shape Model for object detection
abstract
We propose a novel shape model for object detection called Fan Shape Model (FSM). We model contour sample points as rays of final length emanating for a reference point. As in folding fan, its slats, which we call rays, are very flexible. This flexibility allows FSM to tolerate large shape variance. However, the order and the adjacency relation of the slats stay invariant during fan deformation, since the slats are connected with a thin fabric. In analogy, we enforce the order and adjacency relation of the rays to stay invariant during the deformation. Therefore, FSM preserves discriminative power while allowing for a substantial shape deformation. FSM allows also for precise scale estimation during object detection. Thus, there is not need to scale the shape model or image in order to perform object detection. Another advantage of FSM is the fact that it can be applied directly to edge images, since it does not require any linking of edge pixels to edge fragments (contours).
Xinggang Wang, Xiang Bai, Tianyang Ma, Wenyu Liu 0001, Longin Jan Latecki
CVPR5
2012 Coefficient Thresholding with Image Restoration
abstract
During Coefficient thresholding (CT), the last several nonzero DCT coefficients after quantization are dropped if better Rate-Distortion (RD) performance can be achieved. Since image data can be reconstructed by incomplete frequency information with some prior knowledge of image property (e.g., luminance continuity), the quality degradation caused by CT can be sometimes alleviated by some prior knowledge based image restoration. Consider an 8×8 image block I that is represented by 64 transform coefficients of the prediction residual, C1~64=DCT(I-Ipred). When CT is performed, the last few nonzero coefficients of C1~64are dropped as long as the RD cost can be reduced. Then the reconstructed block Ireccan be calculated by Irec1=IDCT(C1~k)+Ipred, where k denotes the index of last nonzero coefficient of the remaining. With some certain image restoration technique, the lost information during CT can be partially recovered, Irec2=RESTORE(DCT(C1~k) )+Ipred). We employ Bilateral Filter (BF) for the image restoration after CT. This CT/BF approach is implemented as a candidate mode in addition to the traditional IDCT mode, and RD optimization is utilized for the mode selection. The CT/BF mode is enabled for the blocks with texture and edges, which saves considerable computational complexity. Moreover, to save the overhead of the mode flags, the mode information are transmitted covertly in terms of the parity of the number of nonzero coefficients like watermarks. Experiments show that the codec with CT/BF improves the quality of decoded video by up to 0.54 dB compared to H.264 high profile.
Wenfei Jiang, Longin Jan Latecki, Zhibo Chen 0001
DCC3
2012 Clustering Aggregation as Maximum-Weight Independent Set
abstract
We formulate clustering aggregation as a special instance of Maximum-Weight Independent Set (MWIS) problem. For a given dataset, an attributed graph is constructed from the union of the input clusterings generated by different underlying clustering algorithms with different parameters. The vertices, which represent the distinct clusters, are weighted by an internal index measuring both cohesion and separation. The edges connect the vertices whose corresponding clusters overlap. Intuitively, an optimal aggregated clustering can be obtained by selecting an optimal subset of non-overlapping clusters partitioning the dataset together. We formalize this intuition as the MWIS problem on the attributed graph, i.e., finding the heaviest subset of mutually non-adjacent vertices. This MWIS problem exhibits a special structure. Since the clusters of each input clustering form a partition of the dataset, the vertices corresponding to each clustering form a maximal independent set (MIS) in the attributed graph. We propose a variant of simulated annealing method that takes advantage of this special structure. Our algorithm starts from each MIS, which is close to a distinct local optimum of the MWIS problem, and utilizes a local search heuristic to explore its neighborhood in order to find the MWIS. Extensive experiments on many challenging datasets show that: 1. our approach to clustering aggregation automatically decides the optimal number of clusters; 2. it does not require any parameter tuning for the underlying clustering algorithms; 3. it can combine the advantages of different underlying clustering algorithms to achieve superior performance; 4. it is robust against moderate or even bad input clusterings.
Nan Li 0021, Longin Jan Latecki
NIPS2
2012 Fusion with Diffusion for Robust Visual Tracking
abstract
A weighted graph is used as an underlying structure of many algorithms like semi-supervised learning and spectral clustering. The edge weights are usually deter-mined by a single similarity measure, but it often hard if not impossible to capture all relevant aspects of similarity when using a single similarity measure. In par-ticular, in the case of visual object matching it is beneficial to integrate different similarity measures that focus on different visual representations. In this paper, a novel approach to integrate multiple similarity measures is pro-posed. First pairs of similarity measures are combined with a diffusion process on their tensor product graph (TPG). Hence the diffused similarity of each pair of ob-jects becomes a function of joint diffusion of the two original similarities, which in turn depends on the neighborhood structure of the TPG. We call this process Fusion with Diffusion (FD). However, a higher order graph like the TPG usually means significant increase in time complexity. This is not the case in the proposed approach. A key feature of our approach is that the time complexity of the dif-fusion on the TPG is the same as the diffusion process on each of the original graphs, Moreover, it is not necessary to explicitly construct the TPG in our frame-work. Finally all diffused pairs of similarity measures are combined as a weighted sum. We demonstrate the advantages of the proposed approach on the task of visual tracking, where different aspects of the appearance similarity between the target object in frame t and target object candidates in frame t+1 are integrated. The obtained method is tested on several challenge video sequences and the experimental results show that it outperforms state-of-the-art tracking methods.
Yu Zhou 0016, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
NIPS4
2012 Dense Neighborhoods on Affinity Graph
Hairong Liu, Xingwei Yang, Longin Jan Latecki, Shuicheng Yan
Int. J. Comput. Vis.3
2012 Contour-based object detection as dominant set computation
Xingwei Yang, Hairong Liu, Longin Jan Latecki
Pattern Recognit.3
2012 Shape matching and classification using height functions
Xiang Bai, Xinge You, Wenyu Liu 0001, Longin Jan Latecki
Pattern Recognit. Lett.5
2011 Size Adaptive Selection of Most Informative Features
abstract
In this paper, we propose a novel method to select the most informativesubset of features, which has little redundancy andvery strong discriminating power. Our proposed approach automaticallydetermines the optimal number of features and selectsthe best subset accordingly by maximizing the averagepairwise informativeness, thus has obvious advantage overtraditional filter methods. By relaxing the essential combinatorialoptimization problem into the standard quadratic programmingproblem, the most informative feature subset canbe obtained efficiently, and a strategy to dynamically computethe redundancy between feature pairs further greatly acceleratesour method through avoiding unnecessary computationsof mutual information. As shown by the extensive experiments,the proposed method can successfully select the mostinformative subset of features, and the obtained classificationresults significantly outperform the state-of-the-art results onmost test datasets.
Si Liu 0001, Hairong Liu, Longin Jan Latecki, Shuicheng Yan, Changsheng Xu, Hanqing Lu
AAAI3
2011 From partial shape matching through local deformation to robust global shape similarity for object detection
abstract
In this paper, we propose a novel framework for contour based object detection. Compared to previous work, our contribution is three-fold. 1) A novel shape matching scheme suitable for partial matching of edge fragments. The shape descriptor has the same geometric units as shape context but our shape representation is not histogram based. 2) Grouping of partial matching hypotheses to object detection hypotheses is expressed as maximum clique inference on a weighted graph. 3) A novel local affine-transformation to utilize the holistic shape information for scoring and ranking the shape similarity hypotheses. Consequently, each detection result not only identifies the location of the target object in the image, but also provides a precise location of its contours, since we transform a complete model contour to the image. Very competitive results on ETHZ dataset, obtained in a pure shape-based framework, demonstrate that our method achieves not only accurate object detection but also precise contour localization on cluttered background.
Tianyang Ma, Longin Jan Latecki
CVPR2
2011 Feature context for image classification and object detection
abstract
In this paper, we presents a new method to encode the spatial information of local image features, which is a natural extension of Shape Context (SC), so we call it Feature Context (FC). Given a position in a image, SC computes histogram of other points belonging to the target binary shape based on their distances and angles to the position. The value of each histogram bin of SC is the number of the shape points in the region assigned to the bin. Thus, SC requires knowing the location of the points of the target shape. In other words, an image point can have only two labels, it belongs to the shape or not. In contrast, FC can be applied to the whole image without knowing the location of the target shape in the image. Each image point can have multiple labels depending on its local features. The value of each histogram bin of FC is a histogram of various features assigned to points in the bin region. We also introduce an efficient coding method to encode the local image features, call Radial Basis Coding (RBC). Combining RBC and FC together, and using a linear SVM classifier, our method is suitable for both image classification and object detection.
Xinggang Wang, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
CVPR4
2011 Particle filter with state permutations for solving image jigsaw puzzles
abstract
We deal with an image jigsaw puzzle problem, which is defined as reconstructing an image from a set of square and non-overlapping image patches. It is known that a general instance of this problem is NP-complete, and it is also challenging for humans, since in the considered setting the original image is not given. Recently a graphical model has been proposed to solve this and related problems. The target label probability function is then maximized using loopy belief propagation. We also formulate the problem as maximizing a label probability function and use exactly the same pairwise potentials. Our main contribution is a novel inference approach in the sampling framework of Particle Filter (PF). Usually in the PF framework it is assumed that the observations arrive sequentially, e.g., the observations are naturally ordered by their time stamps in the tracking scenario. Based on this assumption, the posterior density over the corresponding hidden states is estimated. In the jigsaw puzzle problem all observations (puzzle pieces) are given at once without any particular order. Therefore, we relax the assumption of having ordered observations and extend the PF framework to estimate the posterior density by exploring different orders of observations and selecting the most informative permutations of observations. This significantly broadens the scope of applications of the PF inference. Our experimental results demonstrate that the proposed inference framework significantly outperforms the loopy belief propagation in solving the image jigsaw puzzle problem. In particular, the extended PF inference triples the accuracy of the label assignment compared to that using loopy belief propagation.
Xingwei Yang, Nagesh Adluru, Longin Jan Latecki
CVPR3
2011 Affinity learning on a tensor product graph with applications to shape and image retrieval
abstract
As observed in several recent publications, improved retrieval performance is achieved when pairwise similarities between the query and the database objects are replaced with more global affinities that also consider the relation among the database objects. This is commonly achieved by propagating the similarity information in a weighted graph representing the database and query objects. Instead of propagating the similarity information on the original graph, we propose to utilize the tensor product graph (TPG) obtained by the tensor product of the original graph with itself. By virtue of this construction, not only local but also long range similarities among graph nodes are explicitly represented as higher order relations, making it possible to better reveal the intrinsic structure of the data manifold. In addition, we improve the local neighborhood structure of the original graph in a preprocessing stage. We illustrate the benefits of the proposed approach on shape and image ranking and retrieval tasks. We are able to achieve the bull's eye retrieval score of 99.99% on MPEG-7 shape dataset, which is much higher than the state-of-the-art algorithms.
Xingwei Yang, Longin Jan Latecki
CVPR2
2011 Maximal Cliques that Satisfy Hard Constraints with Application to Deformable Object Model Learning
abstract
We propose a novel inference framework for finding maximal cliques in a weighted graph that satisfy hard constraints. The constraints specify the graph nodes that must belong to the solution as well as mutual exclusions of graph nodes, i.e., sets of nodes that cannot belong to the same solution. The proposed inference is based on a novel particle filter algorithm with state permeations. We apply the inference framework to a challenging problem of learning part-based, deformable object models. Two core problems in the learning framework, matching of image patches and finding salient parts, are formulated as two instances of the problem of finding maximal cliques with hard constraints. Our learning framework yields discriminative part based object models that achieve very good detection rate, and outperform other methods on object classes with large deformation.
Xinggang Wang, Xiang Bai, Xingwei Yang, Wenyu Liu 0001, Longin Jan Latecki
NIPS5
2011 Improving SVM classification on imbalanced time series data sets with ghost points
Suzan Köknar-Tezel, Longin Jan Latecki
Knowl. Inf. Syst.2
2011 Skeleton growing and pruning with bending potential ratio
Wei Shen 0002, Xiang Bai, Longin Jan Latecki
Pattern Recognit.5
2010 Convex shape decomposition
abstract
In this paper, we propose a new shape decomposition method, called convex shape decomposition. We formalize the convex decomposition problem as an integer linear programming problem, and obtain approximate optimal solution by minimizing the total cost of decomposition under some concavity constraints. Our method is based on Morse theory and combines information from multiple Morse functions. The obtained decomposition provides a compact representation, both geometrical and topological, of original object. Our experiments show that such representation is very useful in many applications.
Hairong Liu, Wenyu Liu 0001, Longin Jan Latecki
CVPR3
2010 Two-Step Coding for High Definition Video Compression
abstract
High definition (HD) video has come into people’s life from movie theaters to HDTV. However, the compression of HD videos is a challenging problem due to flicker noise, caused by film grain. The flicker noise significantly limits the applicability of motion estimation (ME), which is a key factor of the efficient video compression in block-based coding standards. Due to the flicker noise, it is difficult to obtain a perfect match between a current block and a reference block. In block-based video coding standards including H.264 a given block is either encoded by inter-frame or intra-frame prediction. We propose a new coding scheme called Two-Step Coding (TSC) that utilizes both for each block. TSC first reduces the resolution of each frame by replacing each block with its DC coefficient of the DCT to the original color values. The flicker noise is greatly reduced in the obtained lower resolution frame, which we call DC frame. The key benefit is that ME becomes very efficient on DC frames, and consequently, the DC frame can be efficiently inter-frame coded. The difference between the original frame and DC frame is actually described by the AC coefficients of the DCT of the original frame. We utilize the existing H.264 tools to combine the intra-frame and inter-frame coded parts of blocks both on the encoder and decoder sides. The key benefit of the proposed TSC in comparison to the most popular standards, in particular, in comparison to H.264 lies in better utilization of inter-frame coding.. Due to flicker noise, H.264 mostly employs intra block coding on HD videos. However, it is well-known that inter-frame coding significantly outperforms intra coding in video compression rate if the temporal correllation is correctly utilized. By reducing each frame to DC frame, TSC makes it possible to apply inter-frame coding. We provide experimental data and analysis to illustrate this fact.
Wenfei Jiang, Wenyu Liu 0001, Longin Jan Latecki, Bing Feng
DCC3
2010 Balancing Deformability and Discriminability for Shape Matching
Haibin Ling, Xingwei Yang, Longin Jan Latecki
ECCV (3)3
2010 Boosting Chamfer Matching by Learning Chamfer Distance Normalization
Tianyang Ma, Xingwei Yang, Longin Jan Latecki
ECCV (5)3
2010 Weakly Supervised Shape Based Object Detection with Particle Filter
Xingwei Yang, Longin Jan Latecki
ECCV (5)2
2010 Robust Clustering as Ensembles of Affinity Relations
abstract
In this paper, we regard clustering as ensembles of k-ary affinity relations and clusters correspond to subsets of objects with maximal average affinity relations. The average affinity relation of a cluster is relaxed and well approximated by a constrained homogenous function. We present an efficient procedure to solve this optimization problem, and show that the underlying clusters can be robustly revealed by using priors systematically constructed from the data. Our method can automatically select some points to form clusters, leaving other points un-grouped; thus it is inherently robust to large numbers of outliers, which has seriously limited the applicability of classical methods. Our method also provides a unified solution to clustering from k-ary affinity relations with k ≥ 2, that is, it applies to both graph-based and hypergraph-based clustering problems. Both theoretical analysis and experimental results show the superiority of our method over classical solutions to the clustering problem, especially when there exists a large number of outliers.
Hairong Liu, Longin Jan Latecki, Shuicheng Yan
NIPS2
2010 Contour based object detection using part bundles
ChengEn Lu, Nagesh Adluru, Haibin Ling, Guangxi Zhu, Longin Jan Latecki
Comput. Vis. Image Underst.5
2010 Skeletonization with Particle Filters
abstract
We present a novel method to obtain high quality skeletons of binary shapes. The obtained skeletons are connected and one pixel thick. They do not require any pruning or any other post-processing. The computation is composed of two major parts. First, a small set of salient contour points is computed. We use Discrete Curve Evolution, but any other robust method could be used. Second, particle filters are used to obtain the skeleton. The main idea is that the particles walk along the skeletal paths between pairs of the salient points. We provide experimental results that clearly demonstrate that the proposed method significantly outperforms other well-known methods for skeleton computation. Moreover, we propose an extension of our method to computing skeletons of gray level images and provide promising experimental results.
Yuchun Tang, Xiang Bai, Xingwei Yang, Shuwei Liu, Longin Jan Latecki
Int. J. Pattern Recognit. Artif. Intell.6
2010 Learning Context-Sensitive Shape Similarity by Graph Transduction
abstract
Shape similarity and shape retrieval are very important topics in computer vision. The recent progress in this domain has been mostly driven by designing smart shape descriptors for providing better similarity measure between pairs of shapes. In this paper, we provide a new perspective to this problem by considering the existing shapes as a group, and study their similarity measures to the query shape in a graph structure. Our method is general and can be built on top of any existing shape similarity measure. For a given similarity measure, a new similarity is learned through graph transduction. The new similarity is learned iteratively so that the neighbors of a given shape influence its final similarity to the query. The basic idea here is related to PageRank ranking, which forms a foundation of Google Web search. The presented experimental results demonstrate that the proposed approach yields significant improvements over the state-of-art shape matching algorithms. We obtained a retrieval rate of 91.61 percent on the MPEG-7 data set, which is the highest ever reported in the literature. Moreover, the learned similarity by the proposed method also achieves promising improvements on both shape classification and shape clustering.
Xiang Bai, Xingwei Yang, Longin Jan Latecki, Wenyu Liu 0001, Zhuowen Tu
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Shape band: A deformable object detection approach
abstract
In this paper, we focus on the problem of detecting/matching a query object in a given image. We propose a new algorithm, shape band, which models an object within a bandwidth of its sketch/contour. The features associated with each point on the sketch are the gradients within the bandwidth. In the detection stage, the algorithm simply scans an input image at various locations and scales for good candidates. We then perform fine scale shape matching to locate the precise object boundaries, also by taking advantage of the information from the shape band. The overall algorithm is very easy to implement, and our experimental results show that it can outperform stat-of-the-art contour based object detection algorithms.
Xiang Bai, Quannan Li, Longin Jan Latecki, Wenyu Liu 0001, Zhuowen Tu
CVPR3
2009 Locally constrained diffusion process on locally densified distance spaces with applications to shape retrieval
abstract
The matching and retrieval of 2D shapes is an important challenge in computer vision. A large number of shape similarity approaches have been developed, with the main focus being the comparison or matching of pairs of shapes. In these approaches, other shapes do not influence the similarity measure of a given pair of shapes. In the proposed approach, other shapes do influence the similarity measure of each pair of shapes, and we show that this influence is beneficial even in the unsupervised setting (without any prior knowledge of shape classes). The influence of other shapes is propagated as a diffusion process on a graph formed by a given set of shapes. However, the classical diffusion process does not perform well in shape space for two reasons: it is unstable in the presence of noise and the underlying local geometry is sparse. We introduce a locally constrained diffusion process which is more stable even if noise is present, and we densify the shape space by adding synthetic points we call 'ghost points'. We present experimental results that demonstrate very significant improvements over state-of-the-art shape matching algorithms. On the MPEG-7 data set, we obtained a bull's-eye retrieval score of 93.32%, which is the highest score ever reported in the literature.
Xingwei Yang, Suzan Köknar-Tezel, Longin Jan Latecki
CVPR3
2009 Active skeleton for non-rigid object detection
abstract
We present a shape-based algorithm for detecting and recognizing non-rigid objects from natural images. The existing literature in this domain often cannot model the objects very well. In this paper, we use the skeleton (medial axis) information to capture the main structure of an object, which has the particular advantage in modeling articulation and non-rigid deformation. Given a set of training samples, a tree-union structure is learned on the extracted skeletons to model the variation in configuration. Each branch on the skeleton is associated with a few part-based templates, modeling the object boundary information. We then apply sum-and-max algorithm to perform rapid object detection by matching the skeleton-based active template to the edge map extracted from a test image. The algorithm reports the detection result by a composition of the local maximum responses. Compared with the alternatives on this topic, our algorithm requires less training samples. It is simple, yet efficient and effective. We show encouraging results on two widely used benchmark image sets: the Weizmann horse dataset [7] and the ETHZ dataset [16].
Xiang Bai, Xinggang Wang, Longin Jan Latecki, Wenyu Liu 0001, Zhuowen Tu
ICCV3
2009 Shape guided contour grouping with particle filters
abstract
We propose a novel framework for contour based object detection and recognition, which we formulate as a joint contour fragment grouping and labeling problem. For a given set of contours of model shapes, we simultaneously perform selection of relevant contour fragments in edge images, grouping of the selected contour fragments, and their matching to the model contours. The inference in all these steps is performed using particle filters (PF) but with static observations. Our approach needs one example shape per class as training data. The PF framework combined with decomposition of model contour fragments to part bundles allows us to implement an intuitive search strategy for the target contour in a clutter of edge fragments. First a rough sketch of the model shape is identified, followed by fine tuning of shape details. We show that this framework yields not only accurate object detections but also localizations in real cluttered images.
ChengEn Lu, Longin Jan Latecki, Nagesh Adluru, Xingwei Yang, Haibin Ling
ICCV2
2009 Improving SVM Classification on Imbalanced Data Sets in Distance Spaces
abstract
Imbalanced data sets present a particular challenge to the data mining community. Often, it is the rare event that is of interest and the cost of misclassifying the rare event is higher than misclassifying the usual event. When the data is highly skewed toward the usual, it can be very difficult for a learning system to accurately detect the rare event. There have been many approaches in recent years for handling imbalanced data sets, from under-sampling the majority class to adding synthetic points to the minority class in feature space. Distances between time series are known to be non-Euclidean and nonmetric, since comparing time series requires warping in time. This fact makes it impossible to apply standard methods like SMOTE to insert synthetic data points in feature spaces. We present an innovative approach that augments the minority class by adding synthetic points in distance spaces. We then use Support Vector Machines for classification. Our experimental results on standard time series show that our synthetic points significantly improve the classification rate of the rare events, and in many cases also improves the overall accuracy of SVM.
Suzan Köknar-Tezel, Longin Jan Latecki
ICDM2
2009 Outlier Detection with Globally Optimal Exemplar-Based GMM
abstract
Outlier detection has recently become an important problem in many data mining applications. In this paper, a novel unsupervised algorithm for outlier detection is proposed. First we apply a provably globally optimal Expectation Maximization (EM) algorithm to fit a Gaussian Mixture Model (GMM) to a given data set. In our approach, a Gaussian is centered at each data point, and hence, the estimated mixture proportions can be interpreted as probabilities of being a cluster center for all data points. The outlier factor at each data point is then defined as a weighted sum of the mixture proportions with weights representing the similarities to other data points. The proposed outlier factor is thus based on global properties of the data set. This is in contrast to most existing approaches to outlier detection, which are strictly local. Our experiments performed on several simulated and real life data sets demonstrate superior performance of the proposed approach. Moreover, we also demonstrate the ability to detect unusual shapes.
Xingwei Yang, Longin Jan Latecki, Dragoljub Pokrajac
SDM2
2009 Contour Grouping Based on Contour-Skeleton Duality
Nagesh Adluru, Longin Jan Latecki
Int. J. Comput. Vis.2
2009 Piecewise Linear Models with Guaranteed Closeness to the Data
abstract
This paper addresses the problem of piecewise linear approximation of point sets without any constraints on the order of data points or the number of model components (line segments). We point out two problems with the maximum likelihood estimate (MLE) that present serious drawbacks in practical applications. One is that the parametric models obtained using a classical MLE framework are not guaranteed to be close to data points. It is typically impossible, in this classical framework, to detect whether a parametric model fits the data well or not. The second problem is related to accurately choosing the optimal number of model components. We first fit a nonparametric density to the data points and use it to define a neighborhood of the data. Observations inside this neighborhood are deemed informative; those outside the neighborhood are deemed uninformative for our purpose. This provides us with a means to recognize when models fail to properly fit the data. We then obtain maximum likelihood estimates by optimizing the Kullback-Leibler Divergence (KLD) between the nonparametric data density restricted to this neighborhood and a mixture of parametric models. We prove that, under the assumption of a reasonably large sample size, the inferred model components are close to their ground-truth model component counterparts. This holds independently of the initial number of assumed model components or their associated parameters. Moreover, in the proposed approach, we are able to estimate the number of significant model components without any additional computation.
Longin Jan Latecki, Marc Sobel, Rolf Lakämper
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 A Video Coding Scheme Based on Joint Spatiotemporal and Adaptive Prediction
abstract
We propose a video coding scheme that departs from traditional Motion Estimation/DCT frameworks and instead uses Karhunen-Loeve Transform (KLT)/Joint Spatiotemporal Prediction framework. In particular, a novel approach that performs joint spatial and temporal prediction simultaneously is introduced. It bypasses the complex H.26x interframe techniques and it is less computationally intensive. Because of the advantage of the effective joint prediction and the image-dependent color space transformation (KLT), the proposed approach is demonstrated experimentally to consistently lead to improved video quality, and in many cases to better compression rates and improved computational speed.
Wenfei Jiang, Longin Jan Latecki, Wenyu Liu 0001, Ken Gorman
IEEE Trans. Image Process.2
2008 Improving Shape Retrieval by Learning Graph Transduction
Xingwei Yang, Xiang Bai, Longin Jan Latecki, Zhuowen Tu
ECCV (4)3
2008 Merging maps of multiple robots
abstract
Merging local maps, acquired by multiple robots, into a global map, (also known as map merging) is one of the important issues faced by virtually all cooperative exploration techniques. We present a novel and simple solution to the problem of map merging by reducing it to the problem of SLAM of a single ¿virtual¿ robot. The individual local maps and their shape information constitute the sensor information for the virtual robot. This approach allows us to adapt the framework of Rao-Blackwellized particle filtering used in SLAM of a single robot for the problem of map merging.
Nagesh Adluru, Longin Jan Latecki, Marc Sobel, Rolf Lakämper
ICPR2
2008 Multiscale Random Fields with Application to Contour Grouping
abstract
We introduce a new interpretation of multiscale random fields (MSRFs) that admits efficient optimization in the framework of regular (single level) random fields (RFs). It is based on a new operator, called append, that combines sets of random variables (RVs) to single RVs. We assume that a MSRF can be decomposed into disjoint trees that link RVs at different pyramid levels. The append operator is then applied to map RVs in each tree structure to a single RV. We demonstrate the usefulness of the proposed approach on a challenging task involving grouping contours of target shapes in images. MSRFs provide a natural representation of multiscale contour models, which are needed in order to cope with unstable contour decompositions. The append operator allows us to find optimal image labels using the classical framework of relaxation labeling, Alternative methods like Markov Chain Monte Carlo (MCMC) could also be used.
Longin Jan Latecki, ChengEn Lu, Marc Sobel, Xiang Bai
NIPS1
2008 Computing Stable Skeletons with Particle Filters
Xiang Bai, Xingwei Yang, Longin Jan Latecki, Yanbo Xu, Wenyu Liu 0001
PRICAI3
2008 A Unified Curvature Definition for Regular, Polygonal, and Digital Planar Curves
Hairong Liu, Longin Jan Latecki, Wenyu Liu 0001
Int. J. Comput. Vis.2
2008 Skeleton-Based Shape Classification Using Path Similarity
abstract
Most of the traditional methods for shape classification are based on contour. They often encounter difficulties when dealing with classes that have large nonlinear variability, especially when the variability is structural or due to articulation. It is well-known that shape representation based on skeletons is superior to contour based representation in such situations. However, approaches to shape similarity based on skeletons suffer from the instability of skeletons, and matching of skeleton graphs is still an open problem. Using a new skeleton pruning method, we are able to obtain stable pruned skeletons even in the presence of significant contour distortions. We also propose a new method for matching of skeleton graphs. In contrast to most existing methods, it does not require converting of skeleton graphs to trees and it does not require any graph editing. Shape classification is done with Bayesian classifier. We present excellent classification results for complete shapes.
Xiang Bai, Xingwei Yang, Deguang Yu, Longin Jan Latecki
Int. J. Pattern Recognit. Artif. Intell.4
2008 Path Similarity Skeleton Graph Matching
abstract
This paper presents a novel framework to for shape recognition based on object silhouettes. The main idea is to match skeleton graphs by comparing the shortest paths between skeleton endpoints. In contrast to typical tree or graph matching methods, we completely ignore the topological graph structure. Our approach is motivated by the fact that visually similar skeleton graphs may have completely different topological structures. The proposed comparison of shortest paths between endpoints of skeleton graphs yields correct matching results in such cases. The skeletons are pruned by contour partitioning with Discrete Curve Evolution, which implies that the endpoints of skeleton branches correspond to visual parts of the objects. The experimental results demonstrate that our method is able to produce correct results in the presence of articulations, stretching, and occlusion.
Xiang Bai, Longin Jan Latecki
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Detection and recognition of contour parts based on shape similarity
Xiang Bai, Xingwei Yang, Longin Jan Latecki
Pattern Recognit.3
2007 Incremental Local Outlier Detection for Data Streams
abstract
Outlier detection has recently become an important problem in many industrial and financial applications. This problem is further complicated by the fact that in many cases, outliers have to be detected from data streams that arrive at an enormous pace. In this paper, an incremental LOF (local outlier factor) algorithm, appropriate for detecting outliers in data streams, is proposed. The proposed incremental LOF algorithm provides equivalent detection performance as the iterated static LOF algorithm (applied after insertion of each data record), while requiring significantly less computational time. In addition, the incremental LOF algorithm also dynamically updates the profiles of data points. This is a very important property, since data profiles may change over time. The paper provides theoretical evidence that insertion of a new data point as well as deletion of an old data point influence only limited number of their closest neighbors and thus the number of updates per such insertion/deletion does not depend on the total number of points TV in the data set. Our experiments performed on several simulated and real life data sets have demonstrated that the proposed incremental LOF algorithm is computationally efficient, while at the same time very successful in detecting outliers and changes of distributional behavior in various data stream applications
Dragoljub Pokrajac, Aleksandar Lazarevic, Longin Jan Latecki
CIDM3
2007 Visual Curvature
abstract
In this paper, we propose a new definition of curvature, called visual curvature. It is based on statistics of the extreme points of the height functions computed over all directions. By gradually ignoring relatively small heights, a single parameter multi-scale curvature is obtained. It does not modify the original contour and the scale parameter has an obvious geometric meaning. The theoretical properties and the experiments presented demonstrate that multi-scale visual curvature is stable, even in the presence of significant noise. In particular, it can deal with contours with significant gaps. We also show a relation between multi-scale visual curvature and convexity of simple closed curves. To our best knowledge, the proposed definition of visual curvature is the first ever that applies to regular curves as defined in differential geometry as well as to turn angles of polygonal curves. Moreover, it yields stable curvature estimates of curves in digital images even under sever distortions.
Hairong Liu, Longin Jan Latecki, Wenyu Liu 0001, Xiang Bai
CVPR2
2007 Contour Grouping Based on Local Symmetry
abstract
The paper deals with grouping of edges to contours of shapes using only local symmetry and continuity. Shape skeletons are used to generate the search space for a version of the Markov Chain Monte Carlo approach utilizing particle filters to find the most likely skeleton. Intuitively this means that grouping of edge segments is performed by walking along the skeleton. The particle search, which is an adapted version of a successful algorithm in robot mapping, is assisted by a reference model of a shape, which is expressed as the sequence of sample points and radii of maximal skeleton disks. This model is sufficiently flexible to represent non-rigid deformations, but restrictive enough to perform well on real, noisy image data. The order of skeleton points (and their corresponding segments) found by the particles defines the grouping.
Nagesh Adluru, Longin Jan Latecki, Rolf Lakämper, Thomas Young, Xiang Bai, Ari D. Gross
ICCV2
2007 Optimal Subsequence Bijection
abstract
We consider the problem of elastic matching of sequences of real numbers. Since both a query and a target sequence may be noisy, i.e., contain some outlier elements, it is desirable to exclude the outlier elements from matching in order to obtain a robust matching performance. Moreover, in many applications like shape alignment or stereo correspondence it is also desirable to have a one-to-one and onto correspondence (bijection) between the remaining elements. We propose an algorithm that determines the optimal subsequence bijection (OSB) of a query and target sequence. The OSB is efficiently computed since we map the problem's solution to a cheapest path in a DAG (directed acyclic graph). We obtained excellent results on standard benchmark time series datasets. We compared OSB to Dynamic Time Warping (DTW) with and without warping window. We do not claim that OSB is always superior to DTW. However, our results demonstrate that skipping outlier elements as done by OSB can significantly improve matching results for many real datasets. Moreover, OSB is particularly suitable for partial matching. We applied it to the object recognition problem when only parts of contours are given. We obtained sequences representing shapes by representing object contours as sequences of curvatures.
Longin Jan Latecki, Qiang Wang 0010, Suzan Köknar-Tezel, Vasileios Megalooikonomou
ICDM1
2007 Skeletonization using SSM of the Distance Transform
abstract
This paper proposes a new approach for skeletonization based on the skeleton strength map (SSM) caculated by Euclidean distance transform of a binary image. After the distance transform and gradient are computed, isotropic diffusion is performed on the gradient vector field and the skeleton strength map is computed from the diffused vector field. A critical point set is then selected from local maxima of the SSM. The critical points are located on significant visual parts of the object. The skeleton is obtained by connecting the critical points with geodesic paths. This approach overcomes intrinsic drawbacks of distance transform based skeletons, since it yields stable and connected skeletons without losing significant visual parts.
Longin Jan Latecki, Quannan Li, Xiang Bai, Wenyu Liu 0001
ICIP (5)1
2007 Skeleton Pruning by Contour Partitioning with Discrete Curve Evolution
abstract
In this paper, we introduce a new skeleton pruning method based on contour partitioning. Any contour partition can be used, but the partitions obtained by Discrete Curve Evolution (DCE) yield excellent results. The theoretical properties and the experiments presented demonstrate that obtained skeletons are in accord with human visual perception and stable, even in the presence of significant noise and shape variations, and have the same topology as the original skeletons. In particular, we have proven that the proposed approach never produces spurious branches, which are common when using the known skeleton pruning methods. Moreover, the proposed pruning method does not displace the skeleton points. Consequently, all skeleton points are centers of maximal disks. Again, many existing methods displace skeleton points in order to produces pruned skeletons.
Xiang Bai, Longin Jan Latecki, Wenyu Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Topological Equivalence between a 3D Object and the Reconstruction of Its Digital Image
abstract
Digitization is not as easy as it looks. If one digitizes a 3D object even with a dense sampling grid, the reconstructed digital object may have topological distortions and, in general, there exists no upper bound for the Hausdorff distance. This explains why so far no algorithm has been known which guarantees topology preservation. However, as we will show, it is possible to repair the obtained digital image in a locally bounded way so that it is homeomorphic and close to the 3D object. The resulting digital object is always well-composed, which has nice implications for a lot of image analysis problems. Moreover, we will show that the surface of the original object is homeomorphic to the result of the marching cubes algorithm. This is really surprising since it means that the well-known topological problems of the marching cubes reconstruction simply do not occur for digital images of r-regular objects. Based on the trilinear interpolation, we also construct a smooth isosurface from the digital image that has the same topology as the original surface. Finally, we give a surprisingly simple topology preserving reconstruction method by using overlapping balls instead of cubical voxels. This is the first approach of digitizing 3D objects which guarantees topology preservation and gives an upper bound for the geometric distortion. Since the output can be chosen as a pure voxel presentation, a union of balls, a reconstruction by trilinear interpolation, a smooth isosurface, or the piecewise linear marching cubes surface, the results are directly applicable to a huge class of image analysis algorithms. Moreover, we show how one can efficiently estimate the volume and the surface area of 3D objects by looking at their digitizations. Measuring volume and surface area of digital objects are important problems in 3D image analysis. Good estimators should be multigrid convergent, i.e., the error goes to zero with increasing sampling density. We will show that every presented reconstruction method can be used for volume estimation and we will give a solution for the much more difficult problem of multigrid-convergent surface area estimation. Our solution is based on simple counting of voxels and we are the first to be able to give absolute bounds for the surface area.
Peer Stelldinger, Longin Jan Latecki, Marcelo Siqueira
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 An elastic partial shape matching technique
Longin Jan Latecki, Vasileios Megalooikonomou, Qiang Wang 0010, Deguang Yu
Pattern Recognit.1
2006 Extended EM for Planar Approximation of 3D Data
abstract
Abstract – The paper deals with fitting of planar patches to 3D laser range data obtained by a mobile robot. The number and the initial position of the patches are unknown, hence their estimation is a challenging problem. It is solved by adding iterated steps of split and merge to a modified Expectation Maximization (EM) algorithm. This allows for precise adjustment of the number of patches, independent from the initial model. The proposed approach overcomes the problem of classical EM, which produces an optimal solution only if the number and position of model components is well estimated. Index Terms- 3D Robot Mapping, EM I.
Rolf Lakämper, Longin Jan Latecki
ICRA2
2006 Polygonal Approximation of Laser Range Data based on Perceptual Grouping and EM
abstract
Our goal is polygonal approximation of laser range data points obtained by a mobile robot. The proposed approach provides a precise estimation of the number of model components (line segments) and their initial parameters independent of their initial values. We use principles of perceptual grouping to evaluate the approximation quality obtained in each expectation maximization (EM) step. By evaluating EM approximation quality we are able to recognize a locally optimal solution, and modify the number of model components and their parameters. Consequently, EM can converge only to a globally optimal solution independent of the initial number of model components and their initial parameters
Longin Jan Latecki, Rolf Lakämper
ICRA1
2006 Polygonal Approximation of Point Sets
Longin Jan Latecki, Rolf Lakämper, Marc Sobel
IWCIA1
2006 New EM derived from Kullback-Leibler divergence
abstract
We introduce a new EM framework in which it is possible not only to optimize the model parameters but also the number of model components. A key feature of our approach is that we use nonparametric density estimation to improve parametric density estimation in the EM framework. While the classical EM algorithm estimates model parameters empirically using the data points themselves, we estimate them using nonparametric density estimates.There exist many possible applications that require optimal adjustment of model components. We present experimental results in two domains. One is polygonal approximation of laser range data, which is an active research topic in robot navigation. The other is grouping of edge pixels to contour boundaries, which still belongs to unsolved problems in computer vision.
Longin Jan Latecki, Marc Sobel, Rolf Lakämper
KDD1
2005 Tracking motion objects in infrared videos
abstract
We propose motion detection and object tracking method that is particularly suitable for infrared videos. Detection of moving objects in infrared videos is based on changing texture in parts of the view field. We estimate the speed of texture change by measuring the spread of texture vectors in the texture space. This method allows us to robustly detect very fast and very slow moving object. Our theoretical and experimental results show that the proposed method significantly outperforms the Stauffer-Grimson approach based on Gaussian mixture model. We observe that the proposed method does not require any post-processing, which is a necessary step for the Stauffer-Grimson approach. Moreover, the object tracking is improved when based on the spatiotemporal texture blocks.
Longin Jan Latecki, Roland Miezianko, Dragoljub Pokrajac
AVSS1
2005 Partial Elastic Matching of Time Series
abstract
We consider the problem of elastic matching of time series. We propose an algorithm that determines a subsequence of a target time series that best matches a query series. In the proposed algorithm, we map the problem of the best matching subsequence to the problem of a cheapest path in a DAG (directed acyclic graph). The proposed approach allows us to also compute the optimal scale and translation of time series values, which is a nontrivial problem in the case of subsequence matching.
Longin Jan Latecki, Vasileios Megalooikonomou, Qiang Wang 0010, Rolf Lakämper, Chotirat (Ann) Ratanamahatana, Eamonn J. Keogh
ICDM1
2005 Incremental multi-robot mapping
abstract
The purpose of this paper is to present a technique to create a global map of robots' surroundings by converting the raw data acquired from a scanning sensor to a compact map composed of just a few generalized polylines (polygonal curves). We propose a new approach to merging robots' maps that is composed of a local geometric process of merging similar line segments (termed discrete segment evolution) with a global statistical control process. In the case of single robot, we are able to incrementally build a map showing the environment the robot has traveled through by merging its polygonal map with actual scans. In the case of a robot team, we are able to identify common parts of their partial maps and if common parts are present construct a joint map of the explored environment.
Rolf Lakämper, Longin Jan Latecki, Diedrich Wolter
IROS2
2005 Elastic Partial Matching of Time Series
Longin Jan Latecki, Vasileios Megalooikonomou, Qiang Wang 0010, Rolf Lakämper, Chotirat (Ann) Ratanamahatana, Eamonn J. Keogh
PKDD1
2005 Optimal partial shape similarity
Longin Jan Latecki, Rolf Lakämper, Diedrich Wolter
Image Vis. Comput.1
2004 A two-stream approach for adaptive rate control in multimedia applications
abstract
We propose a two-stream approach for adaptive rate control in multimedia applications. By monitoring a low-rate monitoring stream, we keep track of the available bandwidth of the network path and dynamically adjust the sending rate of the traffic stream close to the optimal rate. The proposed two-stream approach perfectly meets the requirements of the current best-effort Internet and fits well in multimedia applications. For example, there is no bandwidth overhead for the monitoring stream in peer-to-peer video conferencing, because the monitoring stream is the audio stream. We show in our experiments that both the network and the application can benefit from this approach. The proposed two-stream approach is applicable to monitor the sending rate of the traffic stream over UDP as well as over TCP.
Longin Jan Latecki, Jaiwant Mulik
ICME1
2004 Shape Matching for Robot Mapping
Diedrich Wolter, Longin Jan Latecki
PRICAI2
2003 Detection of Changes in Surveillance Videos
abstract
We provide theoretical and experimental results showing that the dimension of video trajectories is a useful tool to access the mid-level content of videos, such as the appearance or disappearance of an object, and changes in the velocity and direction of moving objects. Moreover, the amount of change is proportional to the size of the objects involved and their speed. All this is achieved by a robust technique of dimensionality computation of video trajectories based on eigenvalues.
Longin Jan Latecki, Xiangdong Wen, Nilesh Ghubade
AVSS1
2003 Tree-structured Partitioning Based on Splitting Histograms of Distances
abstract
We propose a novel clustering algorithm that is similar in spirit to classification trees. The data is recursively split using a criterion that applies a discrete curve evolution method to the histogram of distances. The algorithm can be depicted through tree diagrams with triple splits. Leaf nodes represent either clusters or sets of observations that can not yet be clearly assigned to a cluster. After constructing the tree, unclassified data points are mapped to their closest clusters. The algorithm has several advantages. First, it deals effectively with observations that can not be unambiguously assigned to a cluster by allowing a "margin of error". Second, it automatically determines the number of clusters; apart from the margin of error the user only needs to specify the minimal cluster size but not the number of clusters. Third, it is linear with respect to the number of data points and thus suitable for very large data sets. Experiments involving both simulated and real data from different domains show that the proposed method is effective and efficient.
Longin Jan Latecki, Rajagopal Venugopal, Marc Sobel, Steve Horvat
ICDM1
2003 Better audio performance when video stream is monitored by TCP congestion control
abstract
Conventional wisdom holds that the TCP like congestion control is unsuitable for real-time multimedia conferencing. However, our results clearly show that an audio and video conferencing system that transmits video over TCP (and audio over RTP/UDP) can provide significantly better audio quality to the end user than one built on RTP/UDP alone. We measured audio quality in terms of packet loss, packets arriving too late (for real time play out), average packet delay, and jitter. Our results also clearly indicate that sending video over TCP does not introduce any additional delay in the arrival time of video packets in comparison to RTP/UDP.
Longin Jan Latecki, Kishore Kulkarni, Jaiwant Mulik
ICME1
2003 Topologies for the digital spaces Z2 and Z3
Ulrich Eckhardt, Longin Jan Latecki
Comput. Vis. Image Underst.2
2003 The choice of vantage objects for image retrieval
Christian Hennig, Longin Jan Latecki
Pattern Recognit.2
2002 Recovering a Polygon from Noisy Data
Longin Jan Latecki, Azriel Rosenfeld
Comput. Vis. Image Underst.1
2002 Application of planar shape comparison to object retrieval in image databases
Longin Jan Latecki, Rolf Lakämper
Pattern Recognit.1
2002 Special issue: Shape Representation and Similarity for Image Databases
Longin Jan Latecki, Robert Melter, Ari D. Gross
Pattern Recognit.1
2001 Extraction of key frames from videos by optimal color composition matching and polygon simplification
abstract
A video sequence is first mapped to a sequence of points in a semi-metric space that forms a polyline. We require only that a semi-distance between pairs of points be defined that need not satisfy the triangle inequality. By simplifying the polyline, we obtain a small set of the most relevant key frames that is representative of the whole video sequence. The degree of the simplification is either determined automatically or selected by the user. Using our technique, a viewer can browse a video at the level of summarization that suits his patience level. Applications include the creation of a smart fast-forward function for digital VCRs, and the automatic creation of short summaries or trailers that can be used as previews before videos are downloaded from the Web.
Longin Jan Latecki, Daniel de Wildt, Jianying Hu
MMSP1
2000 Shape Descriptors for Non-Rigid Shapes with a Single Closed Contour
abstract
The Core Experiment CE-Shape-1 for shape descriptors performed for the MPEG-7 standard gave a unique opportunity to compare various shape descriptors for non-rigid shapes with a single closed contour. There are two main differences with respect to other comparison results reported in the literature: (1) For each shape descriptor the experiments were carried out by an institute that is in favor of this descriptor. This implies that the parameters for each system were optimally determined and the implementations were thoroughly rested. (2) It was possible to compare the performance of shape descriptors based on totally different mathematical approaches. A more theoretical comparison of these descriptors seems to be extremely hard. In this paper we report on the MPEG-7 Core Experiment CE-Shape.
Longin Jan Latecki, Rolf Lakämper, Ulrich Eckhardt
CVPR1
2000 Shape Similarity Measure Based on Correspondence of Visual Parts
abstract
A cognitively motivated similarity measure is presented and its properties are analyzed with respect to retrieval of similar objects in image databases of silhouettes of 2D objects. To reduce influence of digitization noise, as well as segmentation errors, the shapes are simplified by a novel process of digital curve evolution. To compute our similarity measure, we first establish the best possible correspondence of visual parts (without explicitly computing the visual parts). Then, the similarity between corresponding parts is computed and aggregated. We applied our similarity measure to shape matching of object contours in various image databases and compared it to well-known approaches in the literature. The experimental results justify that our shape matching procedure gives an intuitive shape correspondence and is stable with respect to noise distortions.
Longin Jan Latecki, Rolf Lakämper
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Convexity Rule for Shape Decomposition Based on Discrete Contour Evolution
Longin Jan Latecki, Rolf Lakämper
Comput. Vis. Image Underst.1
1999 Digitizations preserving shape
Antonio Giraldo, Ari D. Gross, Longin Jan Latecki
Pattern Recognit.3
1999 Digital geometric methods in document image analysis
Ari D. Gross, Longin Jan Latecki
Pattern Recognit.2
1998 Digital Geometric Methods in Image Analysis and Compression
Ari D. Gross, Longin Jan Latecki
ACCV (1)2
1998 Shape similarity measure for image database of occluding contours
abstract
A similarity measure for silhouettes of 2D objects is presented, and its properties are analyzed with respect to retrieval of similar objects in an image database. Our measure profits from a novel approach to subdivision of objects into parts of visual form. To compute our similarity measure, we first establish the best possible correspondence of visual parts, which is based on a correspondence of convex boundary arcs. Then the similarity between corresponding arcs is computed and aggregated. We applied our similarity measure to shape matching of object contours in various image databases and compared it to well-known approaches in the literature. The experimental results justify that our shape matching procedure gives an intuitive shape correspondence and is stable with respect to noise distortions.
Longin Jan Latecki, Rolf Lakämper
WACV1
1998 Supportedness and tameness differentialless geometry of plane curves
Longin Jan Latecki, Azriel Rosenfeld
Pattern Recognit.1
1997 3D Well-Composed Pictures
Longin Jan Latecki
CVGIP Graph. Model. Image Process.1
1997 A Realistic Digitization Model of Straight Lines
Ari D. Gross, Longin Jan Latecki
Comput. Vis. Image Underst.2
1996 Modelling digital straight lines
abstract
We present a realistic mathematical model of a digitized edge which handles both blurring and arbitrary thresholding. We show that a thresholded digital image of a blurred half-plane obtained for some unknown threshold value is equal to the image of a perfectly focused half-plane with the same slope obtained by object boundary quantisation. This result implies that recovering the slope of a blurred half-plane, given its image obtained for some unknown threshold value, reduces to recovering the slope of a perfectly focused half-plane under object boundary quantization. Therefore, the previous results and algorithms for the recovery of straight lines, which mainly assumed object boundary quantization, are also valid if we assume this more realistic digitization model.
Ari D. Gross, Longin Jan Latecki
ICPR2
1996 An Algorithm for a 3D Simplicity Test
Longin Jan Latecki, C. Min Ma
Comput. Vis. Image Underst.1
1995 Digitizations Preserving Topological and Differential Geometric Properties
Ari D. Gross, Longin Jan Latecki
Comput. Vis. Image Underst.2
1995 Well-Composed Sets
Longin Jan Latecki, Ulrich Eckhardt, Azriel Rosenfeld
Comput. Vis. Image Underst.1
1995 Generalized convexity: CP3 and boundaries of convex sets
Longin Jan Latecki, Azriel Rosenfeld, Ruth Silverman
Pattern Recognit.1
1995 Multicolor well-composed pictures
Longin Jan Latecki
Pattern Recognit. Lett.1
1995 Semi-proximity continuous functions in digital images
Longin Jan Latecki, Frank Prokop
Pattern Recognit. Lett.1
1993 Orientation and Qualitative Angle for Spatial Reasoning
Longin Jan Latecki, Ralf Röhrig
IJCAI1
1992 Connection Relations and Quantifier Scope
abstract
A formalism will be presented in this paper which makes it possible to realise the idea of assigning only one scope-ambiguous representation to a sentence that is ambiguous with regard to quantifier scope. The scope determination results in extending this representation with additional context and world knowledge conditions. If there is no scope determining information, the formalism can work further with this scope-ambiguous representation. Thus scope information does not have to be completely determined.
Longin Jan Latecki
ACL1
1992 On Hybrid Reasoning for Processing Spatial Expressions
Longin Jan Latecki, Simone Pribbenow
ECAI1
1991 An Indexing Technique For Implementing Command Relations
Longin Jan Latecki
EACL1