Suya You

dblp:98/3451 · DBLP profile ↗
← Back
96ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0002-6387-7024ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 81 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 31 · 14 since 2021Human-computer interaction and ubiquitous computing · 16 · 4 first-author · 1 since 2021Systems, architecture and hardware · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Descrip3D: Enhancing Large Language Model-based 3D Scene Understanding with Object-Level Text Descriptions
abstract
Understanding 3D scenes goes beyond simply recognizing objects; it requires reasoning about the spatial and semantic relationships between them. Current 3D scenelanguage models often struggle with this relational understanding, particularly when visual embeddings alone do not adequately convey the roles and interactions of objects. In this paper, we introduce Descrip3D, a novel and powerful framework that explicitly encodes the relationships between objects using natural language. Unlike previous methods that rely only on 2D and 3D embeddings, Descrip3D enhances each object with a textual description that captures both its intrinsic attributes and contextual relationships. These relational cues are incorporated into the model through a dual-level integration: embedding fusion and prompt-level injection. This allows for unified reasoning across various tasks such as grounding, captioning, and question answering, all without the need for task-specific heads or additional supervision. When evaluated on five benchmark datasets, including ScanRefer, Multi3DRefer, ScanQA, SQA3D, and Scan2Cap, Descrip3D consistently outperforms strong baseline models, demonstrating the effectiveness of language-guided relational representation for understanding complex indoor scenes.
Jintang Xue, Ganning Zhao, Jie-En Yao, Hong-En Chen, Meida Chen, Suya You, C.-C. Jay Kuo
WACV7
2026 An Integrated Dual-Motion Framework for Camouflaged Object Detection in Videos
abstract
Video Camouflaged Object Detection (VCOD) aims to segment objects visually indistinguishable from their background in unconstrained video sequences. This task is challenging due to the extremely low appearance contrast and motion ambiguity caused by camera shifts, occlusion, and background clutter. Although motion cues provide crucial signals to break camouflage, existing methods typically adopt either implicit motion modeling based on initial semantic representations or explicit modeling based on low-level pixel motion, both of which suffer from information loss or noise sensitivity. To this end, we propose IDM-VCOD, an integrated dual-motion VCOD framework that unifies two complementary motion modeling strategies through a dual-motion integrated design. The implicit path aggregates temporal context via spatio-temporal prediction cubes for temporal neighborhood refinement, while the explicit path explicitly aligns multi-frame backgrounds to highlight foreground motion. A selective activation mechanism adaptively triggers the explicit branch only when implicit predictions are unreliable, enhancing accuracy and efficiency. Extensive experiments on the MoCA-Mask and CAD datasets demonstrate that IDM-VCOD achieves superior detection accuracy and generalization compared to state-of-the-art VCOD methods, while significantly reducing model size and inference cost. Comprehensive ablation studies further validate the effectiveness of the dual-motion design and its robustness in challenging camouflage scenarios.
Hong-Shuo Chen, Zhiruo Zhou, Suya You, Azad M. Madni, C.-C. Jay Kuo
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
abstract
Recent advancements in 2D and multimodal models have achieved remarkable success by leveraging large-scale training on extensive datasets. However, extending these achievements to enable free-form interactions and high-level semantic operations with complex 3D/4D scenes remains challenging. This difficulty stems from the limited availability of large-scale, annotated 3D/4D or multi-view datasets, which are crucial for generalizable vision and language tasks such as open-vocabulary and prompt-based segmentation, language-guided editing, and visual question answering (VQA). In this paper, we introduce Feature4X, a universal framework designed to extend any functionality from 2D vision foundation model into the 4D realm, using only monocular video input, which is widely available from user-generated content. The "X" in Feature4X represents its versatility, enabling any task through adaptable, model-conditioned 4D feature field distillation. At the core of our framework is a dynamic optimization strategy that unifies multiple model capabilities into a single representation. Additionally, to the best of our knowledge, Feature4X is the first method to distill and lift the features of video foundation models (e.g. SAM2, InternVideo2) into an explicit 4D feature field using Gaussian Splatting. Our experiments showcase novel view segment anything, geometric and appearance scene editing, and free-form VQA across all time steps, empowered by LLMs in feedback loops. These advancements broaden the scope of agentic AI applications by providing a foundation for scalable, contextually and spatiotemporally aware systems capable of immersive dynamic 4D scene interaction.
Shijie Zhou 0003, Yijia Weng, Shuwang Zhang, Zhen Wang 0058, Dejia Xu, Zhiwen Fan, Suya You, Zhangyang Wang, Leonidas J. Guibas, Achuta Kadambi
CVPR8
2025 AllTracker: Efficient Dense Point Tracking at High Resolution
abstract
We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing point tracking methods, our approach delivers high-resolution and dense (all-pixel) correspondence fields, which can be visualized as flow maps. Unlike existing optical flow methods, our approach corresponds one frame to hundreds of subsequent frames, rather than just the next frame. We develop a new architecture for this task, blending techniques from existing work in optical flow and point tracking: the model performs iterative inference on low-resolution grids of correspondence estimates, propagating information spatially via 2D convolution layers, and propagating information temporally via pixel-aligned attention layers. The model is fast and parameter-efficient (16 million parameters), and delivers state-of-the-art point tracking accuracy at high resolution (i.e., tracking 768x1024 pixels, on a 40G GPU). A benefit of our design is that we can train jointly on optical flow datasets and point tracking datasets, and we find that doing so is crucial for top performance. We provide an extensive ablation study on our architecture details and training recipe, making it clear which details matter most. Our code and model weights are available at https://alltracker.github.io
Adam W. Harley, Yang You 0004, Xinglong Sun, Nikhil Raghuraman, Yunqi Gu, Sheldon Liang, Wen-Hsuan Chu, Achal Dave, Suya You, Rares Ambrus, Katerina Fragkiadaki, Leonidas J. Guibas
ICCV10
2025 Img2CAD: Reverse Engineering 3D CAD Models from Images through VLM-Assisted Conditional Factorization
abstract
Reverse engineering 3D computer-aided design (CAD) models from images is an important task for many downstream applications including interactive editing, manufacturing, architecture, robotics, etc. The difficulty of the task lies in vast representational disparities between the CAD output and the image input. CAD models are precise, programmatic constructs that involves sequential operations combining discrete command structure with continuous attributes – making it challenging to learn and optimize in an end-to-end fashion. Concurrently, input images introduce inherent challenges such as photometric variability and sensor noise, complicating the reverse engineering process. In this work, we introduce a novel approach that conditionally factorizes the task into two sub-problems. First, we leverage vision-language foundation models (VLMs), a finetuned Llama3.2, to predict the global discrete base structure with semantic information. Second, we propose TrAssembler that conditioned on the discrete structure with semantics predicts the continuous attribute values. To support the training of our TrAssembler, we further constructed an annotated CAD dataset of common objects from ShapeNet. Putting all together, our approach and data demonstrate significant first steps towards CAD-ifying images in the wild. Code and data can be found in https://github.com/qq456cvb/Img2CAD.
Yang You 0004, Mikaela Angelina Uy, Jiaqi Han 0001, Rahul Krishna Thomas, Haotong Zhang 0005, Hansheng Chen 0001, Francis Engelmann, Suya You, Leonidas J. Guibas
SIGGRAPH Asia9
2025 DDS: Decoupled Dynamic Scene-Graph Generation Network
abstract
Scene-graph generation involves creating a structural representation of the relationships between objects in a scene by predicting subject-object-relation triplets from input data. Existing methods show poor performance in detecting triplets outside of a predefined set, primarily due to their reliance on dependent feature learning. To address this issue we propose DDS– a decoupled dynamic scene-graph generation network- that consists of two independent branches that can disentangle extracted features. The key innovation of the current paper is the decoupling of the features representing the relationships from those of the objects, which enables the detection of novel object-relationship combinations. The DDS model is evaluated on three datasets and outperforms previous methods by a significant margin, especially in detecting previously unseen triplets.
A S. M. Iftekhar, Raphael Ruschel, Suya You, B. S. Manjunath
WACV4
2025 Efficient human-object-interaction (EHOI) detection via interaction label coding and Conditional Decision
Tsung-Shan Yang, Chengwei Wei, Suya You, C.-C. Jay Kuo
Comput. Vis. Image Underst.4
2024 Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
abstract
3D scene representations have gained immense popularity in recent years. Methods that use Neural Radiance fields are versatile for traditional tasks such as novel view synthesis. In recent times, some work has emerged that aims to extend the functionality of NeRF beyond view synthesis, for semantically aware tasks such as editing and segmentation using 3D feature field distillation from 2D foundation models. However, these methods have two major limitations: (a) they are limited by the rendering speed of NeRF pipelines, and (b) implicitly represented feature fields suffer from continuity artifacts reducing feature quality. Recently, 3D Gaussian Splatting has shown state-of-the-art performance on real-time radiance field rendering. In this work, we go one step further: in addition to radiance field rendering, we enable 3D Gaussian splatting on arbitrary-dimension semantic features via 2D foundation model distillation. This translation is not straightforward: naively incorporating feature fields in the 3DGS framework encounters significant challenges, notably the disparities in spatial resolution and channel consistency between RGB images and feature maps. We propose architectural and training changes to efficiently avert this problem. Our proposed method is general, and our experiments showcase novel view semantic segmentation, language-guided editing and segment anything through learning feature fields from state-of-the-art 2D foundation models such as SAM and CLIP-LSeg. Across experiments, our distillation method is able to provide comparable or better results, while being significantly faster to both train and render. Additionally, to the best of our knowledge, we are the first method to enable point and bounding-box prompting for radiance field manipulation, by leveraging the SAM model. Project website at: https://feature-3dgs.github.io/.
Shijie Zhou 0003, Sicheng Jiang, Zhiwen Fan, Zehao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, Achuta Kadambi
CVPR8
2024 DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
Shijie Zhou 0003, Zhiwen Fan, Dejia Xu, Pradyumna Chari, Tejas Bharadwaj, Suya You, Zhangyang Wang, Achuta Kadambi
ECCV (6)7
2024 Tokenmotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection VIA Learnable Token Selection
abstract
The area of Video Camouflaged Object Detection (VCOD) presents unique challenges in the field of computer vision due to texture similarities between target objects and their surroundings, as well as irregular motion patterns caused by both objects and camera movement. In this paper, we introduce TokenMotion (TMNet), which employs a transformer-based model to enhance VCOD by extracting motion-guided features using a learnable token selection. Evaluated on the challenging MoCA-Mask dataset, TMNet achieves state-of-the-art performance in VCOD. It outperforms the existing state-of-the-art method by a 12.8% improvement in weighted F-measure, an 8.4% enhancement in S-measure, and a 10.7% boost in mean IoU. The results demonstrate the benefits of utilizing motion-guided features via learnable token selection within a transformer-based framework to tackle the intricate task of VCOD. The code of our work will be available when the paper is accepted.
Zifan Yu, Erfan Bank Tavakoli, Meida Chen, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren
ICASSP4
2024 SemST: Semantically Consistent Multi-Scale Image Translation via Structure-Texture Alignment
abstract
Unsupervised image-to-image translation learns cross-domain image mapping that transfers input from the source domain to output in the target domain while preserving its semantics. One challenge is that different semantic statistics in source and target domains result in content discrepancy known as semantic distortion. To address this problem, a novel I2I method that maintains semantic consistency in translation is proposed and named SemST in this work. SemST reduces semantic distortion by employing contrastive learning and aligning the structural and textural properties of input and output by maximizing their mutual information. Furthermore, a multi-scale approach is introduced to enhance translation performance, thereby enabling the applicability of SemST to domain adaptation in high-resolution images. Experiments show that SemST effectively mitigates semantic distortion and achieves state-of-the-art performance. Also, the application of SemST to domain adaptation is explored. It is demonstrated by preliminary experiments that SemST can be utilized as a beneficial pre-training for the semantic segmentation task.
Ganning Zhao, Wenhui Cui, Suya You, C.-C. Jay Kuo
WACV3
2024 Geometrical Interpretation and Design of Multilayer Perceptrons
abstract
The multilayer perceptron (MLP) neural network is interpreted from the geometrical viewpoint in this work, that is, an MLP partition an input feature space into multiple nonoverlapping subspaces using a set of hyperplanes, where the great majority of samples in a subspace belongs to one object class. Based on this high-level idea, we propose a three-layer feedforward MLP (FF-MLP) architecture for its implementation. In the first layer, the input feature space is split into multiple subspaces by a set of partitioning hyperplanes and rectified linear unit (ReLU) activation, which is implemented by the classical two-class linear discriminant analysis (LDA). In the second layer, each neuron activates one of the subspaces formed by the partitioning hyperplanes with specially designed weights. In the third layer, all subspaces of the same class are connected to an output node that represents the object class. The proposed design determines all MLP parameters in a feedforward one-pass fashion analytically without backpropagation. Experiments are conducted to compare the performance of the traditional backpropagation-based MLP (BP-MLP) and the new FF-MLP. It is observed that the FF-MLP outperforms the BP-MLP in terms of design time, training time, and classification performance in several benchmarking datasets. Our source code is available at https://colab.research.google.com/drive/1Gz0L8AnT4ijrUchrhEXXsnaacrFdenn?usp = sharing.
Ruiyuan Lin, Zhiruo Zhou, Suya You, Raghuveer M. Rao, C.-C. Jay Kuo
IEEE Trans. Neural Networks Learn. Syst.3
2023 ALTO: Alternating Latent Topologies for Implicit 3D Reconstruction
abstract
This work introduces alternating latent topologies (ALTO) for high-fidelity reconstruction of implicit 3D surfaces from noisy point clouds. Previous work identifies that the spatial arrangement of latent encodings is important to recover detail. One school of thought is to encode a latent vector for each point (point latents). Another school of thought is to project point latents into a grid (grid latents) which could be a voxel grid or triplane grid. Each school of thought has tradeoffs. Grid latents are coarse and lose high-frequency detail. In contrast, point latents preserve detail. However, point latents are more difficult to decode into a surface, and quality and runtime suffer. In this paper, we propose ALTO to sequentially alternate between geometric representations, before converging to an easy-to-decode latent. We find that this preserves spatial expressiveness and makes decoding lightweight. We validate ALTO on implicit 3D recovery and observe not only a performance improvement over the state-of-the-art, but a runtime improvement of$3-10\times$. Project website at https://visual.ee.ucla.edu/alto.htm/.
Zhen Wang 0058, Shijie Zhou 0003, Jeong Joon Park, Despoina Paschalidou, Suya You, Gordon Wetzstein, Leonidas J. Guibas, Achuta Kadambi
CVPR5
2023 Enhanced Low-Resolution LiDAR-Camera Calibration via Depth Interpolation and Supervised Contrastive Learning
abstract
Motivated by the increasing application of low-resolution LiDAR, we target the problem of low-resolution LiDAR-camera calibration in this work. The main challenges are two-fold: sparsity and noise in point clouds. To address the problem, we propose to apply depth interpolation to increase the point density and supervised contrastive learning to learn noise-resistant features. The experiments on RELLIS-3D demonstrate that our approach achieves an average mean absolute rotation/translation errors of 0.15cm/0.33° on 32-channel LiDAR point cloud data, which significantly outperforms all reference methods.
Zifan Yu, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren
ICASSP3
2023 SUMMIT: Source-Free Adaptation of Uni-Modal Models to Multi-Modal Targets
abstract
Scene understanding using multi-modal data is necessary in many applications, e.g., autonomous navigation. To achieve this in a variety of situations, existing models must be able to adapt to shifting data distributions without arduous data annotation. Current approaches assume that the source data is available during adaptation and that the source consists of paired multi-modal data. Both these assumptions may be problematic for many applications. Source data may not be available due to privacy, security, or economic concerns. Assuming the existence of paired multi-modal data for training also entails significant data collection costs and fails to take advantage of widely available freely distributed pre-trained uni-modal models. In this work, we relax both of these assumptions by addressing the problem of adapting a set of models trained independently on uni-modal data to a target domain consisting of unlabeled multi-modal data, without having access to the original source dataset. Our proposed approach solves this problem through a switching framework which automatically chooses between two complementary methods of cross-modal pseudo-label fusion – agreement filtering and entropy weighting – based on the estimated domain gap. We demonstrate our work on the semantic segmentation problem. Experiments across seven challenging adaptation scenarios verify the efficacy of our approach, achieving results comparable to, and in some cases outperforming, methods which assume access to source data. Our method achieves an improvement in mIoU of up to 12% over competing baselines. Our code is publicly available at https://github.com/csimo005/SUMMIT.
Cody Simons, Dripta S. Raychaudhuri, Sk Miraj Ahmed, Suya You, Konstantinos Karydis, Amit K. Roy-Chowdhury
ICCV4
2023 LGSQE: Lightweight Generated Sample Quality Evaluation
abstract
Despite prolific work on evaluating generative models, little research has been done on the quality evaluation of an individual generated sample. To address this problem, a lightweight generated sample quality evaluation (LGSQE) method is proposed in this work. In the training stage of LGSQE, a binary classifier is trained on real and synthetic samples, where real and synthetic data are labeled by 0 and 1, respectively. In the inference stage, the classifier assigns soft labels (ranging from 0 to 1) to each generated sample. The value of the soft label indicates the quality level; namely, the quality is better if its soft label is closer to 0. LGSQE can serve as a post-processing module for quality control. Furthermore, LGSQE can be used to evaluate the performance of generative models, such as accuracy, AUC, precision, and recall, by aggregating sample-level quality. Experiments are conducted on several datasets and generative models to demonstrate that LGSQE can preserve the same performance rank order as that predicted by the Fréchet Inception Distance (FID) but with significantly lower complexity.
Ganning Zhao, Vasileios Magoulianitis, Suya You, C.-C. Jay Kuo
ICIP3
2023 TransUPR: A Transformer-based Plug-and-Play Uncertain Point Refiner for LiDAR Point Cloud Semantic Segmentation
abstract
Common image-based LiDAR point cloud semantic segmentation (LiDAR PCSS) approaches have bottlenecks resulting from the boundary-blurring problem of convolution neural networks (CNNs) and quantitation loss of spherical projection. In this work, we propose a transformer-based plug-and-play uncertain point refiner, i.e., TransUPR, to refine selected uncertain points in a learnable manner, which leads to an improved segmentation performance. Uncertain points are sampled from coarse semantic segmentation results of 2D image segmentation where uncertain points are located close to the object boundaries in the 2D range image representation and 3D spherical projection background points. Following that, the geometry and coarse semantic features of uncertain points are aggregated by neighbor points in 3D space without adding expensive computation and memory footprint. Finally, the transformer-based refiner, which contains four stacked self-attention layers, along with an MLP module, is utilized for uncertain point classification on the concatenated features of self-attention layers. As the proposed refiner is independent of 2D CNNs, our TransUPR can be easily integrated into any existing image-based LiDAR PCSS approaches, e.g., CENet. Our TransUPR with the CENet achieves state-of-the-art performance, i.e., 68.2% mean Intersection over Union (mIoU) on the Semantic KITTI benchmark, which provides a performance improvement of 0.6% on the mIoU compared to the original CENet.
Zifan Yu, Meida Chen, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren
IROS4
2023 PrObeD: Proactive Object Detection Wrapper
abstract
Previous research in $2D$ object detection focuses on various tasks, including detecting objects in generic and camouflaged images. These works are regarded as passive works for object detection as they take the input image as is. However, convergence to global minima is not guaranteed to be optimal in neural networks; therefore, we argue that the trained weights in the object detector are not optimal. To rectify this problem, we propose a wrapper based on proactive schemes, PrObeD, which enhances the performance of these object detectors by learning a signal. PrObeD consists of an encoder-decoder architecture, where the encoder network generates an image-dependent signal termed templates to encrypt the input images, and the decoder recovers this template from the encrypted images. We propose that learning the optimum template results in an object detector with an improved detection performance. The template acts as a mask to the input images to highlight semantics useful for the object detector. Finetuning the object detector with these encrypted images enhances the detection performance for both generic and camouflaged. Our experiments on MS-COCO, CAMO, COD$10$K, and NC$4$K datasets show improvement over different detectors after applying PrObeD. Our models/codes are available at https://github.com/vishal3477/Proactive-Object-Detection.
Vishal Asnani, Abhinav Kumar 0004, Suya You, Xiaoming Liu 0002
NeurIPS3
2023 An Online Continuous Semantic Segmentation Framework With Minimal Labeling Efforts
abstract
The annotation load for a new dataset has been greatly decreased using domain adaptation based semantic segmentation, which iteratively constructs pseudo labels on unlabeled target data and retrains the network. However, realistic segmentation datasets are often imbalanced, with pseudo-labels tending to favor certain "head" classes while neglecting other "tail" classes. This can lead to an inaccurate and noisy mask. To address this issue, we propose a novel hard sample mining strategy for an active domain adaptation based semantic segmentation network, with the aim of automatically selecting a small subset of labeled target data to fine-tune the network. By calculating class-wise entropy, we are able to rank the difficulty level of different samples. We use a fusion of focal loss and regional mutual information loss instead of cross-entropy loss for the domain adaptation based semantic segmentation network. Our entire framework has been implemented in real-time using the Robotics Operating System (ROS) with a server PC and a small Unmanned Ground Vehicle (UGV) known as the ROSbot2.0 Pro. This implementation allows ROSbot2.0 Pro to access any type of data at any time, enabling it to perform a variety of tasks with ease. Our approach has been thoroughly evaluated through a series of extensive experiments, which demonstrate its superior performance compared to existing state-of-the-art methods. Remarkably, by using just 20% of hard samples for fine-tuning, our network has achieved a level of performance that is comparable (≈88%) to that of a fully supervised approach, with mIOU scores of 60.51% in the In-house dataset.
Masud Ahmed, Zahid Hasan 0001, Tim Yingling, Eric O'Leary, Sanjay Purushotham, Suya You, Nirmalya Roy
SMARTCOMP6
2022 GADAN: Generative Adversarial Domain Adaptation Network For Debris Detection Using Drone
abstract
In maritime, coastal, and riverine environments debris has become an abundant pollutant, posing a significant threat to aquatic life. Existing debris detection systems are mostly designed to detect debris for specific type of environment. Therefore, we propose GADAN, an autonomous debris detection model fusing RGB and thermal images, and experiment on both over land and aquatic environment. We collected data using the Parrot Anafi Thermal drone from different water streams, and construction sites near the aquatic environment. We employ a generative domain adaptation based architecture to generate the thermal image compatible with the RGB image. We postulate a two-stream network on these thermal and RGB pairs of images separately, and pass the concatenated features to object detecting YOLO model. We also demonstrated the performance of debris detection using only RGB and only thermal images. Our proposed GADAN network incorporating RGB and thermal images outperforms (with a mean average precision of 0.97) RCNN and single modality based models.
Masud Ahmed, Naima Khan, Pretom Roy Ovi, Nirmalya Roy, Sanjay Purushotham, Aryya Gangopadhyay, Suya You
DCOSS7
2022 Not Just Streaks: Towards Ground Truth for Single Image Deraining
Yunhao Ba, Howard Zhang, Ethan Yang, Akira Suzuki 0002, Arnold Pfahnl, Chethan Chinder Chandrappa, Celso de Melo, Suya You, Stefano Soatto, Alex Wong 0001, Achuta Kadambi
ECCV (7)8
2022 Towards Scalable and Efficient Client Selection for Federated Object Detection
abstract
Various computer vision techniques based on deep neural networks have been proposed to detect objects accurately and fast. However, due to the privacy, security and communication bandwidth restrictions of diverse participating parties, it is sometimes prohibitive to train such models on a centralized machine. Federated Learning (FL) provides a promising solution to learn a model from decentralized data. Despite the advances in FL, the diversity of client regions in which they operate and the Non-IID nature of the crowdsourced datasets reduces the accuracy of object detection models significantly. In this paper, we introduce a novel FL object detection system to efficiently train models with heterogeneous client datasets. We propose lightweight client selection methods to learn object detection models faster. Our client selection methods based on the object data distribution at clients achieves up to 74% reduction in required federated rounds compared to conventional approaches. We further extend this method by leveraging the metadata of the training images (e.g., location, direction, depth), to select clients which maximize the coverage of diverse geographical regions. We report on extensive experiments with real datasets.
George Constantinou, Suya You, Cyrus Shahabi
ICPR2
2022 Fake Satellite Image Detection via Parallel Subspace Learning (PSL)
abstract
A new method, called PSL-DefakeHop, is proposed to detect fake satellite images based on the parallel subspace learning (PSL) framework in this work. The DefakeHop method was developed previously for detection of deepfake generated faces under the successive subspace learning (SSL) framework. PSL is proposed to extract features from responses of multiple single-stage filter banks (or called PixelHops), which operate in parallel, and it improves SSL that extracts features from multi-stage cascaded filter banks. PSL has two advantages. First, PSL preserves discriminant features often lie in high-frequency channels, which are however ignored by SSL. Second, decisions from multiple filter banks can be ensembled to further improve detection accuracy. To demonstrate the effectiveness of the proposed PSL-DefakeHop method, we evaluate it on the UW Fake Satellite Image dataset and observe perfect classification performance (i.e., 100% F1 score, precision and recall).
Hong-Shuo Chen, Kaitai Zhang, Shuowen Hu, Suya You, C.-C. Jay Kuo
ISCAS4
2022 GUSOT: Green and Unsupervised Single Object Tracking for Long Video Sequences
abstract
Supervised and unsupervised deep trackers that rely on deep learning technologies are popular in recent years. Yet, they demand high computational complexity and a high memory cost. A green unsupervised single-object tracker, called GUSOT, that aims at object tracking for long videos under a resource-constrained environment is proposed in this work. Built upon a baseline tracker, UHP-SOT++, which works well for short-term tracking, GUSOT contains two additional new modules: 1) lost object recovery, and 2) color-saliency-based shape proposal. They help resolve the tracking loss problem and offer a more flexible object proposal, respectively. Thus, they enable GUSOT to achieve higher tracking accuracy in the long run. We conduct experiments on the large-scale dataset LaSOT with long video sequences, and show that GUSOT offers a lightweight high-performance tracking solution that finds applications in mobile and edge computing platforms.
Zhiruo Zhou, Hongyu Fu, Suya You, C.-C. Jay Kuo
MMSP3
2022 Meta-UDA: Unsupervised Domain Adaptive Thermal Object Detection using Meta-Learning
abstract
Object detectors trained on large-scale RGB datasets are being extensively employed in real-world applications. However, these RGB-trained models suffer a performance drop under adverse illumination and lighting conditions. Infrared (IR) cameras are robust under such conditions and can be helpful in real-world applications. Though thermal cameras are widely used for military applications and increasingly for commercial applications, there is a lack of robust algorithms to robustly exploit the thermal imagery due to the limited availability of labeled thermal data. In this work, we aim to enhance the object detection performance in the thermal domain by leveraging the labeled visible domain data in an Unsupervised Domain Adaptation (UDA) setting. We propose an algorithm agnostic meta-learning framework to improve existing UDA methods instead of proposing a new UDA strategy. We achieve this by meta-learning the initial condition of the detector, which facilitates the adaptation process with fine updates without overfitting or getting stuck at local optima. However, meta-learning the initial condition for the detection scenario is computationally heavy due to long and intractable computation graphs. Therefore, we propose an online meta-learning paradigm which performs online updates resulting in a short and tractable computation graph. To this end, we demonstrate the superiority of our method over many baselines in the UDA setting, producing a state-of-the-art thermal detector for the KAIST and DSIAC datasets.
Vibashan VS, Domenick Poster, Suya You, Shuowen Hu, Vishal M. Patel
WACV3
2021 DefakeHop: A Light-Weight High-Performance Deepfake Detector
abstract
A light-weight high-performance Deepfake detection method, called DefakeHop, is proposed in this work. State-of-the-art Deepfake detection methods are built upon deep neural networks. DefakeHop uses the successive subspace learning (SSL) principle to extracts features automatically from various parts of face images. The features are extracted by channel-wise (c/w) Saab transform and further processed by our feature distillation module using spatial dimension re-duction and soft classification for each channel to get a more concise description of the face. Extensive experiments are conducted to demonstrate the effectiveness of the proposed DefakeHop method. With a small model size of 42,845 parameters, DefakeHop achieves state-of-the-art performance with the area under the ROC curve (AUC) of 100%, 94.95%, and 90.56% on UADFV, Celeb-DF v1, and Celeb-DF v2 datasets, respectively. Our codes are available on GitHub1.
Hong-Shuo Chen, Mozhdeh Rouhsedaghat, Hamza Ghani, Shuowen Hu, Suya You, C.-C. Jay Kuo
ICME5
2021 UHP-SOT: An Unsupervised High-Performance Single Object Tracker
abstract
An unsupervised online object tracking method that exploits both foreground and background correlations is proposed and named UHP-SOT (Unsupervised High-Performance Single Object Tracker) in this work. UHP-SOT consists of three modules: 1) appearance model update, 2) background motion modeling, and 3) trajectory-based box prediction. A state-of-the-art discriminative correlation filters (DCF) based tracker is adopted by UHP-SOT as the first module. We point out shortcomings of using the first module alone such as failure in recovering from tracking loss and inflexibility in object box adaptation and then propose the second and third modules to overcome them. Both are novel in single object tracking (SOT). We test UHP-SOT on two popular object tracking benchmarks, TB-50 and TB-100, and show that it outperforms all previous unsupervised SOT methods, achieves a performance comparable with the best supervised deep-learning-based SOT methods, and operates at a fast speed (i.e. 22.7-32.0 FPS on a CPU).
Zhiruo Zhou, Hongyu Fu, Suya You, Christoph Borel-Donohue, C.-C. Jay Kuo
VCIP3
2021 Low-resolution face recognition in resource-constrained environments
Mozhdeh Rouhsedaghat, Yifan Wang 0018, Shuowen Hu, Suya You, C.-C. Jay Kuo
Pattern Recognit. Lett.4
2021 On Relationship of Multilayer Perceptrons and Piecewise Polynomial Approximators
abstract
The relationship between a multilayer perceptron (MLP) regressor and a piecewise polynomial approximator is investigated in this work. We propose an MLP construction method, including the choice of activation, the specification of neuron numbers and filter weights. Through the construction, a one-to-one correspondence between an MLP and a piecewise polynomial is established. Especially, we point out that the form of nonlinear activation is related to the polynomial order. Since the approximation capability of piecewise polynomials is well understood, our study sheds new light on the universal approximation capability of an MLP.
Ruiyuan Lin, Suya You, Raghuveer M. Rao, C.-C. Jay Kuo
IEEE Signal Process. Lett.2
2020 Modeling Cross-Modal Interaction in a Multi-detector, Multi-modal Tracking Framework
Yiqi Zhong, Suya You, Ulrich Neumann
ACCV (2)2
2020 Pixelhop++: A Small Successive-Subspace-Learning-Based (Ssl-Based) Model For Image Classification
abstract
The successive subspace learning (SSL) principle was developed and used to design an interpretable learning model, known as the PixelHop method, for image classification in our prior work. Here, we propose an improved PixelHop method and call it Pixel-Hop++. First, to make the PixelHop model size smaller, we decouple a joint spatial-spectral input tensor to multiple spatial tensors (one for each spectral component) under the spatial-spectral separability assumption and perform the Saab transform in a channel-wise manner, called the channel-wise (c/w) Saab transform. Second, by performing this operation from one hop to another successively, we construct a channel-decomposed feature tree whose leaf nodes contain features of one dimension (1D). Third, these 1D features are ranked according to their cross-entropy values, which allows us to select a subset of discriminant features for image classification. In Pixel-Hop++, one can control the learning model size of fine-granularity, offering a flexible tradeoff between the model size and the classification performance. We demonstrate the flexibility of Pixel-Hop++ on MNIST, Fashion MNIST, and CIFAR-10 three datasets.
Yueru Chen, Mozhdeh Rouhsedaghat, Suya You, Raghuveer M. Rao, C.-C. Jay Kuo
ICIP3
2020 Object Detection on Monocular Images with Two- Dimensional Canonical Correlation Analysis
abstract
Accurate and robust detection of objects from monocular images is a fundamental vision task. This paper describes a novel approach of holistic scene understanding that can simultaneously achieve multiple tasks of scene reconstruction and object detection from a single monocular image. Rather than pursuing an independent solution for each individual task as most existing work does, we seek a globally optimal solution that holistically resolves the multiple perception and reasoning tasks in an effective manner. The approach explores the complementary properties of multimodal RGB images and depth data to improve scene perception tasks. It uniquely combines the techniques of canonical correlation analysis and deep learning to learn the most correlated features to maximize the modal cross-correlation for improving performance and robustness of object detection in complex environments. Extensive experiments have been conducted to evaluate and demonstrate the performances of proposed approach.
Zifan Yu, Suya You
ICPR2
2020 Visualization, Discriminability and Applications of Interpretable Saak Features
Abinaya Manimaran, Thiyagarajan Ramanathan, Suya You, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.3
2019 Robustness of Saak Transform Against Adversarial Attacks
abstract
Image classification is vulnerable to adversarial attacks. This work investigates the robustness of Saak transform against adversarial attacks towards high performance image classification. We develop a complete image classification system based on multi-stage Saak transform. In the Saak transform domain, clean and adversarial images demonstrate different distributions at different spectral dimensions. Selection of the spectral dimensions at every stage can be viewed as an automatic denoising process. Motivated by this observation, we carefully design strategies of feature extraction, representation and classification that increase adversarial robustness. The performances with well-known datasets and attacks are demonstrated by extensive experimental evaluations.
Thiyagarajan Ramanathan, Abinaya Manimaran, Suya You, C.-C. Jay Kuo
ICIP3
2019 Deep RGB-D Canonical Correlation Analysis For Sparse Depth Completion
abstract
In this paper, we propose our Correlation For Completion Network (CFCNet), an end-to-end deep learning model that uses the correlation between two data sources to perform sparse depth completion. CFCNet learns to capture, to the largest extent, the semantically correlated features between RGB and depth information. Through pairs of image pixels and the visible measurements in a sparse depth map, CFCNet facilitates feature-level mutual transformation of different data sources. Such a transformation enables CFCNet to predict features and reconstruct data of missing depth measurements according to their corresponding, transformed RGB features. We extend canonical correlation analysis to a 2D domain and formulate it as one of our training objectives (i.e. 2d deep canonical correlation, or “2D^2CCA loss"). Extensive experiments validate the ability and flexibility of our CFCNet compared to the state-of-the-art methods on both indoor and outdoor scenes with different real-life sparse patterns. Codes are available at: https://github.com/choyingw/CFCNet.
Yiqi Zhong, Cho-Ying Wu, Suya You, Ulrich Neumann
NeurIPS3
2018 Learning to Prune Filters in Convolutional Neural Networks
abstract
Many state-of-the-art computer vision algorithms use large scale convolutional neural networks (CNNs) as basic building blocks. These CNNs are known for their huge number of parameters, high redundancy in weights, and tremendous computing resource consumptions. This paper presents a learning algorithm to simplify and speed up these CNNs. Specifically, we introduce a “try-and-learn” algorithm to train pruning agents that remove unnecessary CNN filters in a data-driven way. With the help of a novel reward function, our agents removes a significant number of filters in CNNs while maintaining performance at a desired level. Moreover, this method provides an easy control of the tradeoff between network performance and its scale. Performance of our algorithm is validated with comprehensive pruning experiments on several popular CNNs for visual recognition and semantic segmentation tasks.
Qiangui Huang, Shaohua Kevin Zhou, Suya You, Ulrich Neumann
WACV3
2017 Shape Inpainting Using 3D Generative Adversarial Network and Recurrent Convolutional Networks
abstract
Recent advances in convolutional neural networks have shown promising results in 3D shape completion. But due to GPU memory limitations, these methods can only produce low-resolution outputs. To inpaint 3D models with semantic plausibility and contextual details, we introduce a hybrid framework that combines a 3D Encoder-Decoder Generative Adversarial Network (3D-ED-GAN) and a Longterm Recurrent Convolutional Network (LRCN). The 3DED- GAN is a 3D convolutional neural network trained with a generative adversarial paradigm to fill missing 3D data in low-resolution. LRCN adopts a recurrent neural network architecture to minimize GPU memory usage and incorporates an Encoder-Decoder pair into a Long Shortterm Memory Network. By handling the 3D model as a sequence of 2D slices, LRCN transforms a coarse 3D shape into a more complete and higher resolution volume. While 3D-ED-GAN captures global contextual structure of the 3D shape, LRCN localizes the fine-grained details. Experimental results on both real-world and synthetic data show reconstructions from corrupted models result in complete and high-resolution 3D objects.
Weiyue Wang 0002, Qiangui Huang, Suya You, Ulrich Neumann
ICCV3
2017 Self-paced cross-modality transfer learning for efficient road segmentation
abstract
Accurate road segmentation is a prerequisite for autonomous driving. Current state-of-the-art methods are mostly based on convolutional neural networks (CNNs). Nevertheless, their good performance is at expense of abundant annotated data and high computational cost. In this work, we address these two issues by a self-paced cross-modality transfer learning framework with efficient projection CNN. To be specific, with the help of stereo images, we first tackle a relevant but easier task, i.e. free-space detection with well developed unsupervised methods. Then, we transfer these useful but noisy knowledge in depth modality to single RGB modality with self-paced CNN learning. Finally, we only need to fine-tune the CNN with a few annotated images to get good performance. In addition, we propose an efficient projection CNN, which can improve the fine-grained segmentation results with little additional cost. At last, we test our method on KITTI road benchmark. Our proposed method surpasses all published methods at a speed of 15fps.
Weiyue Wang 0002, Naiyan Wang, Suya You, Ulrich Neumann
ICRA4
2017 Mobile collaborative mixed reality for supporting scientific inquiry and visualization of earth science data
abstract
This work seeks to apply the emerging virtual and mixed reality techniques to visual exploration and visualization of earth science data. A novel system is developed to facilitate a collaborative mixed reality visualization, enabling both in-situ and off-site users to simultaneously interact with and visualize science data within mixed reality realm. We implement the prototype system in the context of visualizing earth terrain data. We report our current prototype effort and preliminary results.
Suya You, Charles K. Thompson
VR1
2016 Vehicle detection in urban point clouds with orthogonal-view convolutional neural network
abstract
In this paper, we aim at detecting vehicles from the point clouds scanned from the urban area. Our detection method consists of a segmentation stage and a classification stage. Prior knowledge for vehicles and urban environment is utilized to help the detection process. Specifically, we incorporate curb detection and removal in the segmentation stage. Moreover, our approach is able to estimate the orientation of the candidates and use it to handle the difficult cases such as the vehicles in the parking lot. In order to distinguish the vehicles from other segments among the 3D point cloud candidates, we develop three architectures of the orthogonal-view CNN, which are based on the orthogonal view projections of the candidates. Detailed evaluations and comparisons are performed on a challenging point cloud dataset of urban area.
Jing Huang 0020, Suya You
ICIP2
2016 Point cloud labeling using 3D Convolutional Neural Network
abstract
In this paper, we tackle the labeling problem for 3D point clouds. We introduce a 3D point cloud labeling scheme based on 3D Convolutional Neural Network. Our approach minimizes the prior knowledge of the labeling problem and does not require a segmentation step or hand-crafted features as most previous approaches did. Particularly, we present solutions for large data handling during the training and testing process. Experiments performed on the urban point cloud dataset containing 7 categories of objects show the robustness of our approach.
Jing Huang 0020, Suya You
ICPR2
2015 Pole-like object detection and classification from urban point clouds
abstract
This paper focuses on detecting and classifying pole-like objects from point clouds obtained in urban areas. To achieve our goal, we propose a system consisting of three stages: localization, segmentation and classification. The localization algorithm based on slicing, clustering, pole seed generation and bucket augmentation takes advantage of the unique characteristics of pole-like objects and avoids heavy computation on the feature of every point in traditional methods. Then, the bucket-shaped neighborhood of the segments is integrated and trimmed with region growing algorithms, reducing the noises within candidate's neighborhood. Finally, we introduce a representation of six attributes based on the height and five point classes closely related to the pole categories and apply SVM to classify the candidate objects into 4 categories, including 3 pole categories light, utility pole and sign, and the non-pole category. The performance of our method is demonstrated through comparison with previous works on a large-scale urban dataset.
Jing Huang 0020, Suya You
ICRA2
2015 Change Detection in Laser-Scanned Data of Industrial Sites
abstract
As laser scanners become widely used in 3D data acquisition of industrial sites, one challenging problem emerges: given two data of the same site scanned/modeled at different times, how can we tell the difference between the two? In this paper, we formulate this problem as the 3D change detection problem, and propose a novel method for detecting object-level changes. In general, we notice that the changes can be viewed as the inconsistency between the global alignment and the local alignment. Therefore, we propose a change detection framework that comprises global alignment, local object detection and a novel change detection method. Specifically, we propose a series of change evaluation functions for pair wise change inference, based on which we formulate the many-to-many object change correlation problem as the weighted bipartite matching problem which could be solved efficiently. Finally, we demonstrate the feasibility of our approach through experiments on both synthetic and real industrial datasets.
Jing Huang 0020, Suya You
WACV2
2014 A Unified Framework for Augmented Reality and Knowledge-Based Systems in Maintaining Aircraft
abstract
Aircraft maintenance and training play one of the most important roles in ensuring flight safety. The maintenance process usually involves massive numbers of components and substantial procedural knowledge of maintenance procedures. Maintenance tasks require technicians to follow rigorous procedures to prevent operational errors in the maintenance process. In addition, the maintenance time is a cost-sensitive issue for airlines. This paper proposes intelligent augmented reality (IAR) system to minimize operation errors and time-related costs and help aircraft technicians cope with complex tasks by using an intuitive UI/UX interface for their maintenance tasks. The IAR system is composed mainly of three major modules: 1) the AR module 2) the knowledge-based system (KBS) module 3) a unified platform with an integrated UI/UX module between the AR and KBS modules. The AR module addresses vision-based tracking, annotation, and recognition. The KBS module deals with ontology-based resources and context management. Overall testing of the IAR system is conducted at Korea Air Lines (KAL) hangars. Tasks involving the removal and installation of pitch trimmers in landing gear are selected for benchmarking purposes, and according to the results, the proposed IAR system can help technicians to be more effective and accurate in performing their maintenance tasks.
Kyeong-Jin Oh, Inay Ha, Kee-Sung Lee, Myung-Duk Hong, Ulrich Neumann, Suya You
AAAI7
2014 Segmentation and matching: Towards a robust object detection system
abstract
This paper focuses on detecting parts in laser-scanned data of a cluttered industrial scene. To achieve the goal, we propose a robust object detection system based on segmentation and matching, as well as an adaptive segmentation algorithm and an efficient pose extraction algorithm based on correspondence filtering. We also propose an overlapping-based criterion that exploits more information of the original point cloud than the number-of-matching criterion that only considers key-points. Experiments show how each component works and the results demonstrate the performance of our system compared to the state of the art.
Jing Huang 0020, Suya You
WACV2
2013 Detecting Objects in Scene Point Cloud: A Combinational Approach
abstract
Object detection is a fundamental task in computer vision. As the 3D scanning techniques become popular, directly detecting objects through 3D point cloud of a scene becomes an immediate need. We propose an object detection framework combining learning-Based classification, local descriptor, a new variance of RANSAC imposing rigid-body constraint and an iterative process for multi-object detection in continuous point clouds. The framework not only takes global and local information into account, but also benefits from both learning and empirical methods. The experiments performed on the challenging ground Lidar dataset show the effectiveness of our method.
Jing Huang 0020, Suya You
3DV2
2013 Robust Image Matching with Line Context
abstract
Some features are specially designed to cope with complex illumination variations [1, 2, 3, 4, 5]. These features are based on intensity orders rather than values so that their descriptors are invariant to non-linear intensity changes. However, one limitation of such features is that they are only able to handle monotonic illumination changes, as the examples shown in Figure 1. In this paper, we try to cope with this problem by proposing a new type of local feature, which describes the context information of neighboring line segments. Through extensive experiments and based on existing research, we found that the best pixel grouping level to handle large illumination variations is using curves or lines. Since curves require more sophisticated groupings and can be approximated by multiple line segments, we use line segments as the primitives of our feature descriptor. We call the proposed feature Line Context, since it is inspired by Shape Context. Line Context Detector. It is well known that edges are present at various scales. To detect edges at different scales we use multi-scale Canny edge detector with Gaussian derivatives at several pre-selected scales. Two phases are used to remove unstable edges. The first phase is to apply Laplacian operator. Those edges that do not attain a distinctive extremum over scales will be removed. The second phase is to remove those edges with homogeneous gradients, i.e. the points where the underlying curves have zero curvatures. We apply Harris matrix to achieve this purpose. Segments in Context. The edge pixels are linked to connected curves at different scales. These curves are then fitted by straight line segments. As shown in Figure 2-(a), several cases need to be considered for representing curves with line segments. One curve may be fitted by multiple segments like curve a. Two segments with small gap in between are merged into one larger segment no matter they are on the same scale (curve b and d) or different scales (curve f and g). In the latter case, the merged segment only exists in the lower scale (segment 8). Besides, all the segments in higher levels are also segments (segment 1, 2 and 3) or part of segments (segment 5 and 7) in lower levels. For each keypoint, we need to find line segments in its neighborhood, which is also called context of the feature. The line segments lying inside or partially inside the context are called context segments. The initial scale σ provides an estimate size of searching area. Let dpoint−set(v,S) be the shortest distance from point v to a set of points S. With keypoint k and all-segments {segi}, the segments set L in the context is defined as,
Wei Guan 0005, Suya You
BMVC2
2013 Estimation of camera pose with respect to terrestrial LiDAR data
abstract
In this paper, we present an algorithm that is to estimate the position of a hand-held camera with respect to terrestrial LiDAR data. Our input is a set of 3D range scans with intensities and one or a set of 2D uncalibrated camera images of the scene. The algorithm that automatically registers range scans and 2D images is composed of following steps. In the first step, we project the terrestrial LiDAR onto 2D images according to several preselected viewpoints. Intensity-based features such as SIFT are extracted from these projected images and these features are projected back onto the LiDAR data to obtain their 3D positions. In the second step, we estimate the initial pose of given 2D images from feature correspondences. In the third step, we refine the coarse camera pose obtained from the previous step through iterative matchings and optimization process. We presents results from experiments in several different urban settings.
Wei Guan 0005, Suya You, Guan Pang
WACV2
2012 Efficient matchings in augmented reality application
abstract
With fast growing popularity of smart phones in recent years, augmented reality (AR) becomes more demanding than ever before. However, one of main challenges is that while features like SIFT or SURF are robust in matchings, they are not computationally efficient. In this paper, we propose an efficient matching method for robust features. A distinctive descriptor is also proposed for performance improvements. Besides, we have developed an outdoor augmented reality system that is based on our proposed methods. The system demonstrates that not only it can achieve robust matchings efficiently, it is also capable to handle large occlusions such as passengers and moving vehicles.
Wei Guan 0005, Suya You, Ulrich Neumann
ICIP2
2012 Image relighting and matching with illumination information
abstract
In recent years, interest point based feature such as SIFT and SURF are widely used for image matchings. While these features are robust to changes in scales, rotations, and local geometric deformations, they are usually less able to simultaneously handle viewpoint changes and large illumination changes. In this paper, we will cope with this challenging problem. The basic idea is to relight one of the two images so that the image has similar illumination conditions as the other one. After relighting process, the point based features become effective again in the matching process. Our experiments show that the proposed method has good performance for matching two images with very different illuminations.
Wei Guan 0005, Suya You, Tanasai Sucontphunt, Ulrich Neumann
ICIP2
2012 Efficient matchings and mobile augmented reality
abstract
With the fast-growing popularity of smart phones in recent years, augmented reality (AR) on mobile devices is gaining more attention and becomes more demanding than ever before. However, the limited processors in mobile devices are not quite promising for AR applications that require real-time processing speed. The challenge exists due to the fact that, while fast features are usually not robust enough in matchings, robust features like SIFT or SURF are not computationally efficient. There is always a tradeoff between robustness and efficiency and it seems that we have to sacrifice one for the other. While this is true for most existing features, researchers have been working on designing new features with both robustness and efficiency. In this article, we are not trying to present a completely new feature. Instead, we propose an efficient matching method for robust features. An adaptive scoring scheme and a more distinctive descriptor are also proposed for performance improvements. Besides, we have developed an outdoor augmented reality system that is based on our proposed methods. The system demonstrates that not only it can achieve robust matchings efficiently, it is also capable to handle large occlusions such as passengers and moving vehicles, which is another challenge for many AR applications.
Wei Guan 0005, Suya You, Ulrich Neumann
ACM Trans. Multim. Comput. Commun. Appl.2
2011 GPS-aided recognition-based user tracking system with augmented reality in extreme large-scale areas
abstract
We present a recognition-based user tracking and augmented reality system that works in extreme large scale areas. The system will provide a user who captures an image of a building facade with precise location of the building and augmented information about the building. While GPS cannot provide information about camera poses, it is needed to aid reducing the searching ranges in image database. A patch-retrieval method is used for efficient computations and real-time camera pose recovery. With the patch matching as the prior information, the whole image matching can be done through propagations in an efficient way so that a more stable camera pose can be generated. Augmented information such as building names and locations are then delivered to the user. The proposed system mainly contains two parts, offline database building and online user tracking. The database is composed of images for different locations of interests. The locations are clustered into groups according to their UTM coordinates. An overlapped clustering method is used to cluster these locations in order to restrict the retrieval range and avoid ping pong effects. For each cluster, a vocabulary tree is built for searching the most similar view. On the tracking part, the rough location of the user is obtained from the GPS and the exact location and camera pose are calculated by querying patches of the captured image. The patch property makes the tracking robust to occlusions and dynamics in the scenes. Moreover, due to the overlapped clusters, the system simulates the "soft handoff" feature and avoid frequent swaps in memory resource. Experiments show that the proposed tracking and augmented reality system is efficient and robust in many cases.
Wei Guan 0005, Suya You, Ulrich Neumann
MMSys2
2011 Recognition-driven 3D navigation in large-scale virtual environments
abstract
We present a recognition-driven navigation system for large-scale 3D virtual environments. The proposed system contains three parts, virtual environment reconstruction, feature database building and recognition-based navigation. The virtual environment is reconstructed automatically with LIDAR data and aerial images. The feature database is composed of image patches with features and registered location and orientation information. The database images are taken at different distances from the scenes with various viewing angles, and these images are then partitioned into smaller patches. When a user navigates the real world with a handheld camera, the captured image is used to estimate its location and orientation. These location and orientation information are also reflected in the virtual environment. With the proposed patch approach, the recognition is robust to large occlusions and can be done in real time. Experiments show that our proposed navigation system is efficient and well synchronized with real world navigation.
Wei Guan 0005, Suya You, Ulrich Neumann
VR2
2011 Computationally efficient retrieval-based tracking system and augmented reality for large-scale areas
abstract
We present a retrieval-based tracking system that requires less computational time and cost. The system tracks a user's location through a small portion of an image captured by the camera, and then refines the camera pose by propagating matchings to the whole image. Augmented information such as building names and locations will be delivered to the user. The progressive way to process image data not only can provide the user with location information at real-time speed, but more importantly, it reduces the feature matching time by limiting the searching ranges. The proposed system contains two parts, offline database building and online user tracking. The database is composed of image patches with features and location information. The images are captured at different locations of interests from different viewing angles and distances, and then these images are partitioned into smaller patches. The location of a user can be calculated by querying one or more patches of the captured image. Moreover, the system is capable to handle large occlusions in images due to the patch approach. Experiments show that the proposed tracking system is efficient and robust in many different environments.
Wei Guan 0005, Suya You, Ulrich Neumann
WACV2
2011 Augmented distinctive features for efficient image matching
abstract
Finding corresponding image points is a challenging computer vision problem, especially for confusing scenes with surfaces of low textures or repeated patterns. Despite the well-known challenges of extracting conceptually meaningful high-level matching primitives, many recent works describe high-level image features such as edge groups, lines and regions, which are more distinctive than traditional local appearance based features, to tackle such difficult scenes. In this paper, we propose a different and more general approach, which treats the image matching problem as a recognition problem of spatially related image patch sets. We construct augmented semi-global descriptors (ordinal codes) based on subsets of scale and orientation invariant local keypoint descriptors. Tied ranking problem of ordinal codes is handled by increasingly keypoint sampling around image patch sets. Finally, similarities of augmented features are measured using Spearman correlation coefficient. Our proposed method is compatible with a large range of existing local image descriptors. Experimental results based on standard benchmark datasets and SURF descriptors have demonstrated its distinctiveness and effectiveness.
Quan Wang 0001, Wei Guan 0005, Suya You
WACV3
2010 Explore multiple clues for urban images matching
abstract
Many well-known existing image matching methods are based on local texture analysis, and consequently have difficulty handling low-textured 3D objects, such as those man-made buildings and road networks in urban scenes. In this paper, we propose our urban images matching method utilizing multiple novel clues. Specifically, we explore robust image features generated by interest regions and edge groups extracted from input urban images. Initial correspondences are established based on our features' similarity measurement. Cost ratio, combined with global context information, is used to remove outliers. We believe our hybrid approach is more suitable than traditional texture features for such scenes. The proposed method has been tested using real aerial urban photos under diverse viewing conditions. Experiments report that our registration success rate is nearly doubled compared with classical methods such as and.
Quan Wang 0001, Suya You
ICIP2
2009 A Dynamic Programming Approach to Maximizing Tracks for Structure from Motion
Jonathan Mooser, Suya You, Ulrich Neumann, Raphaël Grasset, Mark Billinghurst
ACCV (2)2
2009 Automatic reconstruction of cities from remote sensor data
abstract
In this paper, we address the complex problem of rapid modeling of large-scale areas and present a novel approach for the automatic reconstruction of cities from remote sensor data. The goal in this work is to automatically create lightweight, watertight polygonal 3D models from LiDAR data (Light Detection and Ranging) captured by an airborne scanner. This is achieved in three steps: preprocessing, segmentation and modeling, as shown in Figure 1. Our main technical contributions in this paper are: (i) a novel, robust, automatic segmentation technique based on the statistical analysis of the geometric properties of the data, which makes no particular assumptions about the input data, thus having no data dependencies, and (ii) an efficient and automatic modeling pipeline for the reconstruction of large-scale areas containing several thousands of buildings. We have extensively tested the proposed approach with several city-size datasets including downtown Baltimore, downtown Denver, the city of Atlanta, downtown Oakland, and we present and evaluate the experimental results.
Charalambos Poullis, Suya You
CVPR2
2009 Wide-baseline image matching using Line Signatures
abstract
We present a wide-baseline image matching approach based on line segments. Line segments are clustered into local groups according to spatial proximity. Each group is treated as a feature called a Line Signature. Similar to local features, line signatures are robust to occlusion, image clutter, and viewpoint changes. The descriptor and similarity measure of line signatures are presented. Under our framework, the feature matching is not only robust against affine distortion but also a considerable range of 3D viewpoint changes for non-planar surfaces. When compared to matching approaches based on existing local features, our method shows improved results with low-texture scenes. Moreover, extensive experiments validate that our method has advantages in matching structured non-planar scenes under large viewpoint changes and illumination variations.
Lu Wang 0013, Ulrich Neumann, Suya You
ICCV3
2009 PTZ camera calibration for Augmented Virtual Environments
abstract
Augmented Virtual Environments (AVE) are very effective in the application of surveillance, in which multiple video streams are projected onto a 3D urban model for better visualization and comprehension of the dynamic scenes. One of the key issues in creating such systems is to estimate the parameters of each camera including the intrinsic parameters and its pose relative to the 3D model. Nowadays, PTZ cameras are popular in this kind of applications. How to rapidly calibrate them at an arbitrary PTZ setting is not clear in the literature. We propose an efficient approach with two steps. In the first step, panoramic images are generated at a set of zooms. The images composing these panoramas are calibrated and stored in a database. In the second step, an image is acquired at an arbitrary PTZ setting. Its best matching image in the database is found by using an efficient local feature recognition technique. Based on this image, the camera parameters at the new PTZ setting can be estimated.
Lu Wang 0013, Suya You, Ulrich Neumann
ICME2
2009 Robust pose estimation in untextured environments for augmented reality applications
abstract
We present a robust camera pose estimation approach for stereo images captured in untextured environments. Unlike most of existing registration algorithms which are point-based and make use of intensities of pixels in the neighborhood, our approach imports line segments in registration process. With line segments as primitives, the proposed algorithm is capable to handle untextured images such as scenes captured in man-made environments, as well as the cases when there are large viewpoint changes or illumination changes. Furthermore, since the proposed algorithm is robust to large base-line stereos, there are improvements on the accuracy of 3D points reconstruction. With well-calculated camera pose and object positions in 3D space, we can embed virtual objects into existing scene with higher accuracy for realistic effects. In our experiments, 2D labels are embedded in the 3D scene space to achieve annotation effects as in AR.
Wei Guan 0005, Lu Wang 0013, Jonathan Mooser, Suya You, Ulrich Neumann
ISMAR4
2009 Automatic Creation of Massive Virtual Cities
abstract
This research effort focuses on the historically-difficult problem of creating large-scale (city size) scene models from sensor data, including rapid extraction and modeling of geometry models. The solution to this problem is sought in the development of a novel modeling system with a fully automatic technique for the extraction of polygonal 3D models from LiDAR (Light Detection And Ranging) data. The result is an accurate 3D model representation of the real-world as shown in Figure 1. We present and evaluate experimental results of our approach for the automatic reconstruction of large U.S. cities.
Charalambos Poullis, Suya You
VR2
2009 Applying robust structure from motion to markerless augmented reality
abstract
We demonstrate a complete system for markerless augmented reality using robust structure from motion. The proposed system includes two main components. The first is a means of learning the appearance of complex 3D objects and augmenting them with virtual annotations. Its output is database of recognizable landmarks along with 3D descriptions of accompanying virtual objects. The second component uses this data to recognize the previously learned landmarks, recover camera pose, and render the associated virtual content. Both components make use of the recently developed subtrack optimization algorithm for structure from motion, which we demonstrate to be a useful tool for both learning the structure of objects and tracking camera pose after recognition. The complete system is demonstrated on several complex real-world examples.
Jonathan Mooser, Suya You, Ulrich Neumann, Quan Wang 0001
WACV2
2009 A vision-based 2D-3D registration system
abstract
In this paper, we propose an automatic system for robust alignment of 2D optical images with 3D LiDAR (light detection and ranging) data. Focusing on applications such as data fusion and rapid updating for GIS (geographic information systems) from diverse sources when accurate georeference is not available, our goal is a vision-based approach to recover the 2D to 3D transformation including orientation, scale and location. Two major challenges of 2D-3D registration systems are different sensors problem and large 3D viewpoint changes. Due to the challenging nature of the problem, we achieve the registration goal through a two-stage solution. The first stage of the proposed system uses a robust region matching method to handle different sensor problems registering 3D data onto 2D images with similar viewing directions. Robustness to 3D viewpoint changes is achieved by the second stage where 2D views with sensor variations compensated by the first stage are trained. Other views are rapidly matched to them and therefore indirectly registered with 3D LiDAR data. Experimental results using four cities' datasets have illustrated the potential of the proposed system.
Quan Wang 0001, Suya You
WACV2
2009 Photorealistic Large-Scale Urban City Model Reconstruction
abstract
The rapid and efficient creation of virtual environments has become a crucial part of virtual reality applications. In particular, civil and defense applications often require and employ detailed models of operations areas for training, simulations of different scenarios, planning for natural or man-made events, monitoring, surveillance, games, and films. A realistic representation of the large-scale environments is therefore imperative for the success of such applications since it increases the immersive experience of its users and helps reduce the difference between physical and virtual reality. However, the task of creating such large-scale virtual environments still remains a time-consuming and manual work. In this work, we propose a novel method for the rapid reconstruction of photorealistic large-scale virtual environments. First, a novel, extendible, parameterized geometric primitive is presented for the automatic building identification and reconstruction of building structures. In addition, buildings with complex roofs containing complex linear and nonlinear surfaces are reconstructed interactively using a linear polygonal and a nonlinear primitive, respectively. Second, we present a rendering pipeline for the composition of photorealistic textures, which unlike existing techniques, can recover missing or occluded texture information by integrating multiple information captured from different optical sensors (ground, aerial, and satellite).
Charalambos Poullis, Suya You
IEEE Trans. Vis. Comput. Graph.2
2008 Supporting range and segment-based hysteresis thresholding in edge detection
abstract
One of the important steps in gradient-based edge detection is thresholding, e.g., hysteresis thresholding in the Canny detector. Traditional approaches only use gradient magnitude as the criterion to select edge pixels. We introduce a novel saliency measure called supporting range. It measures the range along the gradient direction of an edge pixel in which its gradient magnitude is a local maximum. As the result, our approach can detect the edge pixels on object boundaries even if they have very low gradient magnitude. In addition, unlike the Canny detector, the proposed hysteresis thresholding approach is based on segments instead of individual pixels. This also makes the approach more robust to detect edge pixels on weak boundaries.
Lu Wang 0013, Suya You, Ulrich Neumann
ICIP2
2008 CDIKP: A highly-compact local feature descriptor
abstract
A new feature descriptor is presented for object and scene recognition. The new approach, called CDIKP, uniquely combines the scale-invariant feature detection with a robust projection kernel technique to produce highly efficient feature representation. The produced feature descriptors are highly-compact in comparisons to the state-of-the-art, do not require any pre-training step, and show superior advantages in terms of distinctiveness, robustness to occlusions, invariance to scale, and tolerance of geometric distortions. We extensively evaluated the effectiveness of the new approach with various datasets acquired under varying circumstances.
Yun-Ta Tsai, Quan Wang 0001, Suya You
ICPR3
2008 Feature selection for real-time image matching systems
abstract
This paper proposes a general feature selection approach for real-time image matching systems. To demonstrate the idea¿s effectiveness, we focus on the issue of rotational invariance. Most current image matching methods compute and align local image patches to a uniform dominant orientation, which are either too computationally expensive for real-time systems or insufficiently robust. In contrast to current approaches, we combine multiple-view training and feature selection into a unified framework. The most invariant features are selected during an offline training stage. Therefore, no additional computation is needed for online processing. Furthermore the proposed ROTATION INVARIANT FEATURE SELECTION (RIFS) can be easily adapted to similar image matching problems such as scale invariance improvement and kernel selection in feature description. Experimental results show the effectiveness of RIFS using only a small number of training views. The proposed approach is also successfully integrated into an augmented reality application for museum exhibitions.
Quan Wang 0001, Suya You
ICPR2
2008 Large document, small screen: a camera driven scroll and zoom control for mobile devices
abstract
We present a three degree-of-freedom control designed for viewing large documents and images on a mobile device equipped with a camera. Tracking natural features detected in the camera's field of view, we can roughly estimate the motion of the device, using the results to scroll and zoom the current document. Central to our implementation is the manner by which we amplify the motion, allowing the user to scroll through large portions of the document with minimal hand movement. Then, using a Hidden Markov Model, we determine when the user is scrolling, zooming, or some combination of the two, thus providing smoother, more fluid control. We demonstrate a prototype of our 3DOF control that can easily navigate documents that are many times larger than the display area, and show how it might be incorporated into a larger document retrieval application.
Jonathan Mooser, Suya You, Ulrich Neumann
SI3D2
2008 Rapid Creation of Large-scale Photorealistic Virtual Environments
abstract
The rapid and efficient creation of virtual environments has become a crucial part of virtual reality applications. In particular, civil and defense applications often require and employ detailed models of operations areas for training, simulations of different scenarios, planning for natural or man-made events, monitoring, surveillance, games and films. A realistic representation of the large-scale environments is therefore imperative for the success of such applications since it increases the immersive experience of its users and helps reduce the difference between physical and virtual reality. However, the task of creating such large-scale virtual environments still remains a time-consuming and manual work. In this work we propose a novel method for the rapid reconstruction of photorealistic large-scale virtual environments. First, a novel parameterized geometric primitive is presented for the automatic building detection, identification and reconstruction of building structures. In addition, buildings with complex roofs containing non-linear surfaces are reconstructed interactively using a nonlinear primitive. Secondly, we present a rendering pipeline for the composition of photorealistic textures which unlike existing techniques it can recover missing or occluded texture information by integrating multiple information captured from different optical sensors (ground, aerial and satellite).
Charalambos Poullis, Suya You, Ulrich Neumann
VR2
2008 A Vision-Based System For Automatic Detection and Extraction Of Road Networks
abstract
In this paper we present a novel vision-based system for automatic detection and extraction of complex road networks from various sensor resources such as aerial photographs, satellite images, and LiDAR. Uniquely, the proposed system is an integrated solution that merges the power of perceptual grouping theory (Gabor filtering, tensor voting) and optimized segmentation techniques (global optimization using graph-cuts) into a unified framework to address the challenging problems of geospatial feature detection and classification. Firstly, the local precision of the Gabor filters is combined with the global context of the tensor voting to produce accurate classification of the geospatial features. In addition, the tensorial representation used for the encoding of the data eliminates the need for any thresholds, therefore removing any data dependencies. Secondly, a novel orientation-based segmentation is presented which incorporates the classification of the perceptual grouping, and results in segmentations with better defined boundaries and continuous linear segments. Finally, a set of Gaussian-based filters are applied to automatically extract centerline information (magnitude, width and orientation). This information is then used for creating road segments and then transforming them to their polygonal representations.
Charalambos Poullis, Suya You, Ulrich Neumann
WACV2
2007 Real-Time Image Matching Based on Multiple View Kernel Projection
abstract
This paper proposes a novel matching method for realtime finding the correspondences among different images containing the same object. The method utilizes an efficient Kernel Projection scheme to descript the image patch around a detected feature point. In order to achieve invariance and tolerance to geometric distortions, it combines a training stage based on generated synthetic views of the object. The two reliable and efficient methods cooperate together, resulting the core part of our novel multiple view kernel projection method (MVKP). Finally, considering the properties and distribution of the described feature vectors, we search for the best correspondence between two sets of features using a fast filtering vector approximation (FFVA) algorithm, which can be viewed as a fast lower-bound rejection scheme. Extensive experimental results on both synthetic and real data have demonstrated the effectiveness of the proposed approach.
Quan Wang 0001, Suya You
CVPR2
2007 Linear feature extraction using perceptual grouping and graph-cuts
abstract
In this paper we present a novel system for the detection and extraction of road map information from high-resolution satellite imagery.
Charalambos Poullis, Suya You, Ulrich Neumann
GIS2
2007 Semiautomatic registration between ground-level panoramas and an orthorectified aerial image for building modeling
abstract
Aerial imagery and ground-level imagery are two complementary data sources for architectural modeling. How to integrate them is a critical issue in creating complete, photo-realistic and large-scale urban models. We describe a semiautomatic approach of detecting feature correspondences between ground-level images and the building footprint in an orthorectified aerial image. The ground-level images are stitched into panoramas in order to obtain a wide camera field of view. Line segments are extracted from ground-level images. Their corresponding segments on the building footprints are automatically detected through a voting process. Meanwhile the camera pose of the ground-level images is also obtained. Wrong correspondences are corrected through user interaction. Later, the height values of the building roof corners are computed and a piece-wise planar 3D model with photo-realistic facade and roof texture is then created.
Lu Wang 0013, Suya You, Ulrich Neumann
ICCV2
2007 An Augmented Reality Interface for Mobile Information Retrieval
abstract
Recent years have seen growing interest in mobile augmented reality. The ability to retrieve information and display it as virtual content overlaid on top of an image of the real world is a natural extension to a mobile device equipped with a camera and wireless connectivity. Such applications need to address a number of technical hurdles including target recognition and camera pose estimation. Beyond these fundamental challenges, however, there is the problem of data connectivity and user interface presentation. How do we connect mobile clients to multiple data sources that may be changing in real time and provide the user with a flexible interface for navigating through the relevant content? We present an AR user interface framework specifically designed to expose disparate data sources through a single application server. Our proposed system uses an multi-tier architecture to separate back-end data retrieval from front-end graphical presentation and UI event handling. We describe how it might be used to build an oil platform equipment maintenance system, as an example of a collaborative, data-driven mobile application.
Jonathan Mooser, Lu Wang 0013, Suya You, Ulrich Neumann
ICME3
2007 Generating High-Resolution Textures for 3D Virtual Environments using View-Independent Texture Mapping
abstract
Image based modeling and rendering techniques have become increasingly popular for creating and visualizing 3D models from a set of images. Typically, these techniques depend on view-dependent texture mapping to render the textured 3D models in which the texture of novel views is synthesized at runtime according to different view-points. This is computationally expensive and limits their application in domains where efficient computations are required, such as games and virtual reality. In this paper we present an offline technique for creating view-independent texture atlases for 3D models, given a set of registered images. The best texture map resolution is computed by considering the areas of the projected polygons in the images. Texture maps are generated by a weighted composition of all available image information in the scene.Assuming that all surfaces of the model are exhibiting Lambertian reflectance properties, ray-tracing is then employed, for creating the view-independent texture maps. Finally, all the generated texture maps are packed into texture atlases. The result is a 3D model with an associated view-independent texture atlas which can be used efficiently in any application without any knowledge of camera pose information.
Charalambos Poullis, Suya You, Ulrich Neumann
ICME2
2007 A High-Performance Image Matching and Recognition System for Multimedia Applications
abstract
This paper presents a high-performance image matching and recognition system for rapid and robust detection, matching and recognition of scene imagery and objects in varied backgrounds. Advanced image processing and pattern recognition technologies provide the system with object distinctiveness, robustness to occlusions, and invariance to scale and geometric distortions. Extensive experiments and several commercial applications have demonstrated the system's value and utility for multimedia applications.
Suya You, Ulrich Neumann
ICME1
2007 Real-Time Object Tracking for Augmented Reality Combining Graph Cuts and Optical Flow
abstract
We present an efficient and accurate object tracking algorithm based on the concept of graph cut segmentation. The ability to track visible objects in real-time provides an invaluable tool for the implementation of markerless Augmented Reality. Once an object has been detected, it's location in future frames can be used to position virtual content, and thus annotate the environment. Unlike many object tracking algorithms, our approach does not rely on a preexisting 3D model or any other information about the object or its environment. It takes, as input, a set of pixels representing an object in an initial frame and uses a combination of optical flow and graph cut segmentation to determine the corresponding pixels in each future frame. Experiments show that our algorithm robustly tracks objects of disparate shapes and sizes over hundreds of frames, and can even handle difficult cases where an object contains many of the same colors as its background. We further show how this technology can be applied to practical AR applications.
Jonathan Mooser, Suya You, Ulrich Neumann
ISMAR2
2007 Single View Camera Calibration for Augmented Virtual Environments
abstract
Augmented virtual environments (AVE) are very effective in the application of surveillance, in which multiple video streams are projected onto a 3D urban model for better visualization and comprehension of the dynamic scenes. One of the key issues in creating such systems is to estimate the parameters of each camera including the intrinsic parameters and its pose relative to the 3D model. Existing camera pose estimation approaches require known intrinsic parameters and at least three 2D to 3D feature (point or line) correspondences. This cannot always be satisfied in an AVE system. Moreover, due to noise, the estimated camera location may be far from the expectation of the users when the number of correspondences is small. Our approach combines the users' prior knowledge about the camera location and the constraints from the parallel relationship between lines with those from feature correspondences. With at least two feature correspondences, it can always output an estimation of the camera parameters that gives an accurate alignment between the projection of the image (or video) and the 3D model
Lu Wang 0013, Suya You, Ulrich Neumann
VR2
2006 Tricodes: A Barcode-Like Fiducial Design for Augmented Reality Media
abstract
Visual markers, or fiducials, have become one of the most common methods of camera pose estimation in augmented reality (AR) media. Many present day fiducial-based AR systems use arbitrary patterns, such as simple line drawings or alpha-numeric characters, and require that an application be "trained" to recognize its pattern set. These techniques work well on a small scale, but as the number of fiducials grows, accuracy and performance degrade. We describe a new fiducial design called TriCodes that, like a barcode, provides a systematic way of printing and identifying a vast library of patterns. We compare TriCodes to the popular ARToolkit package, demonstrating its advantages in the presence of large numbers of fiducials
Jonathan Mooser, Suya You, Ulrich Neumann
ICME2
2006 Geodec: Enabling Geospatial Decision Making
abstract
The rapid increase in the availability of geospatial data has motivated the effort to seamlessly integrate this information into an information-rich and realistic 3D environment. However, heterogeneous data sources with varying degrees of consistency and accuracy pose a challenge to such efforts. We describe the geospatial decision making (GeoDec) system, which accurately integrates satellite imagery, three-dimensional models, textures and video streams, road data, maps, point data and temporal data. The system also includes a glove-based user interface
Cyrus Shahabi, Yao-Yi Chiang, Kelvin Chung, Kai-Chen Huang, Ali Khoshgozaran, Craig A. Knoblock, Sung Lee, Ulrich Neumann, Ramakant Nevatia, Arjun Rihan, Snehal Thakkar, Suya You
ICME12
2006 Fast Similarity Search for High-Dimensional Dataset
abstract
This paper addresses the challenging problem of rapidly searching and matching high-dimensional features for the applications of multimedia database retrieval and pattern recognition. Most current methods suffer from the problem of dimensionality curse. A number of theoretical and experimental studies lead us to pursue a new approach, called fast filtering vector approximation (FFVA) to tackle the problem. FFVA is a nearest neighbor search technique that facilitates rapidly indexing and recovering the most similar matches to a high-dimensional database of features or spatial data. Extensive experiments have demonstrated effectiveness of the proposed approach
Quan Wang 0001, Suya You
ISM2
2005 Rapid part-based 3D modeling
abstract
An intuitive and easy-to-use 3D modeling system has become more crucial with the rapid growth of computer graphics in our daily lives. Image-based modeling (IBM) has been a popular alternative to pure 3D modelers (e.g. 3D Studio Max, Maya) since its introduction in the late 1990s. However, IBM techniques are inherently very slow and rarely user friendly. Most IBM techniques require either very extensive manual input and/or multiple images. In this paper, we present an IBM technique that gives high level of detail with 1-2 minutes of manipulation from a novice user using only single, un-calibrated image. Our system modifies a generic part-based model of the object under investigation. User inputs are entered via a simple interface and converted into modifications to the whole 3D model. We demonstrate the effectiveness of our modeler by modeling several vehicles, such as SUVs, sedan/hatchback/coupe cars, minivans, trucks and more.
Ismail Oner Sebe, Suya You, Ulrich Neumann
VRST2
2004 A Robust Hybrid Tracking System for Outdoor Augmented Reality
Bolan Jiang, Ulrich Neumann, Suya You
VR3
2004 Colorplate: A Robust Hybrid Tracking System for Outdoor Augmented Reality
Bolan Jiang, Ulrich Neumann, Suya You
VR3
2003 Urban Site Modeling from LiDAR
Suya You, Ulrich Neumann, Pamela Fox
ICCSA (3)1
2003 Augmented Virtual Environments (AVE): Dynamic Fusion of Imagery and 3D Models
abstract
An augmented virtual environment (AVE) fuses dynamic imagery with 3D models. The AVE provides a unique approach to visualize and comprehend multiple streams of temporal data or images. Models are used as a 3D substrate for the visualization of temporal imagery, providing improved comprehension of scene activities. The core elements of AVE systems include model construction, sensor tracking, real-time video/image acquisition, and dynamic texture projection for 3D visualization. This paper focuses on the integration of these components and the results that illustrate the utility and benefits of the resulting augmented virtual environment.
Ulrich Neumann, Suya You, Bolan Jiang, Jong Weon Lee 0001
VR2
2002 Tracking with Omni-Directional Vision for Outdoor AR Systems
abstract
Most pose (3D position and 3D orientation) tracking methods using vision require a priori knowledge about the environment and correspondences between 3D environment features and 2D images. This environmental information is difficult to acquire accurately for large working volumes or may not be available at all, especially for outdoor environments. As a result, most pose tracking methods using vision are designed for small indoor working spaces. We track the pose of a moving camera from 2D images of the world. The pose of a camera is tracked through two 5 degree-of-freedom (DOF) motion estimations, which requires only 2D-to-2D correspondences. Therefore, the presented method can be applied to varied working space sizes including outdoor environments.
Jong Weon Lee 0001, Suya You, Ulrich Neumann
ISMAR2
2001 Fusion of Vision and Gyro Tracking for Robust Augmented Reality Registration
abstract
A novel framework enables accurate augmented reality (AR) registration with integrated inertial gyroscope and vision tracking technologies. The framework includes a two-channel complementary motion filter that combines the low-frequency stability of vision sensors with the high-frequency tracking of gyroscope sensors, hence achieving stable static and dynamic six-degree-of-freedom pose tracking. Our implementation uses an extended Kalman filter (EKF). Quantitative analysis and experimental results show that the fusion method achieves dramatic improvements in tracking stability and robustness over either sensor alone. We also demonstrate a new fiducial design and detection system in our example AR annotation systems that illustrate the behavior and benefits of the new tracking method.
Suya You, Ulrich Neumann
VR1
2000 A video-based augmented reality golf simulator
Alok Govil, Suya You, Ulrich Neumann
ACM Multimedia2
1999 Hybrid Inertial and Vision Tracking for Augmented Reality Registration
abstract
The biggest single obstacle to building effective augmented reality (AR) systems is the lack of accurate wide-area sensors for trackers that report the locations and orientations of objects in an environment. Active (sensor-emitter) tracking technologies require powered-device installation. Limiting their use to prepared areas that are relatively free of natural or man-made interference sources. Vision-based systems can use passive landmarks, but they are more computationally demanding and often exhibit erroneous behavior due to occlusion or numerical instability. Inertial sensors are completely passive, requiring no external devices or targets, however, the drift rates in portable strapdown configurations are too great for practical use. In this paper, we present a hybrid approach to AR tracking that integrates inertial and vision-based technologies. We exploit the complementary nature of the two technologies to compensate for the weaknesses in each component. Analysis and experimental results demonstrate this system's effectiveness.
Suya You, Ulrich Neumann, Ronald T. Azuma
VR1
1999 Tracking in unprepared environments for augmented reality systems
Ronald T. Azuma, Jong Weon Lee 0001, Bolan Jiang, Jun Park, Suya You, Ulrich Neumann
Comput. Graph.5
1999 Natural Feature Tracking for Augmented Reality
abstract
Natural scene features stabilize and extend the tracking range of augmented reality (AR) pose-tracking systems. We develop robust computer vision methods to detect and track natural features in video images. Point and region features are automatically and adaptively selected for properties that lead to robust tracking. A multistage tracking algorithm produces accurate motion estimates, and the entire system operates in a closed-loop that stabilizes its performance and accuracy. We present demonstrations of the benefits of using tracked natural features for AR applications that illustrate direct scene annotation, pose stabilization, and extendible tracking range. Our system represents a step toward integrating vision with graphics to produce robust wide-area augmented realities.
Ulrich Neumann, Suya You
IEEE Trans. Multim.2
1998 Integration of Region Tracking and Optical Flow for Image Motion Estimation
Ulrich Neumann, Suya You
ICIP (3)2
1997 Interactive volume rendering for virtual colonoscopy
abstract
3D virtual colonoscopy has recently been proposed as a non-invasive alternative procedure for the visualization of the human colon. Surface rendering is sufficient for implementing such a procedure to obtain an overview of the interior surface of the colon at interactive rendering speeds. Unfortunately, physicians can not use it to explore tissues beneath the surface to differentiate between benign and malignant structures. In this paper, we present a direct volume rendering approach based on perspective ray casting, as a supplement to the surface navigation. To accelerate the rendering speed, surface-assistant techniques are used to adapt the resampling rates by skipping the empty space inside the colon. In addition, a parallel version of the algorithm has been implemented on a shared-memory multiprocessing architecture. Experiments have been conducted on both simulation and patient data sets.
Suya You, Lichan Hong, Kittiboon Junyaprasert, Arie E. Kaufman, Shigeru Muraki, Yong Zhou 0001, Mark Wax, Zhengrong Liang
IEEE Visualization1
1997 A multi-view face recognition system
Yongyue Zhang, Zhenyun Peng, Suya You, Guangyou Xu
J. Comput. Sci. Technol.3