EDBT 2026 Demo / reviewers in the wild / expert
Sai-Kit Yeung
dblp:144/7479 · also Sai Kit Yeung
· DBLP profile ↗
100ranked-venue papers
6as first author
42since 2021 · last 2026
0000-0001-7974-0607ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 72 · 5 first-author · 28 since 2021Artificial intelligence and machine learning · 63 · 4 first-author · 29 since 2021Systems, architecture and hardware · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent CuesabstractGlass is a prevalent material among solid objects in every-day life, yet segmentation methods struggle to distinguish it from opaque materials due to its transparency and reflection. While it is known that human perception relies on boundary and reflective-object features to distinguish glass objects, the existing literature has not yet sufficiently captured both properties when handling transparent objects. Hence, we propose incorporating both of these powerful visual cues via the Boundary Feature Enhancement and Reflection Feature Enhancement modules in a mutually beneficial way. Our proposed framework, TransCues, is a pyramidal transformer encoder-decoder architecture to segment transparent objects. We empirically show that these two modules can be used together effectively, improving overall performance across various benchmark datasets, including glass object semantic segmentation, mirror object semantic segmentation, and generic segmentation datasets. Our method outperforms the state-of-the-art by a large margin, achieving +4.2% mIoU on Trans10K-v2, +5.6% mIoU on MSD, +10.1% mIoU on RGBD-Mirror, +13.1% mIoU on TROSD, and +8.3% mIoU on Stanford2D3D, showing the effectiveness of our method against glass objects. Tuan-Anh Vu, Hai Nguyen-Truong, Ziqiang Zheng, Binh-Son Hua, Qing Guo 0005, Ivor W. Tsang, Sai-Kit Yeung |
WACV | 7 |
| 2026 | ORCA: Object Recognition and Comprehension for Archiving Marine SpeciesabstractMarine visual understanding is essential for monitoring and protecting marine ecosystems, enabling automatic and scalable biological surveys. However, progress is hindered by limited training data and the lack of a systematic task formulation that aligns domain-specific marine challenges with well-defined computer vision tasks, thereby limiting effective model application. To address this gap, we present ORCA, a multi-modal benchmark for marine research comprising 14,647 images from 478 species, with 42,217 bounding box annotations and 22,321 expert-verified instance captions. The dataset provides fine-grained visual and textual annotations that capture morphology-oriented attributes across diverse marine species. To catalyze methodological advances, we evaluate 18 state-of-the-art models on three tasks: object detection (closed-set and open-vocabulary), instance captioning, and visual grounding. Results highlight key challenges, including species diversity, morphological overlap, and specialized domain demands, underscoring the difficulty of marine understanding. ORCA thus establishes a comprehensive benchmark to advance research in marine domain. Yuk-Kwan Wong, Haixin Liang, Zeyu Ma 0002, Yiwei Chen 0003, Ziqiang Zheng, Rinaldi Gotama, Pascal Sebastian, Lauren D. Sparks, Sai-Kit Yeung |
WACV | 9 |
| 2026 | MarineEval: Assessing the Marine Intelligence of Vision-Language ModelsabstractWe have witnessed promising progress led by large language models (LLMs) and further vision language models (VLMs) in handling various queries as a general-purpose assistant. VLMs, as a bridge to connect the visual world and language corpus, receive both visual content and various text-only user instructions to generate corresponding responses. Though great success has been achieved by VLMs in various fields, in this work, we ask whether the existing VLMs can act as domain experts, accurately answering marine questions, which require significant domain expertise and address special domain challenges/requirements. To comprehensively evaluate the effectiveness and explore the boundary of existing VLMs, we construct the first large-scale marine VLM dataset and benchmark called MarineEval, with 2,000 image-based question-answering pairs. During our dataset construction, we ensure the diversity and coverage of the constructed data: 7 task dimensions and 20 capacity dimensions. The domain requirements are specially integrated into the data construction and further verified by the corresponding marine domain experts. We comprehensively benchmark 17 existing VLMs on our MarineEval and also investigate the limitations of existing models in answering marine research questions. The experimental results reveal that existing VLMs cannot effectively answer the domain-specific questions, and there is still a large room for further performance improvements. We hope our new benchmark and observations will facilitate future research. Yuk-Kwan Wong, Tuan-An To, Ziqiang Zheng, Sai-Kit Yeung |
WACV | 5 |
| 2026 | Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with DiffusionabstractAbstract Text-to-image diffusion techniques have shown exceptional capabilities in producing high-quality, dense visual predictions from open-vocabulary text. This indicates a strong correlation between visual and textual domains in open concepts and that diffusion-based text-to-image models can capture rich and diverse information for computer vision tasks. However, we found that those advantages do not hold for learning of features of camouflaged individuals because of the significant blending between their visual boundaries and their surroundings. In this paper, while leveraging the benefits of diffusion-based techniques and text-image models in open-vocabulary settings, we aim to address a challenging problem in computer vision: open-vocabulary camouflaged instance segmentation (OVCIS). Specifically, we propose a method built upon state-of-the-art diffusion empowered by open-vocabulary to learn multi-scale textual-visual features for camouflaged object representation learning. Such cross-domain representations are desirable in segmenting camouflaged objects where visual cues subtly distinguish the objects from the background, and in segmenting novel object classes which are not seen in training. To enable such powerful representations, we devise complementary modules to effectively fuse cross-domain features, and to engage relevant features towards respective foreground objects. We validate and compare our method with existing ones on several benchmark datasets of camouflaged and generic open-vocabulary instance segmentation. The experimental results confirm the advances of our method over existing ones. We believe that our proposed method would open a new avenue for handling camouflages such as computer vision-based surveillance systems, wildlife monitoring, and military reconnaissance. Tuan-Anh Vu, Duc Thanh Nguyen, Qing Guo 0005, Nhat Chung, Binh-Son Hua, Ivor W. Tsang, Sai-Kit Yeung |
Int. J. Comput. Vis. | 7 |
| 2026 | CamoVid60K: A Large-Scale Video Dataset for Moving Camouflaged Animals UnderstandingabstractAbstract We have been witnessing remarkable success led by the power of neural networks driven by a significant scale of training data in handling various computer vision tasks. However, less attention has been paid to monitoring the camouflaged animals, the masters of hiding themselves in the background. Robust and precise segmentation of camouflaged animals is challenging even for domain experts due to their similarity to the environment. Although several efforts have been made in camouflaged animal image segmentation, to the best of our knowledge, limited work exists on camouflaged animal video understanding (CAVU). Biologists often prefer videos for monitoring and understanding animal behaviors, as videos provide redundant information and temporal consistency. However, the scarcity of labeled video data significantly hinders progress in this area. To address these challenges, we present CamoVid60K , a diverse, large-scale, and accurately annotated video dataset of camouflaged animals. This dataset comprises 218 videos with 62,774 finely annotated frames, covering 70 animal categories, which surpasses all previous datasets in terms of the number of videos/frames and species included. CamoVid60K also offers more diverse downstream tasks in computer vision, such as camouflaged animal classification, detection, and task-specific segmentation (semantic, referring, motion), etc. We have benchmarked several state-of-the-art algorithms on the proposed CamoVid60K dataset, and the experimental results provide valuable insights for future research directions. Our dataset serves as a novel and challenging benchmark to stimulate the development of more powerful camouflaged animal video segmentation algorithms, with substantial room for further improvement. Tuan-Anh Vu, Ziqiang Zheng, Chengyang Song, Qing Guo 0005, Ivor W. Tsang, Sai-Kit Yeung |
Int. J. Comput. Vis. | 6 |
| 2025 | Align3R: Aligned Monocular Depth Estimation for Dynamic VideosabstractRecent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Very recent works address this problem by applying a video diffusion model to generate video depth conditioned on the input video, which is training-expensive and can only produce scale-invariant depth values without camera poses. In this paper, we propose a novel video-depth estimation method called Align3R to estimate temporally consistent depth maps for a dynamic video. Our key idea is to utilize the recent DUSt3R model to align estimated monocular depth maps of different timesteps. First, we fine-tune the DUSt3R model with additional estimated monocular depth as inputs for the dynamic scenes. Then, we apply optimization to reconstruct both depth maps and camera poses. Extensive experiments demonstrate that Align3R estimates consistent video depth and camera poses for a monocular video with superior performance than baseline methods. Jiahao Lu 0001, Zhiyang Dou, Cheng Lin 0001, Zhiming Cui 0001, Zhen Dong 0005, Sai-Kit Yeung, Wenping Wang 0001, Yuan Liu 0025 |
CVPR | 8 |
| 2025 | Color Alignment in DiffusionabstractDiffusion models have shown great promise in synthesizing visually appealing images. However, it remains challenging to condition the synthesis at a fine-grained level, for instance, synthesizing image pixels following some generic color pattern. Existing image synthesis methods often produce contents that fall outside the desired pixel conditions. To address this, we introduce a novel color alignment algorithm that confines the generative process in diffusion models within a given color pattern. Specifically, we project diffusion terms, either imagery samples or latent representations, into a conditional color space to align with the input color distribution. This strategy simplifies the prediction in diffusion models within a color manifold while still allowing plausible structures in generated contents, thus enabling the generation of diverse contents that comply with the target color pattern. Experimental results demonstrate our state-of-the-art performance in conditioning and controlling of color pixels, while maintaining on-par generation quality and diversity in comparison with regular diffusion models. Ka-Chun Shum, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung |
CVPR | 4 |
| 2025 | CoraLSRT: Revisiting Coral Reef Semantic Segmentation by Feature Rectification via Self-Supervised Guidance
Ziqiang Zheng, Yuk-Kwan Wong, Binh-Son Hua, Jianbo Shi, Sai-Kit Yeung |
ICCV | 5 |
| 2025 | SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splattingabstract360-degree cameras streamline data collection for radiance field 3D reconstruction by capturing comprehensive scene data. However, traditional radiance field methods do not address the specific challenges inherent to 360-degree images. We present SC-OmniGS, a novel self-calibrating omnidirectional Gaussian splatting system for fast and accurate omnidirectional radiance field reconstruction using 360-degree images. Rather than converting 360-degree images to cube maps and performing perspective image calibration, we treat 360-degree images as a whole sphere and derive a mathematical framework that enables direct omnidirectional camera pose calibration accompanied by 3D Gaussians optimization. Furthermore, we introduce a differentiable omnidirectional camera model in order to rectify the distortion of real-world data for performance enhancement. Overall, the omnidirectional camera intrinsic model, extrinsic poses, and 3D Gaussians are jointly optimized by minimizing weighted spherical photometric loss. Extensive experiments have demonstrated that our proposed SC-OmniGS is able to recover a high-quality radiance field from noisy camera poses or even no pose prior in challenging scenarios characterized by wide baselines and non-object-centric configurations. The noticeable performance gain in the real-world dataset captured by consumer-grade omnidirectional cameras verifies the effectiveness of our general omnidirectional camera model in reducing the distortion of 360-degree images. Huajian Huang, Yingshu Chen, Tristan Braud, Sai-Kit Yeung |
ICLR | 7 |
| 2025 | MSC: A Marine Wildlife Dataset for Video Understanding with Grounded Segmentation and Clip-Level CaptionsabstractMarine videos present significant challenges for video understanding due to the dynamics of marine objects and the surrounding environment, camera motion, and the complexity of underwater scenes. Existing video captioning datasets, typically focused on generic or human-centric domains, often fail to generalize to the complexities of the marine environment and gain insights about marine life. To address these limitations, we propose a two-stage marine object-oriented video captioning pipeline. We introduce a comprehensive video understanding benchmark that leverages the triplets of video, text, and segmentation masks to facilitate visual grounding and captioning, leading to improved marine video understanding and analysis, and marine video generation. Additionally, we highlight the effectiveness of video splitting in order to detect salient object transitions in scene changes, which significantly enrich the semantics of captioning content. Our dataset and code have been released at https://msc.hkustvgd.com. Quang-Trung Truong, Yuk-Kwan Wong, Vo Hoang Kim Tuyen Dang, Rinaldi Gotama, Duc Thanh Nguyen, Sai-Kit Yeung |
ACM Multimedia | 6 |
| 2025 | TrackingWorld: World-centric Monocular 3D Tracking of Almost All PixelsabstractMonocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall short in separating the camera motion from foreground dynamic motion and cannot densely track newly emerging dynamic subjects in the videos. To address these two limitations, we propose TrackingWorld, a novel pipeline for dense 3D tracking of almost all pixels within a world-centric 3D coordinate system. First, we introduce a tracking upsampler that efficiently lifts the arbitrary sparse 2D tracks into dense 2D tracks. Then, to generalize the current tracking methods to newly emerging objects, we apply the upsampler to all frames and reduce the redundancy of 2D tracks by eliminating the tracks in overlapped regions. Finally, we present an efficient optimization-based framework to back-project dense 2D tracks into world-centric 3D trajectories by estimating the camera poses and the 3D coordinates of these 2D tracks. Extensive evaluations on both synthetic and real-world datasets demonstrate that our system achieves accurate and dense 3D tracking in a world-centric coordinate frame. Jiahao Lu 0001, Weitao Xiong, Jiacheng Deng 0002, Zhiyang Dou, Cheng Lin 0001, Sai-Kit Yeung, Yuan Liu 0025 |
NeurIPS | 8 |
| 2025 | OmniGS: Fast Radiance Field Reconstruction Using Omnidirectional Gaussian SplattingabstractPhotorealistic reconstruction relying on 3D Gaussian Splatting has shown promising potential in various domains. However, the current 3D Gaussian Splatting system only supports radiance field reconstruction using undistorted perspective images. In this paper, we present OmniGS, a novel omnidirectional Gaussian splatting system, to take advantage of omnidirectional images for fast radiance field reconstruction. Specifically, we conduct a theoretical analysis of spherical camera model derivatives in 3D Gaussian Splatting. According to the derivatives, we then implement a new GPU-accelerated omnidirectional rasterizer that directly splats 3D Gaussians onto the equirect-angular screen space for omnidirectional image rendering. We realize differentiable optimization of the omnidirectional radiance field without the requirement of cube-map rectification or tangent-plane approximation. Extensive experiments conducted in egocentric and roaming scenarios demonstrate that our method achieves state-of-the-art reconstruction quality and high rendering speed using omni-directional images. The code will be publicly available at https://github.com/liquorleaf/OmniGS. Huajian Huang, Sai-Kit Yeung |
WACV | 3 |
| 2025 | Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context UnderstandingabstractReferring image segmentation is a challenging task that involves generating pixel-wise segmentation masks based on natural language descriptions. The complexity of this task increases with the intricacy of the sentences provided. Existing methods have relied mostly on visual features to generate the segmentation masks while treating text features as supporting components. However, this under-utilization of text understanding limits the model's capability to fully comprehend the given expressions. In this work, we propose a novel framework that specifically emphasizes object and context comprehension inspired by human cognitive processes through Vision-Aware Text Features. Firstly, we introduce a CLIP Prior module to localize the main object of interest and embed the object heatmap into the query initialization process. Secondly, we propose a combination of two components: Contextual Multimodal Decoder and Meaning Consistency Constraint, to further enhance the coherent and consistent interpretation of language cues with the contextual understanding obtained from the image. Our method achieves significant performance improvements on three benchmark datasets RefCOCO, RefCOCO+ and G-Ref Project page: https://vatex.hkustvgd.com/ Hai Nguyen-Truong, E-Ro Nguyen, Tuan-Anh Vu, Minh-Triet Tran, Binh-Son Hua, Sai-Kit Yeung |
WACV | 6 |
| 2025 | Advances in 3D Neural Stylization: A SurveyabstractAbstract Modern artificial intelligence offers a novel and transformative approach to creating digital art across diverse styles and modalities like images, videos and 3D data, unleashing the power of creativity and revolutionizing the way that we perceive and interact with visual content. This paper reports on recent advances in stylized 3D asset creation and manipulation with the expressive power of neural networks. We establish a taxonomy for neural stylization, considering crucial design choices such as scene representation, guidance data, optimization strategies, and output styles. Building on such taxonomy, our survey first revisits the background of neural stylization on 2D images, and then presents in-depth discussions on recent neural stylization methods for 3D data, accompanied by a benchmark evaluating selected mesh and neural field stylization methods. Based on the insights gained from the survey, we highlight the practical significance, open challenges, future research, and potential impacts of neural stylization, which facilitates researchers and practitioners to navigate the rapidly evolving landscape of 3D content creation using modern artificial intelligence. Yingshu Chen, Guocheng Shao, Ka-Chun Shum, Binh-Son Hua, Sai-Kit Yeung |
Int. J. Comput. Vis. | 5 |
| 2025 | 360VOTS: Visual Object Tracking and Segmentation in Omnidirectional VideosabstractVisual object tracking and segmentation in omnidirectional videos are challenging due to the wide field-of-view and large spherical distortion brought by 360$^{\circ }$∘ images. To alleviate these problems, we introduce a novel representation, extended bounding field-of-view (eBFoV), for target localization and use it as the foundation of a general 360 tracking framework which is applicable for both omnidirectional visual object tracking and segmentation tasks. Building upon our previous work on omnidirectional visual object tracking (360VOT), we propose a comprehensive dataset and benchmark that incorporates a new component called omnidirectional video object segmentation (360VOS). The 360VOS dataset includes 290 sequences accompanied by dense pixel-wise masks and covers a broader range of target categories. To support both the development and evaluation of algorithms in this domain, we divide the dataset into a training subset with 170 sequences and a testing subset with 120 sequences. Furthermore, we tailor evaluation metrics for both omnidirectional tracking and segmentation to ensure rigorous assessment. Through extensive experiments, we benchmark state-of-the-art approaches and demonstrate the effectiveness of our proposed 360 tracking framework and training dataset. Yinzhe Xu, Huajian Huang, Yingshu Chen, Sai-Kit Yeung |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Test-Time Augmentation for 3D Point Cloud Classification and SegmentationabstractData augmentation is a powerful technique to enhance the performance of a deep learning task but has received less attention in 3D deep learning. It is well known that when 3D shapes are sparsely represented with low point density, the performance of the downstream tasks drops significantly. This work explores test-time augmentation (TTA) for 3D point clouds. We are inspired by the recent revolution of learning implicit representation and point cloud upsampling, which can produce high-quality 3D surface reconstruction and proximity-to-surface, respectively. Our idea is to leverage the implicit field reconstruction or point cloud upsampling techniques as a systematic way to augment point cloud data. Mainly, we test both strategies by sampling points from the reconstructed results and using the sampled point cloud as test-time augmented data. We show that both strategies are effective in improving accuracy. We observed that point cloud upsampling for test-time augmentation can lead to more significant performance improvement on downstream tasks such as object classification and segmentation on the ModelNet40, ShapeNet, ScanObjectNN, and SemanticKITTI datasets, especially for sparse point clouds. Tuan-Anh Vu, Srinjay Sarkar, Zhiyuan Zhang 0004, Binh-Son Hua, Sai-Kit Yeung |
3DV | 5 |
| 2024 | 360Loc: A Dataset and Benchmark for Omnidirectional Visual Localization with Cross-Device QueriesabstractPortable 360° cameras are becoming a cheap and efficient tool to establish large visual databases. By capturing omnidirectional views of a scene, these cameras could expedite building environment models that are essential for visual localization. However, such an advantage is often overlooked due to the lack of valuable datasets. This paper introduces a new benchmark dataset, 360Loc, composed of 360° images with ground truth poses for visual localization. We present a practical implementation of 360° mapping combining 360° images with lidar data to generate the ground truth 6DoF poses. 360Loc is the first dataset and benchmark that explores the challenge of cross-device visual positioning, involving 360° reference frames, and query frames from pinhole, ultra-wide FoV fisheye, and 360° cameras. We propose a virtual camera approach to generate lower-FoV query frames from 360° images, which ensures a fair comparison of performance among different query types in visual localization tasks. We also extend this virtual camera approach to feature matching-based and pose regression-based methods to alleviate the performance loss caused by the cross-device domain gap, and evaluate its effectiveness against state-of-the-art base-lines. We demonstrate that omnidirectional visual localization is more robust in challenging large-scale scenes with symmetries and repetitive structures. These results provide new insights into 360-camera mapping and omnidirectional visual localization with cross-device queries. Project Page and dataset: https://huajianup.github.io/research/360Loc/ Huajian Huang, Changkun Liu 0001, Yipeng Zhu, Tristan Braud, Sai-Kit Yeung |
CVPR | 6 |
| 2024 | Photo-SLAM: Real-Time Simultaneous Localization and Photorealistic Mapping for Monocular, Stereo, and RGB-D CamerasabstractThe integration of neural rendering and the SLAM system recently showed promising results in joint localization and photorealistic view reconstruction. However, existing methods, fully relying on implicit representations, are so resource-hungry that they cannot run on portable devices, which deviates from the original intention of SLAM. In this paper, we present Photo-SLAM, a novel SLAM framework with a hyper primitives map. Specifically, we simultaneously exploit explicit geometric features for localization and learn implicit photometric features to represent the texture information of the observed environment. In addition to actively densifying hyper primitives based on geometric features, we further introduce a Gaussian-Pyramid-based training method to progressively learn multi-level features, enhancing photorealistic mapping performance. The extensive experiments with monocular, stereo, and RGB-D datasets prove that our proposed system Photo-SLAM sig-nificantly outperforms current state-of-the-art SLAM systems for online photorealistic mapping, e.g., PSNR is 30% higher and rendering speed is hundreds of times faster in the Replica dataset. Moreover, the Photo-SLAM can run at real-time speed using an embedded platform such as Jet-son AGX Orin, showing the potential of robotics applications. Project Page and code: https://huajianup.github.io/research/Photo-SLAM/. Huajian Huang, Sai-Kit Yeung |
CVPR | 4 |
| 2024 | Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset UpdatesabstractNeural radiance field (NeRF) is an emerging technique for 3D scene reconstruction and modeling. However, current NeRF-based methods are limited in the capabilities of adding or removing objects. This paper fills the aforementioned gap by proposing a new language-driven method for object manipulation in NeRFs through dataset updates. Specifically, to insert an object represented by a set of multi-view images into a background NeRF, we use a text-to-image diffusion model to blend the object into the given background across views. The generated images are then used to update the NeRF so that we can render view-consistent images of the object within the background. To ensure view consistency, we propose a dataset update strategy that prioritizes the radiance field training based on camera poses in a pose-ordered manner. We validate our method in two case studies: object insertion and object removal. Experimental results show that our method can generate photo-realistic results and achieves state-of-the-art performance in NeRF editing. Ka-Chun Shum, Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung |
CVPR | 5 |
| 2024 | CoralSCOP: Segment any COral Image on this PlanetabstractUnderwater visual understanding has recently gained increasing attention within the computer vision community for studying and monitoring underwater ecosystems. Among these, coral reefs play an important and intricate role, often referred to as the rainforests of the sea, due to their rich bio-diversity and crucial environmental impact. Existing coral analysis, due to its technical complexity, requires significant manual work from coral biologists, therefore hindering scalable and comprehensive studies. In this paper, we introduce CoralSCop, the first foundation model designed for the automatic dense segmentation of coral reefs. CoralSCOP is developed to accurately assign labels to different coral entities, addressing the challenges in the semantic analysis of coral imagery. Its main objective is to identify and delineate the irregular boundaries between various coral individuals across different granularities, such as coral/non-coral, growth form, and genus. This task is challenging due to the semantic agnostic nature or fixed limited semantic categories of previous generic segmentation methods, which fail to adequately capture the complex characteristics of coral structures. By introducing a novel parallel semantic branch, CoralSCOP can produce high-quality coral masks with semantics that enable a wide range of downstream coral reef analysis tasks. We demonstrate that CoralSCOP exhibits a strong zero-shot ability to segment unseen coral images. To effectively train our foundation model, we propose CoralMask, a new dataset with 41,297 densely labeled coral images and 330,144 coral masks. We have conducted comprehensive and extensive experiments to demonstrate the advantages of CoralSCOP over existing generalist segmentation algorithms and coral reef analytical approaches. Ziqiang Zheng, Haixin Liang, Binh-Son Hua, Yue Him Wong, Put Ang, Apple Pui Yi Chui, Sai-Kit Yeung |
CVPR | 7 |
| 2024 | StyleCity: Large-Scale 3D Urban Scenes Stylization
Yingshu Chen, Huajian Huang, Tuan-Anh Vu, Ka-Chun Shum, Sai-Kit Yeung |
ECCV (59) | 5 |
| 2024 | MarineInst: A Foundation Model for Marine Image Analysis with Instance Visual Description
Ziqiang Zheng, Yiwei Chen 0003, Tuan-Anh Vu, Binh-Son Hua, Sai-Kit Yeung |
ECCV (2) | 6 |
| 2023 | 360VOT: A New Benchmark Dataset for Omnidirectional Visual Object Trackingabstract360° images can provide an omnidirectional field of view which is important for stable and long-term scene perception. In this paper, we explore 360° images for visual object tracking and perceive new challenges caused by large distortion, stitching artifacts, and other unique attributes of 360° images. To alleviate these problems, we take advantage of novel representations of target localization, i.e., bounding field-of-view, and then introduce a general 360 tracking framework that can adopt typical trackers for omnidirectional tracking. More importantly, we propose a new large-scale omnidirectional tracking benchmark dataset, 360VOT, in order to facilitate future research. 360VOT contains 120 sequences with up to 113K high-resolution frames in equirectangular projection. The tracking targets cover 32 categories in diverse scenarios. Moreover, we provide 4 types of unbiased ground truth, including (rotated) bounding boxes and (rotated) bounding field-of-views, as well as new metrics tailored for 360° images which allow for the accurate evaluation of omnidirectional tracking performance. Finally, we extensively evaluated 20 state-of-the-art visual trackers and provided a new baseline for future comparisons. Homepage: https://360vot.hkustvgd.com Huajian Huang, Yinzhe Xu, Yingshu Chen, Sai-Kit Yeung |
ICCV | 4 |
| 2023 | Locally Stylized Neural Radiance FieldsabstractIn recent years, there has been increasing interest in applying stylization on 3D scenes from a reference style image, in particular onto neural radiance fields (NeRF). While performing stylization directly on NeRF guarantees appearance consistency over arbitrary novel views, it is a challenging problem to guide the transfer of patterns from the style image onto different parts of the NeRF scene. In this work, we propose a stylization framework for NeRF based on local style transfer. In particular, we use a hash-grid encoding to learn the embedding of the appearance and geometry components, and show that the mapping defined by the hash table allows us to control the stylization to a certain extent. Stylization is then achieved by optimizing the appearance branch while keeping the geometry branch fixed. To support local style transfer, we propose a new loss function that utilizes a segmentation network and bipartite matching to establish region correspondences between the style image and the content images obtained from volume rendering. Our experiments show that our method yields plausible stylization results with novel view synthesis while having flexible controllability via manipulating and customizing the region correspondences. Hong-Wing Pang, Binh-Son Hua, Sai-Kit Yeung |
ICCV | 3 |
| 2023 | Conditional 360-degree Image Synthesis for Immersive Indoor Scene DecorationabstractIn this paper, we address the problem of conditional scene decoration for 360° images. Our method takes a 360° background photograph of an indoor scene and generates decorated images of the same scene in the panorama view. To do this, we develop a 360-aware object layout generator that learns latent object vectors in the 360° view to enable a variety of furniture arrangements for an input 360° background image. We use this object layout to condition a generative adversarial network to synthesize images of an input scene. To further reinforce the generation capability of our model, we develop a simple yet effective scene emptier that removes the generated furniture and produces an emptied scene for our model to learn a cyclic constraint. We train the model on the Structure3D dataset and show that our model can generate diverse decorations with controllable object layout. Our method achieves state-of-the-art performance on the Structure3D dataset and generalizes well to the Zillow indoor scene dataset. Our user study confirms the immersive experiences provided by the realistic image quality and furniture layout in our generation results. Our implementation is available at https://github.com/kcshum/neural_360_decoration.git. Ka-Chun Shum, Hong-Wing Pang, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung |
ICCV | 5 |
| 2023 | A Novel Platform to Control Biofouling in Pearl Oysters CultivationabstractThis paper presents a simple yet effective design of a platform to automate the task of shellfish aquaculture, specifically pearl oysters. Compared to traditional methods, our platform can eliminate the tedious task of cleaning the pearl oysters due to fouling. Inspired by the low and high tide characteristics of the intertidal zone, our platform employs an air-water displacement mechanism to periodically float pearl oysters above the water's surface, exposing fouling organisms to air and sunlight. While pearl oysters have developed the ability to stay alive during low tide, these fouling organisms cannot survive after prolonged exposure, thus preventing them from developing. Additionally, the platform provides an alternative approach to grow not only pearl oysters but also various types of shellfish, consequently benefiting the aquaculture industry. We introduce the design of the platform and provide a comprehensive analysis. We also demonstrate the practical deployment of the platform for cultivating pearl oysters. Van-Nhan Tran, Quan-Dung Pham, Tan-Sang Ha, Yue Him Wong, Sai-Kit Yeung |
ICRA | 5 |
| 2023 | Cross-Domain Autonomous Driving Perception Using Contrastive Appearance AdaptationabstractAddressing domain shifts for complex perception tasks in autonomous driving has long been a challenging problem. In this paper, we show that existing domain adaptation methods pay little attention to the content mismatch issue between source and target domains, thus weakening the domain adaptation per-formance and the decoupling of domain-invariant and domain-specific representations. To solve the aforementioned problems, we propose an image-level domain adaptation framework that aims at adapting source-domain images to the target domain with content-aligned source-target image pairs. Our framework consists of three mutually beneficial modules in a cycle: a cross-domain content alignment module to generate source-target pairs with consistent content representations in a self-supervised manner, a reference-guided image synthesis based on the generated content-aligned source-target image pairs, and a contrastive learning module to self-supervise domain-invariant feature extractor. Our contrastive appearance adaptation is task-agnostic and robust to complex perception tasks in autonomous driving. Our proposed method demonstrates state-of-the-art results in cross-domain object detection, semantic segmentation, and depth estimation as well as better image synthesis ability qualitatively and quantitatively. Ziqiang Zheng, Yingshu Chen, Binh-Son Hua, Yang Wu 0001, Sai-Kit Yeung |
IROS | 5 |
| 2023 | CompUDA: Compositional Unsupervised Domain Adaptation for Semantic Segmentation Under Adverse ConditionsabstractIn autonomous driving, performing robust semantic segmentation under adverse weather conditions is a long-standing challenge. Imperfect camera observations under adverse conditions result in images with reduced visibility, which hinders label annotation and semantic scene understanding based on these images. A common solution is to adopt semantic segmentation models trained in a source domain with ground truth labels and perform unsupervised domain adaptation (UDA) from the source domain to an unlabeled target domain that has adverse conditions. Due to imperfect visual observations in the target domain, such adaptation needs special treatment to achieve good performance. In this paper, we propose a new compositional unsupervised domain adaptation (CompUDA) method that disentangles the domain gap based on multiple factors including style, visibility, and image quality. The domain gaps caused by these individual factors can then be addressed separately by introducing the intermediate domains. Specifically, 1) to address the style gap, we perform source-to-intermediate domain adaptation and generate pseudo-labels for self-training in the target domain; 2) to address the visibility gap, we perform a geometry-aligned normal-to-adverse image translation and introduce a synthetic domain; 3) finally, to address the image quality gap between the synthetic and target domain, we perform a synthetic-to-real adaptation based on the generated pseudo-labels. Our compositional unsupervised domain adaptation can be used in conjunction with a wide variety of semantic segmentation methods and result in significant performance improvement across datasets. The codes are available at https://github.com/zhengziqiang/CompUDA. Ziqiang Zheng, Yingshu Chen, Binh-Son Hua, Sai-Kit Yeung |
IROS | 4 |
| 2023 | Marine Video Kit: A New Marine Video Dataset for Content-Based Analysis and Retrieval
Quang-Trung Truong, Tuan-Anh Vu, Tan-Sang Ha, Jakub Lokoc, Yue Him Wong, Ajay Joneja, Sai-Kit Yeung |
MMM (1) | 7 |
| 2023 | Authoring Next-Generation XR Storytelling Experiences using Real-World Scene Data - From Everyday Environments to the Oceanabstractcourse Share on Authoring Next-Generation XR Storytelling Experiences using Real-World Scene Data - From Everyday Environments to the Ocean Authors: Lap-Fai (Craig) Yu GMU GMUView Profile , Sai-Kit Yeung HKUST HKUSTView Profile Authors Info & Claims SA '23: SIGGRAPH Asia 2023 CoursesDecember 2023Article No.: 3Pages 1–39https://doi.org/10.1145/3610538.3614623Published:06 December 2023Publication History 0citation59DownloadsMetricsTotal Citations0Total Downloads59Last 12 Months59Last 6 weeks21 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Lap-Fai Yu, Sai-Kit Yeung |
SIGGRAPH ASIA Courses | 2 |
| 2023 | PointInverter: Point Cloud Reconstruction and Editing via a Generative Model with Shape PriorsabstractIn this paper, we propose a new method for mapping a 3D point cloud to the latent space of a 3D generative adversarial network. Our generative model for 3D point clouds is based on SP-GAN, a state-of-the-art sphere-guided 3D point cloud generator. We derive an efficient way to encode an input 3D point cloud to the latent space of the SP-GAN. Our point cloud encoder can resolve the point ordering issue during inversion, and thus can determine the correspondences between points in the generated 3D point cloud and those in the canonical sphere used by the generator. We show that our method outperforms previous GAN inversion methods for 3D point clouds, achieving state-of-the-art results both quantitatively and qualitatively. Our code is available at https://github.com/hkust-vgd/point_inverter. Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung |
WACV | 4 |
| 2023 | ACNet: Approaching-and-Centralizing Network for Zero-Shot Sketch-Based Image RetrievalabstractThe huge domain gap between sketches and photos poses huge challenges for Sketch-Based Image Retrieval (SBIR). The Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is more generic and practical but brings an even greater challenge: the additional knowledge gap between the seen and unseen categories. In order to simultaneously mitigate both gaps, we propose an Approaching-and-Centralizing Network (termed “ACNet”) to jointly optimize sketch-to-photo synthesis and image retrieval. The retrieval module guides the synthesis module to generate large amounts of diverse photo-like images that help the sketch domain gradually approach the photo domain to eliminate the domain gap, and thus better serves retrieval. Meanwhile, the retrieval module itself centralizes the embeddings of training samples for learning a similarity measurement to eliminate the knowledge gap. Our approach is simple yet effective, which achieves state-of-the-art performance on two widely used ZS-SBIR datasets and surpasses previous methods by a large margin (eg, 8.2% improvement in terms of mAP@all on TU-Berlin Extended dataset). Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Ying Shan, Sai-Kit Yeung |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | Neural Scene Decoration from a Single Photograph
Hong-Wing Pang, Yingshu Chen, Phuoc-Hieu Le, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung |
ECCV (23) | 6 |
| 2022 | RFNet-4D: Joint Object Reconstruction and Flow Estimation from 4D Point Clouds
Tuan-Anh Vu, Duc Thanh Nguyen, Binh-Son Hua, Quang-Hieu Pham, Sai-Kit Yeung |
ECCV (23) | 5 |
| 2022 | Time-of-Day Neural Style Transfer for Architectural PhotographsabstractArchitectural photography is a genre of photography that focuses on capturing a building or structure in the foreground with dramatic lighting in the background. Inspired by recent successes in image-to-image translation methods, we aim to perform style transfer for architectural photographs. However, the special composition in architectural photography poses great challenges for style transfer in this type of photographs. Existing neural style transfer methods treat the architectural images as a single entity, which would generate mismatched chrominance and destroy geometric features of the original architecture, yielding unrealistic lighting, wrong color rendition, and visual artifacts such as ghosting, appearance distortion, or color mismatching. In this paper, we specialize a neural style transfer method for architectural photography. Our method addresses the composition of the foreground and background in an architectural photograph in a two-branch neural network that separately considers the style transfer of the foreground and the background, respectively. Our method comprises a segmentation module, a learning-based image-to-image translation module, and an image blending optimization module. We trained our image-to-image translation neural network with a new dataset of unconstrained outdoor architectural photographs captured at different magic times of a day, utilizing additional semantic information for better chrominance matching and geometry preservation. Our experiments show that our method can produce photorealistic lighting and color rendition on both the foreground and background, and outperforms general image-to-image translation and arbitrary style transfer baselines quantitatively and qualitatively. Our code and data are available at https://github.com/hkust-vgd/architectural_style_transfer. Yingshu Chen, Tuan-Anh Vu, Ka-Chun Shum, Sai-Kit Yeung, Binh-Son Hua |
ICCP | 4 |
| 2022 | SiamX: An Efficient Long-term Tracker Using Cross-level Feature Correlation and Adaptive Tracking SchemeabstractSiamese network based trackers have achieved significant progress in visual object tracking. For the sake of speed, they mainly rely on offline training to learn a mono-level feature correlation between a target template and a search region. During the tracking period, they use a fixed strategy to infer target positions over sequences regardless of target states. However, such approaches are vulnerable in case of long-term challenges e.g. large variance, presence of distractors, fast motion, or target disappearing and the like. In this paper, we propose a new tracking framework, referred to as SiamX, by exploiting cross-level Siamese features to learn robust correlations between the target template and search regions, and also adaptive inference strategies to prevent tracking loss and realize fast target re-localization. Extensive experiments on four benchmarks including VOT-2019, LaSOT, GOT-10k, and TrackingNet show our method significantly enhances the tracker's ability to resist variance and interference, and achieve state-of-the-art results at around 50 FPS. Huajian Huang, Sai-Kit Yeung |
ICRA | 2 |
| 2022 | 360VO: Visual Odometry Using A Single 360 CameraabstractIn this paper, we propose a novel direct visual odometry algorithm to take the advantage of a 360-degree camera for robust localization and mapping. Our system extends direct sparse odometry by using a spherical camera model to process equirectangular images without rectification to attain omnidirectional perception. After adapting mapping and optimization algorithms to the new model, camera parameters, including intrinsic and extrinsic parameters, and 3D mapping can be jointly optimized within the local sliding window. In addition, we evaluate the proposed algorithm using both real world and large-scale simulated scenes for qualitative and quantitative validations. The extensive experiments indicate that our system achieves start of the art results. Huajian Huang, Sai-Kit Yeung |
ICRA | 2 |
| 2022 | 360ST-Mapping: An Online Semantics-Guided Topological Mapping Module for Omnidirectional Visual SLAMabstractAs an abstract representation of the environment structure, a topological map has advantageous properties for path-planning and navigation. Here we proposed an online topological mapping method, 360ST-Mapping, using omnidirectional vision. The 360° field-of-view allows the agent to obtain consistent observation and incrementally extract topological environment information. Moreover, we leverage semantic infor-mation to guide topological place recognition, further improving performance. The topological map possessing semantic infor-mation has the potential to support semantics-related advanced tasks. After integrating the topological mapping module into the omnidirectional visual SLAM system, we conducted extensive experiments in several large-scale indoor scenes and validated the method's effectiveness. Hongji Liu, Huajian Huang, Sai-Kit Yeung |
IROS | 3 |
| 2022 | RIConv++: Effective Rotation Invariant Convolutions for 3D Point Clouds Deep Learning
Zhiyuan Zhang 0004, Binh-Son Hua, Sai-Kit Yeung |
Int. J. Comput. Vis. | 3 |
| 2022 | Patch-Based Uncalibrated Photometric Stereo Under Natural IlluminationabstractThis paper presents a photometric stereo method that works with unknown natural illumination without any calibration objects or initial guess of the target shape. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary orthogonal ambiguity. We further build the patch connections by extracting consistent surface normal pairs via spatial overlaps among patches and intensity profiles. Guided by these connections, the local ambiguities are unified to a global orthogonal one through Markov Random Field optimization and rotation averaging. After applying the integrability constraint, our solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods. Heng Guo 0003, Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Ping Tan 0002, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Minimal Adversarial Examples for Deep Learning on 3D Point CloudsabstractWith recent developments of convolutional neural net-works, deep learning for 3D point clouds has shown significant progress in various 3D scene understanding tasks, e.g., object recognition, semantic segmentation. In a safety-critical environment, it is however not well understood how such deep learning models are vulnerable to adversarial examples. In this work, we explore adversarial attacks for point cloud-based neural networks. We propose a unified formulation for adversarial point cloud generation that can generalise two different attack strategies. Our method generates adversarial examples by attacking the classification ability of point cloud-based networks while considering the perceptibility of the examples and ensuring the minimal level of point manipulations. Experimental results show that our method achieves the state-of-the-art performance with higher than 89% and 90% of attack success rate on synthetic and real-world data respectively, while manipulating only about 4% of the total points. Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung |
ICCV | 4 |
| 2021 | KnitKit: a flexible system for machine knitting of customizable textilesabstractIn this work, we introduce KnitKit , a flexible and customizable system for the computational design and production of functional, multi-material, and three-dimensional knitted textiles. Our system greatly simplifies the knitting of 3D objects with complex, varying patterns that use multiple yarns and stitch patterns by separating the high-level design specification in terms of geometry, stitch patterns, materials or colors from the low-level, machine-specific knitting instruction generation. Starting from a triangular 3D mesh and a 2D texture that specifies knitting patterns on top of the geometry, our system generates the required machine instructions in three major steps. First, the input is processed and the KnitNet data structure is generated. This graph structure serves as an abstract interface between the high-level geometric and knitting configuration and the low-level, machine-specific knitting instructions. Second, a graph rewriting procedure is applied on the KnitNet that produces a sequence of abstract machine actions. Finally, the low-level machine instructions are generated by adapting those abstract actions to a specific machine context. We showcase the potential of this computational approach by designing and fabricating a variety of objects with complex geometries, multiple yarns, and multiple stitch patterns. Georges Nader, Yu Han Quek, Pei Zhi Chia, Oliver Weeger, Sai-Kit Yeung |
ACM Trans. Graph. | 5 |
| 2020 | Global Context Aware Convolutions for 3D Point Cloud UnderstandingabstractRecent advances in deep learning for 3D point clouds have shown great promises in scene understanding tasks thanks to the introduction of convolution operators to consume 3D point clouds directly in a neural network. Point cloud data, however, could have arbitrary rotations, especially those acquired from 3D scanning. Recent works show that it is possible to design point cloud convolutions with rotation invariance property, but such methods generally do not perform as well as translation-invariant only convolution. We found that a key reason is that compared to point coordinates, rotation-invariant features consumed by point cloud convolution are not as distinctive. To address this problem, we propose a novel convolution operator that enhances feature distinction by integrating global context information from the input point cloud to the convolution. To this end, a globally weighted local reference frame is constructed in each point neighborhood in which the local point set is decomposed into bins. Anchor points are generated in each bin to represent global shape features. A convolution can then be performed to transform the points and anchor features into final rotation-invariant features. We conduct several experiments on point cloud classification, part segmentation, shape retrieval, and normals estimation to evaluate our convolution, which achieves state-of-the-art accuracy under challenging rotations. Zhiyuan Zhang 0004, Binh-Son Hua, Yibin Tian, Sai-Kit Yeung |
3DV | 5 |
| 2020 | LCD: Learned Cross-Domain Descriptors for 2D-3D MatchingabstractIn this work, we present a novel method to learn a local cross-domain descriptor for 2D image and 3D point cloud matching. Our proposed method is a dual auto-encoder neural network that maps 2D and 3D input into a shared latent space representation. We show that such local cross-domain descriptors in the shared embedding are more discriminative than those obtained from individual training in 2D and 3D domains. To facilitate the training process, we built a new dataset by collecting ≈ 1.4 millions of 2D-3D correspondences with various lighting conditions and settings from publicly available RGB-D scenes. Our descriptor is evaluated in three main experiments: 2D-3D matching, cross-domain retrieval, and sparse-to-dense depth estimation. Experimental results confirm the robustness of our approach as well as its competitive performance not only in solving cross-domain tasks but also in being able to generalize to solve sole 2D and 3D tasks. Our dataset and code are released publicly at https://hkust-vgd.github.io/lcd. Quang-Hieu Pham, Mikaela Angelina Uy, Binh-Son Hua, Duc Thanh Nguyen, Gemma Roig, Sai-Kit Yeung |
AAAI | 6 |
| 2020 | SideInfNet: A Deep Neural Network for Semi-Automatic Semantic Segmentation with Side Information
Jing Yu Koh, Duc Thanh Nguyen, Quang-Trung Truong, Sai-Kit Yeung, Alexander Binder |
ECCV (24) | 4 |
| 2020 | Dual-SLAM: A framework for robust single camera navigationabstractSLAM (Simultaneous Localization And Mapping) seeks to provide a moving agent with real-time self-localization. To achieve real-time speed, SLAM incrementally propagates position estimates. This makes SLAM fast but also makes it vulnerable to local pose estimation failures. As local pose estimation is ill-conditioned, local pose estimation failures happen regularly, making the overall SLAM system brittle. This paper attempts to correct this problem. We note that while local pose estimation is ill-conditioned, pose estimation over longer sequences is well-conditioned. Thus, local pose estimation errors eventually manifest themselves as mapping inconsistencies. When this occurs, we save the current map and activate two new SLAM threads. One processes incoming frames to create a new map and the other, recovery thread, backtracks to link new and old maps together. This creates a Dual-SLAM framework that maintains real-time performance while being robust to local pose estimation failures. Evaluation on benchmark datasets shows Dual-SLAM can reduce failures by a dramatic 88%. Huajian Huang, Wen-Yan Lin, Sai-Kit Yeung |
IROS | 5 |
| 2020 | GMS: Grid-Based Motion Statistics for Fast, Ultra-robust Feature CorrespondenceabstractAbstract Feature matching aims at generating correspondences across images, which is widely used in many computer vision tasks. Although considerable progress has been made on feature descriptors and fast matching for initial correspondence hypotheses, selecting good ones from them is still challenging and critical to the overall performance. More importantly, existing methods often take a long computational time, limiting their use in real-time applications. This paper attempts to separate true correspondences from false ones at high speed. We term the proposed method (GMS) grid-based motion Statistics, which incorporates the smoothness constraint into a statistic framework for separation and uses a grid-based implementation for fast calculation. GMS is robust to various challenging image changes, involving in viewpoint, scale, and rotation. It is also fast, e.g., take only 1 or 2 ms in a single CPU thread, even when 50K correspondences are processed. This has important implications for real-time applications. What’s more, we show that incorporating GMS into the classic feature matching and epipolar geometry estimation pipeline can significantly boost the overall performance. Finally, we integrate GMS into the well-known ORB-SLAM system for monocular initialization, resulting in a significant improvement. Jiawang Bian, Wen-Yan Lin, Yun Liu 0011, Le Zhang 0001, Sai-Kit Yeung, Ming-Ming Cheng, Ian D. Reid 0001 |
Int. J. Comput. Vis. | 5 |
| 2020 | Ambiguity-Free Radiometric Calibration for Internet Photo CollectionsabstractRadiometrically calibrating nonlinear images from Internet photo collections makes photometric analysis applicable not only to lab data but also to big image data in the wild. However, conventional calibration methods cannot be directly applied to such photo collections. This paper presents a method to jointly perform radiometric calibration for a set of nonlinear images in Internet photo collections. By incorporating the consistency of scene reflectance of corresponding pixels across nonlinear images, the proposed method first estimates radiometric response functions of all the nonlinear images up to a unique exponential ambiguity using a rank minimization framework. The ambiguity is then resolved using the linear edge color blending constraint. Quantitative evaluation using both synthetic and real-world data shows the effectiveness of the proposed method. Zhipeng Mo, Boxin Shi, Sai-Kit Yeung, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Rotation Invariant Convolutions for 3D Point Clouds Deep LearningabstractRecent progresses in 3D deep learning has shown that it is possible to design special convolution operators to consume point cloud data. However, a typical drawback is that rotation invariance is often not guaranteed, resulting in networks that generalizes poorly to arbitrary rotations. In this paper, we introduce a novel convolution operator for point clouds that achieves rotation invariance. Our core idea is to use low-level rotation invariant geometric features such as distances and angles to design a convolution operator for point cloud learning. The well-known point ordering problem is also addressed by a binning approach seamlessly built into the convolution. This convolution operator then serves as the basic building block of a neural network that is robust to point clouds under 6-DoF transformations such as translation and rotation. Our experiment shows that our method performs with high accuracy in common scene understanding tasks such as object classification and segmentation. Compared to previous and concurrent works, most importantly, our method is able to generalize and achieve consistent results across different scenarios in which training and testing can contain arbitrary rotations. Our implementation is publicly available at our project page. Zhiyuan Zhang 0004, Binh-Son Hua, David W. Rosen, Sai-Kit Yeung |
3DV | 4 |
| 2019 | JSIS3D: Joint Semantic-Instance Segmentation of 3D Point Clouds With Multi-Task Pointwise Networks and Multi-Value Conditional Random FieldsabstractDeep learning techniques have become the to-go models for most vision-related tasks on 2D images. However, their power has not been fully realised on several tasks in 3D space, e.g., 3D scene understanding. In this work, we jointly address the problems of semantic and instance segmentation of 3D point clouds. Specifically, we develop a multi-task pointwise network that simultaneously performs two tasks: predicting the semantic classes of 3D points and embedding the points into high-dimensional vectors so that points of the same object instance are represented by similar embeddings. We then propose a multi-value conditional random field model to incorporate the semantic and instance labels and formulate the problem of semantic and instance segmentation as jointly optimising labels in the field model. The proposed method is thoroughly evaluated and compared with existing methods on different indoor scene datasets including S3DIS and SceneNN. Experimental results showed the robustness of the proposed joint semantic-instance segmentation scheme over its single components. Our method also achieved state-of-the-art performance on semantic segmentation. Quang-Hieu Pham, Duc Thanh Nguyen, Binh-Son Hua, Gemma Roig, Sai-Kit Yeung |
CVPR | 5 |
| 2019 | Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataabstractDeep learning techniques for point cloud data have demonstrated great potentials in solving classical problems in 3D computer vision such as 3D object classification and segmentation. Several recent 3D object classification methods have reported state-of-the-art performance on CAD model datasets such as ModelNet40 with high accuracy (~92\%). Despite such impressive results, in this paper, we argue that object classification is still a challenging task when objects are framed with real-world settings. To prove this, we introduce ScanObjectNN, a new real-world point cloud object dataset based on scanned indoor scene data. From our comprehensive benchmark, we show that our dataset poses great challenges to existing point cloud classification techniques as objects from real-world scans are often cluttered with background and/or are partial due to occlusions. We identify three key open problems for point cloud object classification, and propose new point cloud classification neural networks that achieve state-of-the-art performance on classifying objects with cluttered background. Our dataset and code are publicly available in our project page https://hkust-vgd.github.io/scanobjectnn/. Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung |
ICCV | 5 |
| 2019 | ShellNet: Efficient Point Cloud Convolutional Neural Networks Using Concentric Shells StatisticsabstractDeep learning with 3D data has progressed significantly since the introduction of convolutional neural networks that can handle point order ambiguity in point cloud data. While being able to achieve good accuracies in various scene understanding tasks, previous methods often have low training speed and complex network architecture. In this paper, we address these problems by proposing an efficient end-to-end permutation invariant convolution for point cloud deep learning. Our simple yet effective convolution operator named ShellConv uses statistics from concentric spherical shells to define representative features and resolve the point order ambiguity, allowing traditional convolution to perform on such features. Based on ShellConv we further build an efficient neural network named ShellNet to directly consume the point clouds with larger receptive fields while maintaining less layers. We demonstrate the efficacy of ShellNet by producing state-of-the-art results on object classification, object part segmentation, and semantic scene segmentation while keeping the network very fast to train. Zhiyuan Zhang 0004, Binh-Son Hua, Sai-Kit Yeung |
ICCV | 3 |
| 2019 | Force-based Heterogeneous Traffic Simulation for Autonomous Vehicle TestingabstractRecent failures in real-world self-driving tests have suggested a paradigm shift from directly learning in real-world roads to building a high-fidelity driving simulator as an alternative, effective, and safe tool to handle intricate traffic environments in urban areas. To date, traffic simulation can construct virtual urban environments with various weather conditions, day and night, and traffic control for autonomous vehicle testing. However, mutual interactions between autonomous vehicles and pedestrians are rarely modeled in existing simulators. Besides vehicles and pedestrians, the usage of personal mobility devices is increasing in congested cities as an alternative to the traditional transport system. A simulator that considers all potential road-users in a realistic urban environment is urgently desired. In this work, we propose a novel, extensible, and microscopic method to build heterogenous traffic simulation using the force-based concept. This force-based approach can accurately replicate the sophisticated behaviors of various road users and their interactions through a simple and unified way. Furthermore, we validate our approach through simulation experiments and comparisons to the popular simulators currently used for research and development of autonomous vehicles. Qianwen Chao, Xiaogang Jin 0001, Hen-Wei Huang, Shaohui Foong, Lap-Fai Yu, Sai-Kit Yeung |
ICRA | 6 |
| 2019 | Real-Time Progressive 3D Semantic Segmentation for Indoor ScenesabstractThe widespread adoption of autonomous systems such as drones and assistant robots has created a need for real-time high-quality semantic scene segmentation. In this paper, we propose an efficient yet robust technique for on-the-fly dense reconstruction and semantic segmentation of 3D indoor scenes. To guarantee (near) real-time performance, our method is built atop an efficient super-voxel clustering method and a conditional random field with higher-order constraints from structural and object cues, enabling progressive dense semantic segmentation without any precomputation. We extensively evaluate our method on different indoor scenes including kitchens, offices, and bedrooms in the SceneNN and ScanNet datasets and show that our technique consistently produces state-of-the-art segmentation results in both qualitative and quantitative experiments. Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung |
WACV | 4 |
| 2019 | 3D articulated skeleton extraction using a single consumer-grade depth camera
Xuequan Lu, Zhigang Deng 0001, Jun Luo 0001, Wenzhi Chen, Sai-Kit Yeung, Ying He 0001 |
Comput. Vis. Image Underst. | 5 |
| 2019 | A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric StereoabstractClassic photometric stereo is often extended to deal with real-world materials and work with unknown lighting conditions for practicability. To quantitatively evaluate non-Lambertian and uncalibrated photometric stereo, a photometric stereo image dataset containing objects of various shapes with complex reflectance properties and high-quality ground truth normals is still missing. In this paper, we introduce the 'DiLiGenT' dataset with calibrated Directional Lightings, objects of General reflectance with different shininess, and 'ground Truth' normals from high-precision laser scanning. We use our dataset to quantitatively evaluate state-of-the-art photometric stereo methods for general materials and unknown lighting conditions, selected from a newly proposed photometric stereo taxonomy emphasizing non-Lambertian and uncalibrated methods. The dataset and evaluation results are made publicly available, and we hope it can serve as a benchmark platform that inspires future research. Boxin Shi, Zhipeng Mo, Dinglong Duan, Sai-Kit Yeung, Ping Tan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Unsupervised Articulated Skeleton Extraction From Point Set Sequences Captured by a Single Depth CameraabstractHow to robustly and accurately extract articulated skeletons from point set sequences captured by a single consumer-grade depth camera still remains to be an unresolved challenge to date. To address this issue, we propose a novel, unsupervised approach consisting of three contributions (steps): (i) a non-rigid point set registration algorithm to first build one-to-one point correspondences among the frames of a sequence; (ii) a skeletal structure extraction algorithm to generate a skeleton with reasonable numbers of joints and bones; (iii) a skeleton joints estimation algorithm to achieve accurate joints. At the end, our method can produce a quality articulated skeleton from a single 3D point sequence corrupted with noise and outliers. The experimental results show that our approach soundly outperforms state of the art techniques, in terms of both visual quality and accuracy. Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Zhigang Deng 0001, Wenzhi Chen |
AAAI | 3 |
| 2018 | Pointwise Convolutional Neural NetworksabstractDeep learning with 3D data such as reconstructed point clouds and CAD models has received great research interests recently. However, the capability of using point clouds with convolutional neural network has been so far not fully explored. In this paper, we present a convolutional neural network for semantic segmentation and object recognition with 3D point clouds. At the core of our network is point-wise convolution, a new convolution operator that can be applied at each point of a point cloud. Our fully convolutional network design, while being surprisingly simple to implement, can yield competitive accuracy in both semantic segmentation and object recognition task. Binh-Son Hua, Minh-Khoi Tran, Sai-Kit Yeung |
CVPR | 3 |
| 2018 | Uncalibrated Photometric Stereo Under Natural IlluminationabstractThis paper presents a photometric stereo method that works with unknown natural illuminations without any calibration object. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary rotation ambiguity. Our method connects the resulting patches and unifies the local ambiguities to a global rotation one through angular distance propagation defined over the whole surface. After applying the integrability constraint, our final solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods. Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Yasuyuki Matsushita |
CVPR | 4 |
| 2018 | Self-Calibrating Polarising Radiometric CalibrationabstractWe present a self-calibrating polarising radiometric calibration method. From a set of images taken from a single viewpoint under different unknown polarising angles, we recover the inverse camera response function and the polarising angles relative to the first angle. The problem is solved in an integrated manner, recovering both of the unknowns simultaneously. The method exploits the fact that the intensity of polarised light should vary sinusoidally as the polarising filter is rotated, provided that the response is linear. It offers the first solution to demonstrate the possibility of radiometric calibration through polarisation. We evaluate the accuracy of our proposed method using synthetic data and real world objects captured using different cameras. The self-calibrated results were found to be comparable with those from multiple exposure sequence. Daniel Teo, Boxin Shi, Yinqiang Zheng, Sai-Kit Yeung |
CVPR | 4 |
| 2018 | Urban Zoning Using Higher-Order Markov Random Fields on Multi-View Imagery Data
Tian Feng 0001, Quang-Trung Truong, Duc Thanh Nguyen, Jing Yu Koh, Lap-Fai Yu, Alexander Binder, Sai-Kit Yeung |
ECCV (8) | 7 |
| 2018 | CODE: Coherence Based Decision Boundaries for Feature CorrespondenceabstractA key challenge in feature correspondence is the difficulty in differentiating true and false matches at a local descriptor level. This forces adoption of strict similarity thresholds that discard many true matches. However, if analyzed at a global level, false matches are usually randomly scattered while true matches tend to be coherent (clustered around a few dominant motions), thus creating a coherence based separability constraint. This paper proposes a non-linear regression technique that can discover such a coherence based separability constraint from highly noisy matches and embed it into a correspondence likelihood model. Once computed, the model can filter the entire set of nearest neighbor matches (which typically contains over 90 percent false matches) for true matches. We integrate our technique into a full feature correspondence system which reliably generates large numbers of good quality correspondences over wide baselines where previous techniques provide few or no matches. Wen-Yan Lin, Fan Wang 0010, Ming-Ming Cheng, Sai-Kit Yeung, Philip Torr 0001, Minh N. Do, Jiangbo Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Language-driven synthesis of 3D scenes from scene databasesabstractWe introduce a novel framework for using natural language to generate and edit 3D indoor scenes, harnessing scene semantics and text-scene grounding knowledge learned from large annotated 3D scene databases. The advantage of natural language editing interfaces is strongest when performing semantic operations at the sub-scene level, acting on groups of objects. We learn how to manipulate these sub-scenes by analyzing existing 3D scenes. We perform edits by first parsing a natural language command from the user and transforming it into a semantic scene graph that is used to retrieve corresponding sub-scenes from the databases that match the command. We then augment this retrieved sub-scene by incorporating other objects that may be implied by the scene context. Finally, a new 3D scene is synthesized by aligning the augmented sub-scene with the user's current scene, where new objects are spliced into the environment, possibly triggering appropriate adjustments to the existing scene arrangement. A suggestive modeling interface with multiple interpretations of user commands is used to alleviate ambiguities in natural language. We conduct studies comparing our approach against both prior text-to-scene work and artist-made scenes and find that our method significantly outperforms prior work and is comparable to handmade scenes even when complex and varied natural sentences are used. Rui Ma 0011, Akshay Gadi Patil, Matthew Fisher, Manyi Li, Sören Pirk, Binh-Son Hua, Sai-Kit Yeung, Xin Tong 0001, Leonidas J. Guibas, Hao (Richard) Zhang |
ACM Trans. Graph. | 7 |
| 2018 | GPF: GMM-Inspired Feature-Preserving Point Set FilteringabstractPoint set filtering, which aims at reconstructing noise-free point sets from their corresponding noisy inputs, is a fundamental problem in 3D geometry processing. The main challenge of point set filtering is to preserve geometric features of the underlying geometry while at the same time removing the noise. State-of-the-art point set filtering methods still struggle with this issue: some are not designed to recover sharp features, and others cannot well preserve geometric features, especially fine-scale features. In this paper, we propose a novel approach for robust feature-preserving point set filtering, inspired by the Gaussian Mixture Model (GMM). Taking a noisy point set and its filtered normals as input, our method can robustly reconstruct a high-quality point set which is both noise-free and feature-preserving. Various experiments show that our approach can soundly outperform the selected state-of-the-art methods, in terms of both filtering quality and reconstruction accuracy. Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Wenzhi Chen, Matthias Zwicker |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | A Robust 3D-2D Interactive Tool for Scene Segmentation and AnnotationabstractRecent advances of 3D acquisition devices have enabled large-scale acquisition of 3D scene data. Such data, if completely and well annotated, can serve as useful ingredients for a wide spectrum of computer vision and graphics works such as data-driven modeling and scene understanding, object detection and recognition. However, annotating a vast amount of 3D scene data remains challenging due to the lack of an effective tool and/or the complexity of 3D scenes (e.g. clutter, varying illumination conditions). This paper aims to build a robust annotation tool that effectively and conveniently enables the segmentation and annotation of massive 3D data. Our tool works by coupling 2D and 3D information via an interactive framework, through which users can provide high-level semantic annotation for objects. We have experimented our tool and found that a typical indoor scene could be well segmented and annotated in less than 30 minutes by using the tool, as opposed to a few hours if done manually. Along with the tool, we created a dataset of over a hundred 3D scenes associated with complete annotations using our tool. Both the tool and dataset will be available at http://scenenn.net. Duc Thanh Nguyen, Binh-Son Hua, Lap-Fai Yu, Sai-Kit Yeung |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | GMS: Grid-Based Motion Statistics for Fast, Ultra-Robust Feature CorrespondenceabstractIncorporating smoothness constraints into feature matching is known to enable ultra-robust matching. However, such formulations are both complex and slow, making them unsuitable for video applications. This paper proposes GMS (Grid-based Motion Statistics), a simple means of encapsulating motion smoothness as the statistical likelihood of a certain number of matches in a region. GMS enables translation of high match numbers into high match quality. This provides a real-time, ultra-robust correspondence system. Evaluation on videos, with low textures, blurs and wide-baselines show GMS consistently out-performs other real-time matchers and can achieve parity with more sophisticated, much slower techniques. Jiawang Bian, Wen-Yan Lin, Yasuyuki Matsushita, Sai-Kit Yeung, Tan-Dat Nguyen, Ming-Ming Cheng |
CVPR | 4 |
| 2017 | Radiometric Calibration for Internet Photo CollectionsabstractRadiometrically calibrating the images from Internet photo collections brings photometric analysis from lab data to big image data in the wild, but conventional calibration methods cannot be directly applied to such image data. This paper presents a method to jointly perform radiometric calibration for a set of images in an Internet photo collection. By incorporating the consistency of scene reflectance for corresponding pixels in multiple images, the proposed method estimates radiometric response functions of all the images using a rank minimization framework. Our calibration aligns all response functions in an image set up to the same exponential ambiguity in a robust manner. Quantitative results using both synthetic and real data show the effectiveness of the proposed method. Zhipeng Mo, Boxin Shi, Sai-Kit Yeung, Yasuyuki Matsushita |
CVPR | 3 |
| 2017 | A Novel Riemannian Metric Based on Riemannian Structure and Scaling Information for Fixed Low-Rank Matrix CompletionabstractRiemannian optimization has been widely used to deal with the fixed low-rank matrix completion problem, and Riemannian metric is a crucial factor of obtaining the search direction in Riemannian optimization. This paper proposes a new Riemannian metric via simultaneously considering the Riemannian geometry structure and the scaling information, which is smoothly varying and invariant along the equivalence class. The proposed metric can make a tradeoff between the Riemannian geometry structure and the scaling information effectively. Essentially, it can be viewed as a generalization of some existing metrics. Based on the proposed Riemanian metric, we also design a Riemannian nonlinear conjugate gradient algorithm, which can efficiently solve the fixed low-rank matrix completion problem. By experimenting on the fixed low-rank matrix completion, collaborative filtering, and image and video recovery, it illustrates that the proposed method is superior to the state-of-the-art methods on the convergence efficiency and the numerical performance. Shasha Mao, Licheng Jiao, Tian Feng 0001, Sai-Kit Yeung |
IEEE Trans. Cybern. | 5 |
| 2017 | Approximate dissectionsabstractA geometric dissection is a set of pieces which can be assembled in different ways to form distinct shapes. Dissections are used as recreational puzzles because it is striking when a single set of pieces can construct highly different forms. Existing techniques for creating dissections find pieces that reconstruct two input shapes exactly. Unfortunately, these methods only support simple, abstract shapes because an excessive number of pieces may be needed to reconstruct more complex, naturalistic shapes. We introduce a dissection design technique that supports such shapes by requiring that the pieces reconstruct the shapes only approximately. We find that, in most cases, a small number of pieces suffices to tightly approximate the input shapes. We frame the search for a viable dissection as a combinatorial optimization problem, where the goal is to search for the best approximation to the input shapes using a given number of pieces. We find a lower bound on the tightness of the approximation for a partial dissection solution, which allows us to prune the search space and makes the problem tractable. We demonstrate our approach on several challenging examples, showing that it can create dissections between shapes of significantly greater complexity than those supported by previous techniques. Noah Duncan, Lap-Fai Yu, Sai-Kit Yeung, Demetri Terzopoulos |
ACM Trans. Graph. | 3 |
| 2016 | SceneNN: A Scene Meshes Dataset with aNNotationsabstractSeveral RGB-D datasets have been publicized over the past few years for facilitating research in computer vision and robotics. However, the lack of comprehensive and fine-grained annotation in these RGB-D datasets has posed challenges to their widespread usage. In this paper, we introduce SceneNN, an RGB-D scene dataset consisting of 100 scenes. All scenes are reconstructed into triangle meshes and have per-vertex and per-pixel annotation. We further enriched the dataset with fine-grained information such as axis-aligned bounding boxes, oriented bounding boxes, and object poses. We used the dataset as a benchmark to evaluate the state-of-the-art methods on relevant research problems such as intrinsic decomposition and shape completion. Our dataset and annotation tools are available at http://www.scenenn.net. Binh-Son Hua, Quang-Hieu Pham, Duc Thanh Nguyen, Minh-Khoi Tran, Lap-Fai Yu, Sai-Kit Yeung |
3DV | 6 |
| 2016 | A Field Model for Repairing 3D ShapesabstractThis paper proposes a field model for repairing 3D shapes constructed from multi-view RGB data. Specifically, we represent a 3D shape in a Markov random field (MRF) in which the geometric information is encoded by random binary variables and the appearance information is retrieved from a set of RGB images captured at multiple viewpoints. The local priors in the MRF model capture the local structures of object shapes and are learnt from 3D shape templates using a convolutional deep belief network. Repairing a 3D shape is formulated as the maximum a posteriori (MAP) estimation in the corresponding MRF. Variational mean field approximation technique is adopted for the MAP estimation. The proposed method was evaluated on both artificial data and real data obtained from reconstruction of practical scenes. Experimental results have shown the robustness and efficiency of the proposed method in repairing noisy and incomplete 3D shapes. Duc Thanh Nguyen, Binh-Son Hua, Minh-Khoi Tran, Quang-Hieu Pham, Sai-Kit Yeung |
CVPR | 5 |
| 2016 | A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric StereoabstractRecent progress on photometric stereo extends the technique to deal with general materials and unknown illumination conditions. However, due to the lack of suitable benchmark data with ground truth shapes (normals), quantitative comparison and evaluation is difficult to achieve. In this paper, we first survey and categorize existing methods using a photometric stereo taxonomy emphasizing on non-Lambertian and uncalibrated methods. We then introduce the 'DiLiGenT' photometric stereo image dataset with calibrated Directional Lightings, objects of General reflectance, and 'ground Truth' shapes (normals). Based on our dataset, we quantitatively evaluate state-of-the-art photometric stereo methods for general non-Lambertian materials and unknown lightings to analyze their strengths and limitations. Boxin Shi, Zhipeng Mo, Dinglong Duan, Sai-Kit Yeung, Ping Tan 0002 |
CVPR | 5 |
| 2016 | View-Aware Image Object Compositing and Synthesis from Multiple Sources
Xiang Chen 0001, Weiwei Xu 0003, Sai-Kit Yeung, Kun Zhou 0001 |
J. Comput. Sci. Technol. | 3 |
| 2016 | Interchangeable components for hands-on assembly based modellingabstractInterchangeable components allow an object to be easily reconfigured, but usually reveal that the object is composed of parts. In this work, we present a computational approach for the design of components which are interchangeable, but also form objects with a coherent appearance which conceals their composition from parts. These components allow a physical realization of Assembly Based Modelling, a popular virtual modelling paradigm in which new models are constructed from the parts of existing ones. Given a collection of 3D models and a segmentation that specifies the component connectivity, our approach generates the components by jointly deforming and partitioning the models. We determine the component boundaries by evolving a set of closed contours on the input models to maximize the contours' geometric similarity. Next, we efficiently deform the input models to enforce both C0 and C1 continuity between components while minimizing deviation from their original appearance. The user can guide our deformation scheme to preserve desired features. We demonstrate our approach on several challenging examples, showing that our components can be physically reconfigured to assemble a large variety of coherent shapes. Noah Duncan, Lap-Fai Yu, Sai-Kit Yeung |
ACM Trans. Graph. | 3 |
| 2016 | Crowd-driven mid-scale layout designabstractWe propose a novel approach for designing mid-scale layouts by optimizing with respect to human crowd properties. Given an input layout domain such as the boundary of a shopping mall, our approach synthesizes the paths and sites by optimizing three metrics that measure crowd flow properties: mobility, accessibility, and coziness. While these metrics are straightforward to evaluate by a full agent-based crowd simulation, optimizing a layout usually requires hundreds of evaluations, which would require a long time to compute even using the latest crowd simulation techniques. To overcome this challenge, we propose a novel data-driven approach where nonlinear regressors are trained to capture the relationship between the agent-based metrics, and the geometrical and topological features of a layout. We demonstrate that by using the trained regressors, our approach can synthesize crowd-aware layouts and improve existing layouts with better crowd flow properties. Tian Feng 0001, Lap-Fai Yu, Sai-Kit Yeung, KangKang Yin, Kun Zhou 0001 |
ACM Trans. Graph. | 3 |
| 2016 | 3D Navigation on Impossible Figures via Dynamically Reconfigurable MazeabstractPrevious research on impossible figures focuses extensively on single view modeling and rendering. Existing computer games that employ impossible figures as navigation maze for gaming either use a fixed third-person view with axonometric projection to retain the figure's impossibility perception, or simply break the figure's impossibility upon view changes. In this paper, we present a new approach towards 3D gaming with impossible figures, delivering for the first time navigation in 3D mazes constructed from impossible figures. Such result cannot be achieved by previous research work in modeling impossible figures. To deliver seamless gaming navigation and interaction, we propose i) a set of guiding principles for bringing out subtle perceptions and ii) a novel computational approach to construct 3D structures from impossible figure images and then to dynamically construct the impossible-figure maze subjected to user's view. In the end, we demonstrate and discuss our method with a variety of generic maze types. Chi-Fu William Lai, Sai-Kit Yeung, Xiaoqi Yan, Chi-Wing Fu, Chi-Keung Tang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | The Clutterpalette: An Interactive Tool for Detailing Indoor ScenesabstractWe introduce the Clutterpalette, an interactive tool for detailing indoor scenes with small-scale items. When the user points to a location in the scene, the Clutterpalette suggests detail items for that location. In order to present appropriate suggestions, the Clutterpalette is trained on a dataset of images of real-world scenes, annotated with support relations. Our experiments demonstrate that the adaptive suggestions presented by the Clutterpalette increase modeling speed and enhance the realism of indoor scenes. Lap-Fai Yu, Sai-Kit Yeung, Demetri Terzopoulos |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Unbounded High Dynamic Range Photography Using a Modulo CameraabstractThis paper presents a novel framework to extend the dynamic range of images called Unbounded High Dynamic Range (UHDR) photography with a modulo camera. A modulo camera could theoretically take unbounded radiance levels by keeping only the least significant bits. We show that with limited bit depth, very high radiance levels can be recovered from a single modulus image with our newly proposed unwrapping algorithm for natural images. We can also obtain an HDR image with details equally well preserved for all radiance levels by merging the least number of modulus images. Synthetic experiment and experiment with a real modulo camera show the effectiveness of the proposed approach. Hang Zhao 0021, Boxin Shi, Christy Fernandez-Cull, Sai-Kit Yeung, Ramesh Raskar |
ICCP | 4 |
| 2015 | An MRF-Poselets Model for Detecting Highly Articulated HumansabstractDetecting highly articulated objects such as humans is a challenging problem. This paper proposes a novel part-based model built upon poselets, a notion of parts, and Markov Random Field (MRF) for modelling the human body structure under the variation of human poses and viewpoints. The problem of human detection is then formulated as maximum a posteriori (MAP) estimation in the MRF model. Variational mean field method, a robust statistical inference, is adopted to approximate the MAP estimation. The proposed method was evaluated and compared with existing methods on different test sets including H3D and PASCAL VOC 2007-2009. Experimental results have favourbly shown the robustness of the proposed method in comparison to the state-of-the-art. Duc Thanh Nguyen, Minh-Khoi Tran, Sai-Kit Yeung |
ICCV | 3 |
| 2015 | Fill and Transfer: A Simple Physics-Based Approach for Containability ReasoningabstractThe visual perception of object affordances has emerged as a useful ingredient for building powerful computer vision and robotic applications. In this paper we introduce a novel approach to reason about liquid containability - the affordance of containing liquid. Our approach analyzes container objects based on two simple physical processes: the Fill and Transfer of liquid. First, it reasons about whether a given 3D object is a liquid container and its best filling direction. Second, it proposes directions to transfer its contained liquid to the outside while avoiding spillage. We compare our simplified model with a common fluid dynamics simulation and demonstrate that our algorithm makes human-like choices about the best directions to fill containers and transfer liquid from them. We apply our approach to reason about the containability of several real-world objects acquired using a consumer-grade depth camera. Lap-Fai Yu, Noah Duncan, Sai-Kit Yeung |
ICCV | 3 |
| 2015 | Efficient direct rendering of deforming surfaces via shared subdivision trees
Fuchang Liu, Sai-Kit Yeung, Markus Gross 0001 |
Comput. Aided Des. | 3 |
| 2015 | Matching-constrained active contours with affine-invariant shape prior
Junyan Wang 0002, Sai-Kit Yeung, Kap Luk Chan |
Comput. Vis. Image Underst. | 2 |
| 2015 | Normal Estimation of a Transparent Object Using a VideoabstractReconstructing transparent objects is a challenging problem. While producing reasonable results for quite complex objects, existing approaches require custom calibration or somewhat expensive labor to achieve high precision. When an overall shape preserving salient and fine details is sufficient, we show in this paper a significant step toward solving the problem when the object's silhouette is available and simple user interaction is allowed, by using a video of a transparent object shot under varying illumination. Specifically, we estimate the normal map of the exterior surface of a given solid transparent object, from which the surface depth can be integrated. Our technical contribution lies in relating this normal estimation problem to one of graph-cut segmentation. Unlike conventional formulations, however, our graph is dual-layered, since we can see a transparent object's foreground as well as the background behind it. Quantitative and qualitative evaluation are performed to verify the efficacy of this practical solution. Sai-Kit Yeung, Tai-Pang Wu, Chi-Keung Tang, Tony F. Chan, Stanley J. Osher |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Weighted classifier ensemble based on quadratic form
Shasha Mao, Licheng Jiao, Shuiping Gou, Bo Chen 0001, Sai-Kit Yeung |
Pattern Recognit. | 6 |
| 2015 | Zoomorphic designabstractZoomorphic shapes are man-made shapes that possess the form or appearance of an animal. They have desirable aesthetic properties, but are difficult to create using conventional modeling tools. We present a method for creating zoomorphic shapes by merging a man-made shape and an animal shape. To identify a pair of shapes that are suitable for merging, we use an efficient graph kernel based technique. We formulate the merging process as a continuous optimization problem where the two shapes are deformed jointly to minimize an energy function combining several design factors. The modeler can adjust the weighting between these factors to attain high-level control over the final shape produced. A novel technique ensures that the zoomorphic shape does not violate the design restrictions of the man-made shape. We demonstrate the versatility and effectiveness of our approach by generating a wide variety of zoomorphic shapes. Noah Duncan, Lap-Fai Yu, Sai-Kit Yeung, Demetri Terzopoulos |
ACM Trans. Graph. | 3 |
| 2014 | Photometric Stereo Using Internet ImagesabstractPhotometric stereo using unorganized Internet images is very challenging, because the input images are captured under unknown general illuminations, with uncontrolled cameras. We propose to solve this difficult problem by a simple yet effective approach that makes use of a coarse shape prior. The shape prior is obtained from multi-view stereo and will be useful in twofold: resolving the shape-light ambiguity in uncalibrated photometric stereo and guiding the estimated normals to produce the high quality 3D surface. By assuming the surface albedo is not highly contrasted, we also propose a novel linear approximation of the nonlinear camera responses with our normal estimation algorithm. We evaluate our method using synthetic data and demonstrate the surface improvement on real data over multi-view stereo results. Boxin Shi, Kenji Inose, Yasuyuki Matsushita, Ping Tan 0002, Sai-Kit Yeung, Katsushi Ikeuchi |
3DV | 5 |
| 2014 | Sub-pixel Layout for Super-Resolution with Images in the Octic Group
Boxin Shi, Hang Zhao 0021, Moshe Ben-Ezra, Sai-Kit Yeung, Christy Fernandez-Cull, R. Hamilton Shepard, Christopher Barsi, Ramesh Raskar |
ECCV (1) | 4 |
| 2013 | Shading-Based Shape Refinement of RGB-D ImagesabstractWe present a shading-based shape refinement algorithm which uses a noisy, incomplete depth map from Kinect to help resolve ambiguities in shape-from-shading. In our framework, the partial depth information is used to overcome bas-relief ambiguity in normals estimation, as well as to assist in recovering relative albedos, which are needed to reliably estimate the lighting environment and to separate shading from albedo. This refinement of surface normals using a noisy depth map leads to high-quality 3D surfaces. The effectiveness of our algorithm is demonstrated through several challenging real-world examples. Lap-Fai Yu, Sai-Kit Yeung, Yu-Wing Tai, Stephen Lin 0001 |
CVPR | 2 |
| 2013 | Outdoor photometric stereoabstractWe introduce a framework for outdoor photometric stereo utilizing natural environmental illumination. Our framework extends beyond existing photometric stereo methods intended for laboratory environments to encompass robust outdoor operation in the real world. In this paper, we motivate our framework, describe the components of its processing pipeline, and assess its performance in synthetic experiments as well as in natural experiments including objects in outdoor environments with complex real-world illuminations. Lap-Fai Yu, Sai-Kit Yeung, Yu-Wing Tai, Demetri Terzopoulos, Tony F. Chan |
ICCP | 2 |
| 2013 | Transmural Imaging of Ventricular Action Potentials and Post-Infarction Scars in Swine HeartsabstractThe problem of using surface data to reconstruct transmural electrophysiological (EP) signals is intrinsically ill-posed without a unique solution in its unconstrained form. Incorporating physiological spatiotemporal priors through probabilistic integration of dynamic EP models, we have previously developed a Bayesian approach to transmural electrophysiological imaging (TEPI) using body-surface electrocardiograms. In this study, we generalize TEPI to using electrical signals collected from heart surfaces, and we test its feasibility on two pre-clinical swine models provided through the STACOM 2011 EP simulation Challenge. Since this new application of TEPI does not require whole-body imaging, there may be more immediate potential in EP laboratories where it could utilize catheter mapping data and produce transmural information for therapy guidance. Another focus of this study is to investigate the consistency among three modalities in delineating scar after myocardial infarction: TEPI, electroanatomical voltage mapping (EAVM), and magnetic resonance imaging (MRI). Our preliminary data demonstrate that, compared to the low-voltage scar area in EAVM, the 3-D electrical scar volume detected by TEPI is more consistent with anatomical scar volume delineated in MRI. Furthermore, TEPI could complement anatomical imaging by providing EP functional features related to both scar and healthy tissue. Fady Dawoud, Sai-Kit Yeung, Ken C. L. Wong, Huafeng Liu 0003, Albert C. Lardo |
IEEE Trans. Medical Imaging | 3 |
| 2012 | A Closed-Form Solution to Tensor Voting: Theory and ApplicationsabstractWe prove a closed-form solution to tensor voting (CFTV): Given a point set in any dimensions, our closed-form solution provides an exact, continuous, and efficient algorithm for computing a structure-aware tensor that simultaneously achieves salient structure detection and outlier attenuation. Using CFTV, we prove the convergence of tensor voting on a Markov random field (MRF), thus termed as MRFTV, where the structure-aware tensor at each input site reaches a stationary state upon convergence in structure propagation. We then embed structure-aware tensor into expectation maximization (EM) for optimizing a single linear structure to achieve efficient and robust parameter estimation. Specifically, our EMTV algorithm optimizes both the tensor and fitting parameters and does not require random sampling consensus typically used in existing robust statistical techniques. We performed quantitative evaluation on its accuracy and robustness, showing that EMTV performs better than the original TV and other state-of-the-art techniques in fundamental matrix estimation for multiview stereo matching. The extensions of CFTV and EMTV for extracting multiple and nonlinear structures are underway. Tai-Pang Wu, Sai-Kit Yeung, Jiaya Jia, Chi-Keung Tang, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | DressUp!: outfit synthesis through automatic optimizationabstractWe present an automatic optimization approach to outfit synthesis. Given the hair color, eye color, and skin color of the input body, plus a wardrobe of clothing items, our outfit synthesis system suggests a set of outfits subject to a particular dress code. We introduce a probabilistic framework for modeling and applying dress codes that exploits a Bayesian network trained on example images of real-world outfits. Suitable outfits are then obtained by optimizing a cost function that guides the selection of clothing items to maximize the color compatibility and dress code suitability. We demonstrate our approach on the four most common dress codes:Casual, Sportswear, Business-Casual, andBusiness. A perceptual study validated on multiple resultant outfits demonstrates the efficacy of our framework. Lap-Fai Yu, Sai-Kit Yeung, Demetri Terzopoulos, Tony F. Chan |
ACM Trans. Graph. | 2 |
| 2011 | Adequate reconstruction of transparent objects on a shoestring budgetabstractReconstructing transparent objects is a challenging problem. While producing reasonable results for quite complex objects, existing approaches require custom calibration or somewhat expensive labor to achieve high precision. On the other hand, when an overall shape preserving salient and fine details is sufficient, we show in this paper a significant step toward solving the problem on a shoestring budget, by using only a video camera, a moving spotlight, and a small chrome sphere. Specifically, the problem we address is to estimate the normal map of the exterior surface of a given solid transparent object, from which the surface depth can be integrated. Our technical contribution lies in relating this normal reconstruction problem to one of graph-cut segmentation. Unlike conventional formulations, however, our graph is dual-layered, since we can see a transparent object's foreground as well as the background behind it. Quantitative and qualitative evaluation are performed to verify the efficacy of this practical solution. Sai-Kit Yeung, Tai-Pang Wu, Chi-Keung Tang, Tony F. Chan, Stanley J. Osher |
CVPR | 1 |
| 2011 | Matting and compositing of transparent and refractive objectsabstractThis article introduces a new approach for matting and compositing transparent and refractive objects in photographs. The key to our work is an image-based matting model, termed the Attenuation-Refraction Matte (ARM), that encodes plausible refractive properties of a transparent object along with its observed specularities and transmissive properties. We show that an object's ARM can be extracted directly from a photograph using simple user markup. Once extracted, the ARM is used to paste the object onto a new background with a variety of effects, including compound compositing, Fresnel effect, scene depth, and even caustic shadows. User studies find our results favorable to those obtained with Photoshop as well as perceptually valid in most cases. Our approach allows photo editing of transparent and refractive objects in a manner that produces realistic effects previously only possible via 3D models or environment matting. Sai-Kit Yeung, Chi-Keung Tang, Michael S. Brown, Sing Bing Kang |
ACM Trans. Graph. | 1 |
| 2011 | Make it home: automatic optimization of furniture arrangementabstractWe present a system that automatically synthesizes indoor scenes realistically populated by a variety of furniture objects. Given examples of sensibly furnished indoor scenes, our system extracts, in advance, hierarchical and spatial relationships for various furniture objects, encoding them into priors associated with ergonomic factors, such as visibility and accessibility, which are assembled into a cost function whose optimization yields realistic furniture arrangements. To deal with the prohibitively large search space, the cost function is optimized by simulated annealing using a Metropolis-Hastings state search step. We demonstrate that our system can synthesize multiple realistic furniture arrangements and, through a perceptual study, investigate whether there is a significant difference in the perceived functionality of the automatically synthesized results relative to furniture arrangements produced by human designers. Lap-Fai Yu, Sai-Kit Yeung, Chi-Keung Tang, Demetri Terzopoulos, Tony F. Chan, Stanley J. Osher |
ACM Trans. Graph. | 2 |
| 2010 | Quasi-dense 3D reconstruction using tensor-based multiview stereoabstractWe propose tensor-based multiview stereo (TMVS) for quasi-dense 3D reconstruction from uncalibrated images. Our work is inspired by the patch-based multiview stereo (PMVS), a state-of-the-art technique in multiview stereo reconstruction. The effectiveness of PMVS is attributed to the use of 3D patches in the match-propagate-filter MVS pipeline. Our key observation is: PMVS has not fully utilized the valuable 3D geometric cue available in 3D patches which are oriented points. This paper combines the complementary advantages of photoconsistency, visibility and geometric consistency enforcement in MVS via the use of 3D tensors, where our closed-form solution to tensor voting provides a unified approach to implement the match-propagate-filter pipeline. Using PMVS as the implementation backbone where TMVS is built, we provide qualitative and quantitative evaluation to demonstrate how TMVS significantly improve the MVS pipeline. Tai-Pang Wu, Sai-Kit Yeung, Jiaya Jia, Chi-Keung Tang |
CVPR | 2 |
| 2010 | Modeling and rendering of impossible figuresabstractThis article introduces an optimization approach for modeling and rendering impossible figures. Our solution is inspired by how modeling artists construct physical 3D models to produce a valid 2D view of an impossible figure. Given a set of 3D locally possible parts of the figure, our algorithm automatically optimizes a view-dependent 3D model, subject to the necessary 3D constraints for rendering the impossible figure at the desired novel viewpoint. A linear and constrained least-squares solution to the optimization problem is derived, thereby allowing an efficient computation and rendering new views of impossible figures at interactive rates. Once the optimized model is available, a variety of compelling rendering effects can be applied to the impossible figure. Tai-Pang Wu, Chi-Wing Fu, Sai-Kit Yeung, Jiaya Jia, Chi-Keung Tang |
ACM Trans. Graph. | 3 |
| 2008 | Enforcing stochastic inverse consistency in non-rigid image registration and matchingabstractThis paper presents a new method to enforce inverse consistency in nonrigid image registration and matching. Conventional approaches assume diffeomorphic transformation, implicitly or explicitly. However, the inherent smoothness constraint discourages discontinuity consideration. We propose a post-processing algorithm that integrates the input forward and backward fields, which are output by existing registration/matching algorithms, to produce more robust results. Given such a pair of input fields, our algorithm alternately refines the fields by tensor belief propagation, and enforces inverse consistency in stochastic sense by generalized total least squares fitting. To show the efficacy of our stochastic inverse consistency approach, we first present results on very noisy fields. We then demonstrate improvement on existing stereo matching where occlusion is naturally handled by localizing violations of inverse consistency. Finally, we propose a novel application on image stitching, where stochastic inverse consistency is employed in structure deformation, in order to seamlessly align overlapping images with severe misalignment in structure and intensity. Sai-Kit Yeung, Chi-Keung Tang, Josien P. W. Pluim, Max A. Viergever, Albert C. S. Chung, Helen C. Shen |
CVPR | 1 |
| 2008 | Extracting smooth and transparent layers from a single imageabstractLayer decomposition from a single image is an under-constrained problem, because there are more unknowns than equations. This paper studies a slightly easier but very useful alternative where only the background layer has substantial image gradients and structures. We propose to solve this useful alternative by an expectation-maximization (EM) algorithm that employs the hidden markov model (HMM), which maintains spatial coherency of smooth and overlapping layers, and helps to preserve image details of the textured background layer. We demonstrate that, using a small amount of user input, various seemingly unrelated problems in computational photography can be effectively addressed by solving this alternative using our EM-HMM algorithm. Sai-Kit Yeung, Tai-Pang Wu, Chi-Keung Tang |
CVPR | 1 |
| 2005 | Stochastic Inverse Consistency in Medical Image Registration
Sai-Kit Yeung |
MICCAI (2) | 1 |