EDBT 2026 Demo / reviewers in the wild / expert
Kwan-Yee Kenneth Wong
dblp:w/KwanYeeKennethWong · also Kwan-Yee K. Wong
· DBLP profile ↗
112ranked-venue papers
11as first author
39since 2021 · last 2026
0000-0001-8560-9007ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 87 · 8 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 7 first-author · 26 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LooC: Effective Low-Dimensional Codebook for Compositional Vector QuantizationabstractVector quantization (VQ) is a prevalent and fundamental technique that discretizes continuous feature vectors by approximating them using a codebook. As the diversity and complexity of data and models continue to increase, there is an urgent need for high-capacity, yet more compact VQ methods. This paper aims to reconcile this conflict by presenting a new approach called LooC, which utilizes an effective Low-dimensional codebook for Compositional vector quantization. Firstly, LooC introduces a parameter-efficient codebook by reframing the relationship between codevectors and feature vectors, significantly expanding its solution space. Instead of individually matching codevectors with feature vectors, LooC treats them as lower-dimensional compositional units within feature vectors and combines them, resulting in a more compact codebook with improved performance. Secondly, LooC incorporates a parameter-free extrapolation-by-interpolation mechanism to enhance and smooth features during the VQ process, which allows for better preservation of details and fidelity in feature approximation. The design of LooC leads to full codebook usage, effectively utilizing the compact codebook while avoiding the problem of collapse. Thirdly, LooC can serve as a plug-and-play module for existing methods for different downstream tasks based on VQ. Finally, extensive evaluations on different tasks, datasets, and architectures demonstrate that LooC outperforms existing VQ methods, achieving state-of-the-art performance with a significantly smaller codebook. Jie Li 0040, Kwan-Yee Kenneth Wong, Kai Han 0001 |
WACV | 2 |
| 2026 | Learning Coherent Portrait-to-Anime Translation via Latent Cyclic TransformationabstractTranslating real portrait video into anime is an application of interest to both consumers and researchers. However, anime differs considerably from portraits, making portrait-to-anime translation challenging. Existing StyleGAN-based portrait stylization works assume that the portrait and stylized generators share the same latent space, but this assumption fails in the style of anime due to the large domain gap. Moreover, directly applying them to each video frame often leads to undesirable temporal inconsistencies. In this paper, we argue that two latent spaces with a large domain gap cannot be shared but can be related by a transformation, and develop a cyclic transformation network to connect the two spaces with two cycle constraints. This provides high-quality translation for each frame. We extend our framework to video transformation by proposing a novel frame interpolation constraint which ensures that in-between frames can be interpolated from their neighboring frames, guaranteeing temporal coherence across translated frames. Together with latent code smoothing regularization, this provides temporally coherent video-to-anime translation. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods both qualitatively and quantitatively. Yangyang Xu 0003, Shengfeng He, Kwan-Yee Kenneth Wong, Ping Luo 0002 |
Comput. Vis. Media | 3 |
| 2025 | Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language NavigationabstractLLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) task. However, existing LLM-based methods often focus only on solving high-level task planning by selecting nodes in predefined navigation graphs for movements, overlooking low-level control in navigation scenarios. To bridge this gap, we propose AO-Planner, a novel Affordances-Oriented Planner for continuous VLN task. Our AO-Planner integrates various foundation models to achieve affordances-oriented low-level motion planning and high-level decision-making, both performed in a zero-shot setting. Specifically, we employ a Visual Affordances Prompting (VAP) approach, where the visible ground is segmented by SAM to provide navigational affordances, based on which the LLM selects potential candidate waypoints and plans low-level paths towards selected waypoints. We further propose a high-level PathAgent which marks planned paths into the image input and reasons the most probable path by comprehending all environmental information. Finally, we convert the selected path into 3D coordinates using camera intrinsic parameters and depth information, avoiding challenging 3D predictions for LLMs. Experiments on the challenging R2R-CE and RxR-CE datasets show that AO-Planner achieves state-of-the-art zero-shot performance (8.8% improvement on SPL). Our method can also serve as a data annotator to obtain pseudo-labels, distilling its waypoint prediction ability into a learning-based predictor. This new predictor does not require any waypoint data from the simulator and achieves 47% SR competing with supervised methods. We establish an effective connection between LLM and 3D world, presenting novel prospects for employing foundation models in low-level motion control. Bingqian Lin, Xinmin Liu, Lin Ma 0002, Xiaodan Liang, Kwan-Yee Kenneth Wong |
AAAI | 6 |
| 2025 | DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous DrivingabstractMultimodal large language models (MLLMs) possess the ability to comprehend visual images or videos, and show impressive reasoning ability thanks to the vast amounts of pretrained knowledge, making them highly suitable for autonomous driving applications. Unlike the previous work, DriveGPT4-V1, which focused on open-loop tasks, this study explores the capabilities of LLMs in enhancing closed-loop autonomous driving. DriveGPT4-V2 processes camera images and vehicle states as input to generate low-level control signals for end-to-end vehicle operation. A multi-view visual tokenizer (MV-VT) is employed enabling DriveGPT4-V2 to perceive the environment with an extensive range while maintaining critical details. The model architecture has been refined to improve decision prediction and inference speed. To further enhance the performance, an additional expert LLM is trained for online imitation learning. The expert LLM, sharing a similar structure with DriveGPT4-V2, can access privileged information about surrounding objects for more robust and reliable predictions. Experimental results show that DriveGPT4-V2 outperforms all baselines on the challenging CARLA Longest6 benchmark. The code and data of DriveGPT4-V2 will be publicly available. Zhenhua Xu 0003, Yujia Zhang 0003, Zhuoling Li, Kwan-Yee Kenneth Wong, Hengshuang Zhao |
CVPR | 6 |
| 2025 | ArtiFade: Learning to Generate High-quality Subject from Blemished ImagesabstractSubject-driven text-to-image generation has demonstrated remarkable advancements in its ability to learn and capture characteristics of a subject using only a limited number of images. However, existing methods commonly rely on high-quality images for training and often struggle to generate reasonable images when the input images are blemished by artifacts. This is primarily attributed to the inadequate capability of current techniques in distinguishing subject-related features from disruptive artifacts. In this paper, we introduce ArtiFade to tackle this issue and successfully generate high-quality artifact-free images from blemished datasets. Specifically, ArtiFade exploits fine-tuning of a pre-trained text-to-image model, aiming to remove artifacts. The elimination of artifacts is achieved by utilizing a specialized dataset that encompasses both unblemished images and their corresponding blemished counterparts during fine-tuning. ArtiFade also ensures the preservation of the original generative capabilities inherent within the diffusion model, thereby enhancing the overall performance of subject-driven methods in generating high-quality and artifact-free images. We further devise evaluation benchmarks tailored for this task. Through extensive qualitative and quantitative experiments, we demonstrate the generalizability of ArtiFade in effective artifact removal under both in-distribution and out-of-distribution scenarios. Shuya Yang, Shaozhe Hao, Kwan-Yee Kenneth Wong |
CVPR | 4 |
| 2025 | Rethinking Cross-Modal Interaction in Multimodal Diffusion TransformersabstractMultimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-driven visual generation. However, even state-of-the-art MM-DiT models like FLUX struggle with achieving precise alignment between text prompts and generated content. We identify two key issues in the attention mechanism of MM-DiT, namely 1) the suppression of cross-modal attention due to token imbalance between visual and textual modalities and 2) the lack of timestep-aware attention weighting, which hinder the alignment. To address these issues, we propose \textbf{Temperature-Adjusted Cross-modal Attention (TACA)}, a parameter-efficient method that dynamically rebalances multimodal interactions through temperature scaling and timestep-dependent adjustment. When combined with LoRA fine-tuning, TACA significantly enhances text-image alignment on the T2I-CompBench benchmark with minimal computational overhead. We tested TACA on state-of-the-art models like FLUX and SD3.5, demonstrating its ability to improve image-text alignment in terms of object appearance, attribute binding, and spatial relationships. Our findings highlight the importance of balancing cross-modal attention in improving semantic fidelity in text-to-image diffusion models. Our codes are publicly available at \href{https://github.com/Vchitect/TACA} Zhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen 0009, Wangmeng Zuo, Ziwei Liu 0002, Kwan-Yee Kenneth Wong |
ICCV | 7 |
| 2025 | Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
Zhengyao Lv, Chenyang Si, Tianlin Pan, Zhaoxi Chen 0009, Kwan-Yee Kenneth Wong, Yu Qiao 0001, Ziwei Liu 0002 |
ICCV | 5 |
| 2025 | AvatarGO: Zero-shot 4D Human-Object Interaction Generation and AnimationabstractRecent advancements in diffusion models have led to significant improvements in the generation and animation of 4D full-body human-object interactions (HOI). Nevertheless, existing methods primarily focus on SMPL-based motion generation, which is limited by the scarcity of realistic large-scale interaction data. This constraint affects their ability to create everyday HOI scenes. This paper addresses this challenge using a zero-shot approach with a pre-trained diffusion model. Despite this potential, achieving our goals is difficult due to the diffusion model's lack of understanding of ''where'' and ''how'' objects interact with the human body. To tackle these issues, we introduce **AvatarGO**, a novel framework designed to generate animatable 4D HOI scenes directly from textual inputs. Specifically, **1)** for the ''where'' challenge, we propose **LLM-guided contact retargeting**, which employs Lang-SAM to identify the contact body part from text prompts, ensuring precise representation of human-object spatial relations. **2)** For the ''how'' challenge, we introduce **correspondence-aware motion optimization** that constructs motion fields for both human and object models using the linear blend skinning function from SMPL-X. Our framework not only generates coherent compositional motions, but also exhibits greater robustness in handling penetration issues. Extensive experiments with existing methods validate AvatarGO's superior generation and animation capabilities on a variety of human-object pairs and diverse poses. As the first attempt to synthesize 4D avatars with object interactions, we hope AvatarGO could open new doors for human-centric 4D content creation. Liang Pan, Kai Han 0001, Kwan-Yee Kenneth Wong, Ziwei Liu 0002 |
ICLR | 4 |
| 2025 | BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation CapabilitiesabstractWe introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation capabilities. BiGR is the first conditional generative model that unifies generation and discrimination within the same framework.
BiGR features a binary tokenizer, a masked modeling mechanism, and a binary transcoder for binary code prediction.
Additionally, we introduce a novel entropy-ordered sampling method to enable efficient image generation.
Extensive experiments validate BiGR's superior performance in generation quality, as measured by FID-50k, and representation capabilities, as evidenced by linear-probe accuracy.
Moreover, BiGR showcases zero-shot generalization across various vision tasks, enabling applications such as image inpainting, outpainting, editing, interpolation, and enrichment, without the need for structural modifications. Our findings suggest that BiGR unifies generative and discriminative tasks effectively, paving the way for further advancements in the field. We further enable BiGR to perform text-to-image generation, showcasing its potential for broader applications. Shaozhe Hao, Xuantong Liu, Xianbiao Qi, Bojia Zi, Rong Xiao 0003, Kai Han 0001, Kwan-Yee Kenneth Wong |
ICLR | 8 |
| 2025 | FasterCache: Training-Free Video Diffusion Model Acceleration with High QualityabstractIn this paper, we present \textbf{\textit{FasterCache}}, a novel training-free strategy designed to accelerate the inference of video diffusion models with high-quality generation. By analyzing existing cache-based methods, we observe that \textit{directly reusing adjacent-step features degrades video quality due to the loss of subtle variations}. We further perform a pioneering investigation of the acceleration potential of classifier-free guidance (CFG) and reveal significant redundancy between conditional and unconditional features within the same timestep. Capitalizing on these observations, we introduce FasterCache to substantially accelerate diffusion-based video generation. Our key contributions include a dynamic feature reuse strategy that preserves both feature distinction and temporal continuity, and CFG-Cache which optimizes the reuse of conditional and unconditional outputs to further enhance inference speed without compromising video quality. We empirically evaluate FasterCache on recent video diffusion models. Experimental results show that FasterCache can significantly accelerate video generation (\eg 1.67$\times$ speedup on Vchitect-2.0) while keeping video quality comparable to the baseline, and consistently outperform existing methods in both inference speed and video quality. \textit{Our code will be made public upon publication.} Zhengyao Lv, Chenyang Si, Yu Qiao 0001, Ziwei Liu 0002, Kwan-Yee Kenneth Wong |
ICLR | 7 |
| 2025 | DiffusionMat: Alpha Matting as Deterministic Sequential Refinement Learning
Yangyang Xu 0003, Shengfeng He, Wenqi Shao, Yong Du 0003, Kwan-Yee Kenneth Wong, Yu Qiao 0001, Jun Yu 0002, Ping Luo 0002 |
ACM Multimedia | 5 |
| 2025 | SPC: Evolving Self-Play Critic via Adversarial Games for LLM ReasoningabstractEvaluating the step-by-step reliability of large language model (LLM) reasoning, such as Chain-of-Thought, remains challenging due to the difficulty and cost of obtaining high-quality step-level supervision. In this paper, we introduce Self-Play Critic (SPC), a novel approach where a critic model evolves its ability to assess reasoning steps through adversarial self-play games, eliminating the need for manual step-level annotation. SPC involves fine-tuning two copies of a base model to play two roles, namely a "sneaky generator" that deliberately produces erroneous steps designed to be difficult to detect, and a "critic" that analyzes the correctness of reasoning steps. These two models engage in an adversarial game in which the generator aims to fool the critic, while the critic model seeks to identify the generator's errors. Using reinforcement learning based on the game outcomes, the models iteratively improve; the winner of each confrontation receives a positive reward and the loser receives a negative reward, driving continuous self-evolution. Experiments on three reasoning process benchmarks (ProcessBench, PRM800K, DeltaBench) demonstrate that our SPC progressively enhances its error detection capabilities (e.g., accuracy increases from 70.8% to 77.7% on ProcessBench) and surpasses strong baselines, including distilled R1 model. Furthermore, SPC can guide the test-time search of diverse LLMs and significantly improve their mathematical reasoning performance on MATH500 and AIME2024, surpassing those guided by state-of-the-art process reward models. Bang Zhang, Ruotian Ma, Peisong Wang 0002, Xiaodan Liang, Zhaopeng Tu, Kwan-Yee Kenneth Wong |
NeurIPS | 8 |
| 2025 | VipDiff: Towards Coherent and Diverse Video Inpainting via Training-Free Denoising Diffusion ModelsabstractRecent video inpainting methods have achieved encouraging improvements by leveraging optical flow to guide pixel propagation from reference frames, either in the image space or feature space. However, they would produce severe artifacts when the masked area is too large and no pixel correspondences could be found. Recently, denoising diffusion models have demonstrated impressive performance in generating diverse and high-quality images, and have been exploited in a number of works for image inpainting. These methods, however, cannot be applied directly to videos to produce temporal-coherent inpainting results. In this paper, we propose a training-free framework, named VipDiff, for conditioning diffusion model on the reverse diffusion process to produce temporal-coherent inpainting results without requiring any training data or fine-tuning the pre-trained models. VipDiff takes optical flow as guidance to extract valid pixels from reference frames to serve as constraints in optimizing the randomly sampled Gaussian noise, and uses the generated results for further pixel propagation and conditional generation. VipDiff also allows for generating diverse video inpainting results over different sampled noise. Experiments demonstrate that our VipDiff outperforms state-of-the-art methods in terms of both spatial-temporal coherence and fidelity. Chaohao Xie, Kai Han 0001, Kwan-Yee Kenneth Wong |
WACV | 3 |
| 2025 | RIGID: Recurrent GAN Inversion and Editing of Real Face Videos and Beyond
Yangyang Xu 0003, Shengfeng He, Kwan-Yee Kenneth Wong, Ping Luo 0002 |
Int. J. Comput. Vis. | 3 |
| 2024 | MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language NavigationabstractEmbodied agents equipped with GPT as their brains have exhibited extraordinary decisionmaking and generalization abilities across various tasks.However, existing zero-shot agents for vision-and-language navigation (VLN) only prompt GPT-4 to select potential locations within localized environments, without constructing an effective "global-view" for the agent to understand the overall environment.In this work, we present a novel map-guided GPT-based agent, dubbed MapGPT, which introduces an online linguistic-formed map to encourage global exploration.Specifically, we build an online map and incorporate it into the prompts that include node information and topological relationships, to help GPT understand the spatial environment.Benefiting from this design, we further propose an adaptive planning mechanism to assist the agent in performing multi-step path planning based on a map, systematically exploring multiple candidate nodes or sub-goals step by step.Extensive experiments demonstrate that our MapGPT is applicable to both GPT-4 and GPT-4V, achieving state-of-the-art zero-shot performance on R2R and REVERIE simultaneously (∼10% and ∼12% improvements in SR), and showcasing the newly emergent global thinking and path planning abilities of the GPT. Bingqian Lin, Zhenhua Chai, Xiaodan Liang, Kwan-Yee Kenneth Wong |
ACL (1) | 6 |
| 2024 | DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion ModelsabstractWe present DreamAvatar, a text-and-shape guided framework for generating high-quality 3D human avatars with controllable poses. While encouraging results have been reported by recent methods on text-guided 3D common object generation, generating high-quality human avatars remains an open challenge due to the complexity of the human body's shape, pose, and appearance. We propose DreamAvatar to tackle this challenge, which utilizes a train-able NeRF for predicting density and color for 3D points and pretrained text-to-image diffusion models for providing 2D self-supervision. Specifically, we leverage the SMPL model to provide shape and pose guidance for the generation. We introduce a dual-observation-space design that involves the joint optimization of a canonical space and a posed space that are related by a learnable deformation field. This facilitates the generation of more complete textures and geometry faithful to the target pose. We also jointly optimize the losses computed from the full body and from the zoomed-in 3D head to alleviate the common multi-face “Janus” problem and improve facial details in the generated avatars. Extensive evaluations demonstrate that DreamAvatar significantly outperforms existing meth-ods, establishing a new state-of-the-art for text-and-shape guided 3D human avatar generation. Yan-Pei Cao 0001, Kai Han 0001, Ying Shan, Kwan-Yee Kenneth Wong |
CVPR | 5 |
| 2024 | PLACE: Adaptive Layout-Semantic Fusion for Semantic Image SynthesisabstractRecent advancements in large-scale pre-trained text-to-image models have led to remarkable progress in semantic image synthesis. Nevertheless, synthesizing high-quality images with consistent semantics and layout remains a challenge. In this paper, we propose the adaPtive LAyout-semantiC fusion modulE (PLACE) that harnesses pre-trained models to alleviate the aforementioned issues. Specifically, we first employ the layout control map to faithfully represent layouts in the feature space. Subsequently, we combine the layout and semantic features in a timestep-adaptive manner to synthesize images with realistic details. During fine-tuning, we propose the Semantic Alignment (SA) loss to further enhance layout alignment. Additionally, we introduce the Layout-Free Prior Preservation (LFP) loss, which leverages unlabeled data to maintain the priors of pre-trained models, thereby improving the visual quality and semantic consistency of synthesized images. Extensive experiments demonstrate that our approach performs favorably in terms of visual quality, semantic consistency, and layout alignment. The source code and model are available at PLACE. Zhengyao Lv, Yuxiang Wei 0001, Wangmeng Zuo, Kwan-Yee Kenneth Wong |
CVPR | 4 |
| 2024 | ConceptExpress: Harnessing Diffusion Models for Single-Image Unsupervised Concept Extraction
Shaozhe Hao, Kai Han 0001, Zhengyao Lv, Kwan-Yee Kenneth Wong |
ECCV (59) | 5 |
| 2024 | InsMapper: Exploring Inner-Instance Information for Vectorized HD Mapping
Zhenhua Xu 0003, Kwan-Yee Kenneth Wong, Hengshuang Zhao |
ECCV (34) | 2 |
| 2024 | Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
Shaozhe Hao, Bojia Zi, Huaizhe Xu, Kwan-Yee Kenneth Wong |
ECCV (81) | 5 |
| 2024 | Deep Variational Network Toward Blind Image RestorationabstractBlind image restoration (IR) is a common yet challenging problem in computer vision. Classical model-based methods and recent deep learning (DL)-based methods represent two different methodologies for this problem, each with their own merits and drawbacks. In this paper, we propose a novel blind image restoration method, aiming to integrate both the advantages of them. Specifically, we construct a general Bayesian generative model for the blind IR, which explicitly depicts the degradation process. In this proposed model, a pixel-wise non-i.i.d. Gaussian distribution is employed to fit the image noise. It is with more flexibility than the simple i.i.d. Gaussian or Laplacian distributions as adopted in most of conventional methods, so as to handle more complicated noise types contained in the image degradation. To solve the model, we design a variational inference algorithm where all the expected posteriori distributions are parameterized as deep neural networks to increase their model capability. Notably, such an inference algorithm induces a unified framework to jointly deal with the tasks of degradation estimation and image restoration. Further, the degradation information estimated in the former task is utilized to guide the latter IR process. Experiments on two typical blind IR tasks, namely image denoising and super-resolution, demonstrate that the proposed method achieves superior performance over current state-of-the-arts. Zongsheng Yue, Hongwei Yong, Qian Zhao 0002, Lei Zhang 0006, Deyu Meng, Kwan-Yee Kenneth Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | SeSDF: Self-Evolved Signed Distance Field for Implicit 3D Clothed Human ReconstructionabstractWe address the problem of clothed human reconstruction from a single image or uncalibrated multiview images. existing methods struggle with reconstructing detailed geometry of a clothed human and often require a calibrated setting for multiview reconstruction. We propose a flexible framework which, by leveraging the parametric SMPL-X model, can take an arbitrary number of input images to reconstruct a clothed human model under an uncalibrated setting. At the core of our framework is our novel self-evolved signed distance field (SeSDF) module which allows the framework to learn to deform the signed distance field (SDF) derived from the fitted SMPL-X model, such that detailed geometry reflecting the actual clothed human can be encoded for better reconstruction. Besides, we propose a simple method for self-calibration of multiview images via the fitted SMPL-X parameters. This lifts the requirement of tedious manual calibration and largely increases the flexibility of our method. Further, we introduce an effective occlusion-aware feature fusion strategy to account for the most useful features to reconstruct the human model. We thoroughly evaluate our framework on public benchmarks, demonstrating significant superiority over the state-of-the-arts both qualitatively and quantitatively. Kai Han 0001, Kwan-Yee Kenneth Wong |
CVPR | 3 |
| 2023 | Learning Attention as Disentangler for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) aims at learning visual concepts (i.e., attributes and objects) from seen compositions and combining concept knowledge into unseen compositions. The key to CZSL is learning the disentanglement of the attribute-object composition. To this end, we propose to exploit cross-attentions as compositional disentanglers to learn disentangled concept embeddings. For example, if we want to recognize an unseen composition “yellow flower”, we can learn the attribute concept “yellow” and object concept “flower” from different yellow objects and different flowers respectively. To further constrain the disentanglers to learn the concept of interest, we employ a regularization at the attention level. Specifically, we adapt the earth mover's distance (EMD) as a feature similarity metric in the cross-attention module. Moreover, benefiting from concept disentanglement, we improve the inference process and tune the prediction score by combining multiple concept probabilities. Comprehensive experiments on three CZSL benchmark datasets demonstrate that our method significantly outperforms previous works in both closed- and open-world settings, establishing a new state-of-the-art. Project page: https://haoosz.github.io/ade-czsl/ Shaozhe Hao, Kai Han 0001, Kwan-Yee Kenneth Wong |
CVPR | 3 |
| 2023 | RIGID: Recurrent GAN Inversion and Editing of Real Face VideosabstractGAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a unified recurrent framework, named Recurrent vIdeo GAN Inversion and eDiting (RIGID), to explicitly and simultaneously enforce temporally coherent GAN inversion and facial editing of real videos. Our approach models the temporal relations between current and previous frames from three aspects. To enable a faithful real video reconstruction, we first maximize the inversion fidelity and consistency by learning a temporal compensated latent code. Second, we observe incoherent noises lie in the high-frequency domain that can be disentangled from the latent space. Third, to remove the inconsistency after attribute manipulation, we propose an in-between frame composition constraint such that the arbitrary frame must be a direct composite of its neighboring frames. Our unified framework learns the inherent coherence between input frames in an end-to-end manner, and therefore it is agnostic to a specific attribute and can be applied to arbitrary editing of the same video without re-training. Extensive experiments demonstrate that RIGID outperforms state-of-the-art methods qualitatively and quantitatively in both inversion and editing tasks. The deliverables can be found in https://cnnlstm.github.io/RIGID. Yangyang Xu 0003, Shengfeng He, Kwan-Yee Kenneth Wong, Ping Luo 0002 |
ICCV | 3 |
| 2023 | HeadSculpt: Crafting 3D Head Avatars with TextabstractRecently, text-guided 3D generative methods have made remarkable advancements in producing high-quality textures and geometry, capitalizing on the proliferation of large vision-language and image diffusion models.
However, existing methods still struggle to create high-fidelity 3D head avatars in two aspects:
(1) They rely mostly on a pre-trained text-to-image diffusion model whilst missing the necessary 3D awareness and head priors.
This makes them prone to inconsistency and geometric distortions in the generated avatars.
(2) They fall short in fine-grained editing. This is primarily due to the inherited limitations from the pre-trained 2D image diffusion models, which become more pronounced when it comes to 3D head avatars.
In this work, we address these challenges by introducing a versatile coarse-to-fine pipeline dubbed HeadSculpt for crafting (i.e., generating and editing) 3D head avatars from textual prompts.
Specifically, we first equip the diffusion model with 3D awareness by leveraging landmark-based control and a learned textual embedding representing the back view appearance of heads, enabling 3D-consistent head avatar generations.
We further propose a novel identity-aware editing score distillation strategy to optimize a textured mesh with a high-resolution differentiable rendering technique.
This enables identity preservation while following the editing instruction.
We showcase HeadSculpt's superior fidelity and editing capabilities through comprehensive experiments and comparisons with existing methods. Kai Han 0001, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song, Tao Xiang 0002, Kwan-Yee Kenneth Wong |
NeurIPS | 8 |
| 2023 | Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsabstractText-to-Image diffusion models have made tremendous progress over the past two years, enabling the generation of highly realistic images based on open-domain text descriptions. However, despite their success, text descriptions often struggle to adequately convey detailed controls, even when composed of long and complex texts. Moreover, recent studies have also shown that these models face challenges in understanding such complex texts and generating the corresponding images. Therefore, there is a growing need to enable more control modes beyond text description. In this paper, we introduce Uni-ControlNet, a unified framework that allows for the simultaneous utilization of different local controls (e.g., edge maps, depth map, segmentation masks) and global controls (e.g., CLIP image embeddings) in a flexible and composable manner within one single model. Unlike existing methods, Uni-ControlNet only requires the fine-tuning of two additional adapters upon frozen pre-trained text-to-image diffusion models, eliminating the huge cost of training from scratch. Moreover, thanks to some dedicated adapter designs, Uni-ControlNet only necessitates a constant number (i.e., 2) of adapters, regardless of the number of local or global controls used. This not only reduces the fine-tuning costs and model size, making it more suitable for real-world deployment, but also facilitate composability of different conditions. Through both quantitative and qualitative comparisons, Uni-ControlNet demonstrates its superiority over existing methods in terms of controllability, generation quality and composability. Code is available at https://github.com/ShihaoZhaoZSH/Uni-ControlNet. Dongdong Chen 0001, Yen-Chun Chen 0001, Jianmin Bao, Shaozhe Hao, Lu Yuan 0001, Kwan-Yee Kenneth Wong |
NeurIPS | 7 |
| 2023 | Semi-supervised Cycle-GAN for face photo-sketch translation in the wild
Chaofeng Chen, Wei Liu 0091, Xiao Tan 0001, Kwan-Yee Kenneth Wong |
Comput. Vis. Image Underst. | 4 |
| 2023 | Deep Face Video Inpainting via UV MappingabstractThis paper addresses the problem of face video inpainting. Existing video inpainting methods target primarily at natural scenes with repetitive patterns. They do not make use of any prior knowledge of the face to help retrieve correspondences for the corrupted face. They therefore only achieve sub-optimal results, particularly for faces under large pose and expression variations where face components appear very differently across frames. In this paper, we propose a two-stage deep learning method for face video inpainting. We employ 3DMM as our 3D face prior to transform a face between the image space and the UV (texture) space. In Stage I, we perform face inpainting in the UV space. This helps to largely remove the influence of face poses and expressions and makes the learning task much easier with well aligned face features. We introduce a frame-wise attention module to fully exploit correspondences in neighboring frames to assist the inpainting task. In Stage II, we transform the inpainted face regions back to the image space and perform face video refinement that inpaints any background regions not covered in Stage I and also refines the inpainted face regions. Extensive experiments have been carried out which show our method can significantly outperform methods based merely on 2D information, especially for faces under large pose and expression variations. Project page: https://ywq.github.io/FVIP. Wenqi Yang, Zhenfang Chen, Chaofeng Chen, Guanying Chen, Kwan-Yee Kenneth Wong |
IEEE Trans. Image Process. | 5 |
| 2022 | JIFF: Jointly-aligned Implicit Face Function for High Quality Single View Clothed Human ReconstructionabstractThis paper addresses the problem of single view 3D human reconstruction. Recent implicit function based methods have shown impressive results, but they fail to recover fine face details in their reconstructions. This largely degrades user experience in applications like 3D telepresence. In this paper, we focus on improving the quality of face in the reconstruction and propose a novel Jointly-aligned Implicit Face Function (JIFF) that combines the merits of the implicit function based approach and model based approach. We employ a 3D morphable face model as our shape prior and compute space-aligned 3D features that capture detailed face geometry information. Such space-aligned 3D features are combined with pixel-aligned 2D features to jointly predict an implicit face function for high quality face reconstruction. We further extend our pipeline and introduce a coarse-to-fine architecture to predict high quality texture for our detailed face model. Extensive evaluations have been carried out on public datasets and our proposed JIFF has demonstrates superior performance (both quantitatively and qualitatively) over existing state-of-the-arts. Guanying Chen, Kai Han 0001, Wenqi Yang, Kwan-Yee Kenneth Wong |
CVPR | 5 |
| 2022 | Blind Image Super-resolution with Elaborate Degradation Modeling on Noise and KernelabstractWhile researches on model-based blind single image super-resolution (SISR) have achieved tremendous successes recently, most of them do not consider the image degradation sufficiently. Firstly, they always assume image noise obeys an independent and identically distributed (i.i.d.) Gaussian or Laplacian distribution, which largely underestimates the complexity of real noise. Secondly, previous commonly-used kernel priors (e.g., normalization, sparsity) are not effective enough to guarantee a rational kernel solution, and thus degenerates the performance of subsequent SISR task. To address the above issues, this paper proposes a model-based blind SISR method under the probabilistic framework, which elaborately models image degradation from the perspectives of noise and blur kernel. Specifically, instead of the traditional i.i.d. noise assumption, a patch-based non-i.i.d. noise model is proposed to tackle the complicated real noise, expecting to increase the degrees of freedom of the model for noise representation. As for the blur kernel, we novelly construct a concise yet effective kernel generator, and plug it into the proposed blind SISR method as an explicit kernel prior (EKP). To solve the proposed model, a theoretically grounded Monte Carlo EM algorithm is specifically designed. Comprehensive experiments demonstrate the superiority of our method over current state-of-the-arts on synthetic and real datasets. The source code is available at https://github.com/zsyOAOA/BSRDM. Zongsheng Yue, Qian Zhao 0002, Jianwen Xie, Lei Zhang 0006, Deyu Meng, Kwan-Yee Kenneth Wong |
CVPR | 6 |
| 2022 | PS-NeRF: Neural Inverse Rendering for Multi-view Photometric Stereo
Wenqi Yang, Guanying Chen, Chaofeng Chen, Zhenfang Chen, Kwan-Yee Kenneth Wong |
ECCV (1) | 5 |
| 2022 | A Unified Framework for Masked and Mask-Free Face Recognition Via Feature RectificationabstractFace recognition under ideal conditions is now considered a well-solved problem with advances in deep learning. Recognizing faces under occlusion, however, still remains a challenge. Existing techniques often fail to recognize faces with both the mouth and nose covered by a mask, which is now very common under the COVID-19 pandemic. Common approaches to tackle this problem include 1) discarding information from the masked regions during recognition and 2) restoring the masked regions before recognition. Very few works considered the consistency between features extracted from masked faces and from their mask-free counterparts. This resulted in models trained for recognizing masked faces of-ten showing degraded performance on mask-free faces. In this paper, we propose a unified framework, named Face Feature Rectification Network (FFR-Net), for recognizing both masked and mask-free faces alike. We introduce rectification blocks to rectify features extracted by a state-of-the-art recognition model, in both spatial and channel dimensions, to minimize the distance between a masked face and its mask-free counterpart in the rectified feature space. Experiments show that our unified framework can learn a rectified feature space for recognizing both masked and mask-free faces effectively, achieving state-of-the-art results. Project code: https://github.com/haoosz/FFR-Net Shaozhe Hao, Chaofeng Chen, Zhenfang Chen, Kwan-Yee Kenneth Wong |
ICIP | 4 |
| 2022 | S3-NeRF: Neural Reflectance Field from Shading and Shadow under a Single ViewpointabstractIn this paper, we address the "dual problem" of multi-view scene reconstruction in which we utilize single-view images captured under different point lights to learn a neural scene representation. Different from existing single-view methods which can only recover a 2.5D scene representation (i.e., a normal / depth map for the visible surface), our method learns a neural reflectance field to represent the 3D geometry and BRDFs of a scene. Instead of relying on multi-view photo-consistency, our method exploits two information-rich monocular cues, namely shading and shadow, to infer scene geometry. Experiments on multiple challenging datasets show that our method is capable of recovering 3D geometry, including both visible and invisible parts, of a scene from single-view images. Thanks to the neural reflectance field representation, our method is robust to depth discontinuities. It supports applications like novel-view synthesis and relighting. Our code and model can be found at https://ywq.github.io/s3nerf. Wenqi Yang, Guanying Chen, Chaofeng Chen, Zhenfang Chen, Kwan-Yee Kenneth Wong |
NeurIPS | 5 |
| 2022 | Deep Photometric Stereo for Non-Lambertian SurfacesabstractThis paper addresses the problem of photometric stereo, in both calibrated and uncalibrated scenarios, for non-Lambertian surfaces based on deep learning. We first introduce a fully convolutional deep network for calibrated photometric stereo, which we call PS-FCN. Unlike traditional approaches that adopt simplified reflectance models to make the problem tractable, our method directly learns the mapping from reflectance observations to surface normal, and is able to handle surfaces with general and unknown isotropic reflectance. At test time, PS-FCN takes an arbitrary number of images and their associated light directions as input and predicts a surface normal map of the scene in a fast feed-forward pass. To deal with the uncalibrated scenario where light directions are unknown, we introduce a new convolutional network, named LCNet, to estimate light directions from input images. The estimated light directions and the input images are then fed to PS-FCN to determine the surface normals. Our method does not require a pre-defined set of light directions and can handle multiple images in an order-agnostic manner. Thorough evaluation of our approach on both synthetic and real datasets shows that it outperforms state-of-the-art methods in both calibrated and uncalibrated scenarios. Guanying Chen, Kai Han 0001, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee Kenneth Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Progressive Semantic-Aware Style Transformation for Blind Face RestorationabstractFace restoration is important in face image processing, and has been widely studied in recent years. However, previous works often fail to generate plausible high quality (HQ) results for real-world low quality (LQ) face images. In this paper, we propose a new progressive semantic-aware style transformation framework, named PSFR-GAN, for face restoration. Specifically, instead of using an encoder-decoder framework as previous methods, we formulate the restoration of LQ face images as a multi-scale progressive restoration procedure through semantic-aware style transformation. Given a pair of LQ face image and its corresponding parsing map, we first generate a multi-scale pyramid of the inputs, and then progressively modulate different scale features from coarse-to-fine in a semantic-aware style transfer way. Compared with previous networks, the proposed PSFR-GAN makes full use of the semantic (parsing maps) and pixel (LQ images) space information from different scales of input pairs. In addition, we further introduce a semantic aware style loss which calculates the feature style loss for each semantic region individually to improve the details of face textures. Finally, we pretrain a face parsing network which can generate decent parsing maps from real-world LQ face images. Experiment results show that our model trained with synthetic data can not only produce more realistic high-resolution results for synthetic LQ inputs but also generalize better to natural LQ face images compared with state-of-the-art methods. Chaofeng Chen, Xiaoming Li 0002, Lingbo Yang, Xianhui Lin, Lei Zhang 0006, Kwan-Yee Kenneth Wong |
CVPR | 6 |
| 2021 | HDR Video Reconstruction: A Coarse-to-fine Network and A Real-world Benchmark DatasetabstractHigh dynamic range (HDR) video reconstruction from sequences captured with alternating exposures is a very challenging problem. Existing methods often align low dynamic range (LDR) input sequence in the image space using optical flow, and then merge the aligned images to produce HDR output. However, accurate alignment and fusion in the image space are difficult due to the missing details in the over-exposed regions and noise in the under-exposed regions, resulting in unpleasing ghosting artifacts. To enable more accurate alignment and HDR fusion, we introduce a coarse-to-fine deep learning framework for HDR video reconstruction. Firstly, we perform coarse alignment and pixel blending in the image space to estimate the coarse HDR video. Secondly, we conduct more sophisticated alignment and temporal fusion in the feature space of the coarse HDR video to produce better reconstruction. Considering the fact that there is no publicly available dataset for quantitative and comprehensive evaluation of HDR video reconstruction methods, we collect such a benchmark dataset, which contains 97 sequences of static scenes and 184 testing pairs of dynamic scenes. Extensive experiments show that our method outperforms previous state-of-the-art methods. Our code and dataset can be found at https://guanyingc.github.io/DeepHDRVideo. Guanying Chen, Chaofeng Chen, Shi Guo, Zhetong Liang, Kwan-Yee Kenneth Wong, Lei Zhang 0006 |
ICCV | 5 |
| 2021 | Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning
Zhenfang Chen, Jiayuan Mao, Jiajun Wu 0001, Kwan-Yee Kenneth Wong, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 4 |
| 2021 | Learning Spatial Attention for Face Super-ResolutionabstractGeneral image super-resolution techniques have difficulties in recovering detailed face structures when applying to low resolution face images. Recent deep learning based methods tailored for face images have achieved improved performance by jointly trained with additional task such as face parsing and landmark prediction. However, multi-task learning requires extra manually labeled data. Besides, most of the existing works can only generate relatively low resolution face images (e.g., 128×128 ), and their applications are therefore limited. In this paper, we introduce a novel SPatial Attention Residual Network (SPARNet) built on our newly proposed Face Attention Units (FAUs) for face super-resolution. Specifically, we introduce a spatial attention mechanism to the vanilla residual blocks. This enables the convolutional layers to adaptively bootstrap features related to the key face structures and pay less attention to those less feature-rich regions. This makes the training more effective and efficient as the key face structures only account for a very small portion of the face image. Visualization of the attention maps shows that our spatial attention network can capture the key face structures well even for very low resolution faces (e.g., 16×16 ). Quantitative comparisons on various kinds of metrics (including PSNR, SSIM, identity similarity, and landmark detection) demonstrate the superiority of our method over current state-of-the-arts. We further extend SPARNet with multi-scale discriminators, named as SPARNetHD, to produce high resolution results (i.e., 512×512 ). We show that SPARNetHD trained with synthetic data can not only produce high quality and high resolution outputs for synthetically degraded face images, but also show good generalization ability to real world low quality face images. Codes are available at https://github.com/chaofengc/Face-SPARNet. Chaofeng Chen, Dihong Gong, Hao Wang 0050, Zhifeng Li 0001, Kwan-Yee Kenneth Wong |
IEEE Trans. Image Process. | 5 |
| 2021 | Fixed Viewpoint Mirror Surface Reconstruction Under an Uncalibrated CameraabstractThis paper addresses the problem of mirror surface reconstruction, and proposes a solution based on observing the reflections of a moving reference plane on the mirror surface. Unlike previous approaches which require tedious calibration, our method can recover the camera intrinsics, the poses of the reference plane, as well as the mirror surface from the observed reflections of the reference plane under at least three unknown distinct poses. We first show that the 3D poses of the reference plane can be estimated from the reflection correspondences established between the images and the reference plane. We then form a bunch of 3D lines from the reflection correspondences, and derive an analytical solution to recover the line projection matrix. We transform the line projection matrix to its equivalent camera projection matrix, and propose a cross-ratio based formulation to optimize the camera projection matrix by minimizing reprojection errors. The mirror surface is then reconstructed based on the optimized cross-ratio constraint. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and accuracy of our method. Kai Han 0001, Miaomiao Liu 0001, Dirk Schnieders, Kwan-Yee Kenneth Wong |
IEEE Trans. Image Process. | 4 |
| 2020 | Cops-Ref: A New Dataset and Task on Compositional Referring Expression ComprehensionabstractReferring expression comprehension (REF) aims at identifying a particular object in a scene by a natural language expression. It requires joint reasoning over the textual and visual domains to solve the problem. Some popular referring expression datasets, however, fail to provide an ideal test bed for evaluating the reasoning ability of the models, mainly because 1) their expressions typically describe only some simple distinctive properties of the object and 2) their images contain limited distracting information. To bridge the gap, we propose a new dataset for visual reasoning in context of referring expression comprehension with two main features. First, we design a novel expression engine rendering various reasoning logics that can be flexibly combined with rich visual properties to generate expressions with varying compositionality. Second, to better exploit the full reasoning chain embodied in an expression, we propose a new test setting by adding additional distracting images containing objects sharing similar properties with the referent, thus minimising the success rate of reasoning-free cross-domain alignment. We evaluate several state-of-the-art REF models, but find none of them can achieve promising performance. A proposed modular hard mining strategy performs the best but still leaves substantial room for improvement. Zhenfang Chen, Peng Wang 0023, Lin Ma 0002, Kwan-Yee Kenneth Wong, Qi Wu 0001 |
CVPR | 4 |
| 2020 | What Is Learned in Deep Uncalibrated Photometric Stereo?
Guanying Chen, Michael Waechter, Boxin Shi, Kwan-Yee Kenneth Wong, Yasuyuki Matsushita |
ECCV (14) | 4 |
| 2019 | Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in VideoabstractIn this paper, we address a novel task, namely weakly-supervised spatio-temporally grounding natural sentence in video.Specifically, given a natural sentence and a video, we localize a spatio-temporal tube in the video that semantically corresponds to the given sentence, with no reliance on any spatio-temporal annotations during training.First, a set of spatiotemporal tubes, referred to as instances, are extracted from the video.We then encode these instances and the sentence using our proposed attentive interactor which can exploit their fine-grained relationships to characterize their matching behaviors.Besides a ranking loss, a novel diversity loss is introduced to train the proposed attentive interactor to strengthen the matching behaviors of reliable instance-sentence pairs and penalize the unreliable ones.Moreover, we also contribute a dataset, called VID-sentence, based on the Im-ageNet video object detection dataset, to serve as a benchmark for our task.Extensive experimental results demonstrate the superiority of our model over the baseline approaches.Our code and the constructed VID-sentence dataset are available at: https://github.com/ JeffCHEN2017/WSSTG.git. Zhenfang Chen, Lin Ma 0002, Wenhan Luo, Kwan-Yee Kenneth Wong |
ACL (1) | 4 |
| 2019 | Self-Calibrating Deep Photometric Stereo NetworksabstractThis paper proposes an uncalibrated photometric stereo method for non-Lambertian scenes based on deep learning. Unlike previous approaches that heavily rely on assumptions of specific reflectances and light source distributions, our method is able to determine both shape and light directions of a scene with unknown arbitrary reflectances observed under unknown varying light directions. To achieve this goal, we propose a two-stage deep learning architecture, called SDPS-Net, which can effectively take advantage of intermediate supervision, resulting in reduced learning difficulty compared to a single-stage model. Experiments on both synthetic and real datasets show that our proposed approach significantly outperforms previous uncalibrated photometric stereo methods. Guanying Chen, Kai Han 0001, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee Kenneth Wong |
CVPR | 5 |
| 2019 | Learning Local Similarity with Spatial Relations for Object RetrievalabstractMany state-of-the-art object retrieval algorithms aggregate activations of convolutional neural networks into a holistic compact feature, and utilize global similarity for an efficient nearest neighbor search. However, holistic features are often insufficient for representing small objects of interest in gallery images, and global similarity drops most of the spatial relations in the images. In this paper, we propose an end-to-end local similarity learning framework to tackle these problems. By applying a correlation layer to the locally aggregated features, we compute a local similarity that can not only handle small objects, but also capture spatial relations between the query and gallery images. We further reduce the memory and storage footprints of our framework by quantizing local features. Our model can be trained using only synthetic data, and achieve competitive performance. Extensive experiments on challenging benchmarks demonstrate that our local similarity learning framework outperforms previous global similarity based methods. Zhenfang Chen, Zhanghui Kuang, Wayne Zhang 0001, Kwan-Yee Kenneth Wong |
ACM Multimedia | 4 |
| 2019 | Learning Transparent Object Matting
Guanying Chen, Kai Han 0001, Kwan-Yee Kenneth Wong |
Int. J. Comput. Vis. | 3 |
| 2018 | Char-Net: A Character-Aware Neural Network for Distorted Scene Text RecognitionabstractIn this paper, we present a Character-Aware Neural Network (Char-Net) for recognizing distorted scene text. Our Char-Net is composed of a word-level encoder, a character-level encoder, and a LSTM-based decoder. Unlike previous work which employed a global spatial transformer network to rectify the entire distorted text image, we take an approach of detecting and rectifying individual characters. To this end, we introduce a novel hierarchical attention mechanism (HAM) which consists of a recurrent RoIWarp layer and a character-level attention layer. The recurrent RoIWarp layer sequentially extracts a feature region corresponding to a character from the feature map produced by the word-level encoder, and feeds it to the character-level encoder which removes the distortion of the character through a simple spatial transformer and further encodes the character region. The character-level attention layer then attends to the most relevant features of the feature map produced by the character-level encoder and composes a context vector, which is finally fed to the LSTM-based decoder for decoding. This approach of adopting a simple local transformation to model the distortion of individual characters not only results in an improved efficiency, but can also handle different types of distortion that are hard, if not impossible, to be modelled by a single global transformation. Experiments have been conducted on six public benchmark datasets. Our results show that Char-Net can achieve state-of-the-art performance on all the benchmarks, especially on the IC-IST which contains scene text with large distortion. Code will be made available. Wei Liu 0091, Chaofeng Chen, Kwan-Yee Kenneth Wong |
AAAI | 3 |
| 2018 | Semi-supervised Learning for Face Sketch Synthesis in the Wild
Chaofeng Chen, Wei Liu 0091, Xiao Tan 0001, Kwan-Yee Kenneth Wong |
ACCV (1) | 4 |
| 2018 | SAFE: Scale Aware Feature Encoder for Scene Text Recognition
Wei Liu 0091, Chaofeng Chen, Kwan-Yee Kenneth Wong |
ACCV (2) | 3 |
| 2018 | TOM-Net: Learning Transparent Object Matting From a Single ImageabstractThis paper addresses the problem of transparent object matting. Existing image matting approaches for transparent objects often require tedious capturing procedures and long processing time, which limit their practical use. In this paper, we first formulate transparent object matting as a refractive flow estimation problem. We then propose a deep learning framework, called TOM-Net, for learning the refractive flow. Our framework comprises two parts, namely a multi-scale encoder-decoder network for producing a coarse prediction, and a residual network for refinement. At test time, TOM-Net takes a single image as input, and outputs a matte (consisting of an object mask, an attenuation mask and a refractive flow field) in a fast feed-forward pass. As no off-the-shelf dataset is available for transparent object matting, we create a large-scale synthetic dataset consisting of 178K images of transparent objects rendered in front of images sampled from the Microsoft COCO dataset. We also collect a real dataset consisting of 876 samples using 14 transparent objects and 60 background images. Promising experimental results have been achieved on both synthetic and real data, which clearly demonstrate the effectiveness of our approach. Guanying Chen, Kai Han 0001, Kwan-Yee Kenneth Wong |
CVPR | 3 |
| 2018 | PS-FCN: A Flexible Learning Framework for Photometric Stereo
Guanying Chen, Kai Han 0001, Kwan-Yee Kenneth Wong |
ECCV (9) | 3 |
| 2018 | Face Sketch Synthesis with Style Transfer Using Pyramid Column FeatureabstractIn this paper, we propose a novel framework based on deep neural networks for face sketch synthesis from a photo. Imitating the process of how artists draw sketches, our framework synthesizes face sketches in a cascaded manner. A content image is first generated that outlines the shape of the face and the key facial features. Textures and shadings are then added to enrich the details of the sketch. We utilize a fully convolutional neural network (FCNN) to create the content image, and propose a style transfer approach to introduce textures and shadings based on a newly proposed pyramid column feature. We demonstrate that our style transfer approach based on the pyramid column feature can not only preserve more sketch details than the common style transfer method, but also surpasses traditional patch based methods. Quantitative and qualitative evaluations suggest that our framework outperforms other state-of-the-arts methods, and can also generalize well to different test images. Chaofeng Chen, Xiao Tan 0001, Kwan-Yee Kenneth Wong |
WACV | 3 |
| 2018 | Dense Reconstruction of Transparent Objects by Altering Incident Light Paths Through Refraction
Kai Han 0001, Kwan-Yee Kenneth Wong, Miaomiao Liu 0001 |
Int. J. Comput. Vis. | 2 |
| 2017 | SCNet: Learning Semantic CorrespondenceabstractThis paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearance only. We propose instead a convolutional neural network architecture, called SCNet, for learning a geometrically plausible model for semantic correspondence. SCNet uses region proposals as matching primitives, and explicitly incorporates geometric consistency in its loss function. It is trained on image pairs obtained from the PASCAL VOC 2007 keypoint dataset, and a comparative evaluation on several standard benchmarks demonstrates that the proposed approach substantially outperforms both recent deep learning architectures and previous methods based on hand-crafted features. Kai Han 0001, Rafael S. Rezende, Bumsub Ham, Kwan-Yee Kenneth Wong, Minsu Cho, Cordelia Schmid, Jean Ponce |
ICCV | 4 |
| 2016 | Single View 3D Reconstruction under an Uncalibrated Camera and an Unknown Mirror SphereabstractIn this paper, we develop a novel self-calibration method for single view 3D reconstruction using a mirror sphere. Unlike other mirror sphere based reconstruction methods, our method needs neither the intrinsic parameters of the camera, nor the position and radius of the sphere be known. Based on eigen decomposition of the matrix representing the conic image of the sphere and enforcing a repeated eignvalue constraint, we derive an analytical solution for recovering the focal length of the camera given its principal point. We then introduce a robust algorithm for estimating both the principal point and the focal length of the camera by minimizing the differences between focal lengths estimated from multiple images of the sphere. We also present a novel approach for estimating both the principal point and focal length of the camera in the case of just one single image of the sphere. With the estimated camera intrinsic parameters, the position(s) of the sphere can be readily retrieved from the eigen decomposition(s) and a scaled 3D reconstruction follows. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and accuracy of our approach. Kai Han 0001, Kwan-Yee Kenneth Wong, Xiao Tan 0001 |
3DV | 2 |
| 2016 | STAR-Net: A SpaTial Attention Residue Network for Scene Text Recognition
Wei Liu 0091, Chaofeng Chen, Kwan-Yee Kenneth Wong, Zhizhong Su, Junyu Han |
BMVC | 3 |
| 2016 | Mirror Surface Reconstruction under an Uncalibrated CameraabstractThis paper addresses the problem of mirror surface reconstruction, and a solution based on observing the reflections of a moving reference plane on the mirror surface is proposed. Unlike previous approaches which require tedious work to calibrate the camera, our method can recover both the camera intrinsics and extrinsics together with the mirror surface from reflections of the reference plane under at least three unknown distinct poses. Our previous work has demonstrated that 3D poses of the reference plane can be registered in a common coordinate system using reflection correspondences established across images. This leads to a bunch of registered 3D lines formed from the reflection correspondences. Given these lines, we first derive an analytical solution to recover the camera projection matrix through estimating the line projection matrix. We then optimize the camera projection matrix by minimizing reprojection errors computed based on a cross-ratio formulation. The mirror surface is finally reconstructed based on the optimized cross-ratio constraint. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and accuracy of our method. Kai Han 0001, Kwan-Yee Kenneth Wong, Dirk Schnieders, Miaomiao Liu 0001 |
CVPR | 2 |
| 2016 | A two-stage detector for hand detection in ego-centric videosabstractWe propose a two-stage detector that can not only detect and localize hands, but also provide fine-detailed information in the bounding box of hand in an efficient fashion. In the first stage, hand bounding box proposals are generated from a pixel-level hand probability map. Next, each hand proposal is evaluated by a Multi-task Convolutional Neural Network to filter out false positives and obtain fine shape and landmark information. Through experiments, we demonstrate that our method is efficient and robust to detect hands with their shape and landmark information, and our system can also be flexibly combined with other detection methods to handle a new scene. Further experiment shows that our Multi-task CNN can also be extended to hand gesture classification with a large performance increase. Wei Liu 0091, Xuhui Jia, Kwan-Yee Kenneth Wong |
WACV | 4 |
| 2016 | Guided image completion by confidence propagation
Xiao Tan 0001, Changming Sun, Kwan-Yee Kenneth Wong, Tuan D. Pham |
Pattern Recognit. | 3 |
| 2015 | A fixed viewpoint approach for dense reconstruction of transparent objectsabstractThis paper addresses the problem of reconstructing the surface shape of transparent objects. The difficulty of this problem originates from the viewpoint dependent appearance of a transparent object, which quickly makes reconstruction methods tailored for diffuse surfaces fail disgracefully. In this paper, we develop a fixed viewpoint approach for dense surface reconstruction of transparent objects based on refraction of light. We introduce a simple setup that allows us alter the incident light paths before light rays enter the object, and develop a method for recovering the object surface based on reconstructing and triangulating such incident light paths. Our proposed approach does not need to model the complex interactions of light as it travels through the object, neither does it assume any parametric form for the shape of the object nor the exact number of refractions and reflections taken place along the light paths. It can therefore handle transparent objects with a complex shape and structure, with unknown and even inhomogeneous refractive index. Experimental results on both synthetic and real data are presented which demonstrate the feasibility and accuracy of our proposed approach. Kai Han 0001, Kwan-Yee Kenneth Wong, Miaomiao Liu 0001 |
CVPR | 2 |
| 2015 | Structured forests for pixel-level hand detection and hand part labelling
Xuhui Jia, Kwan-Yee Kenneth Wong |
Comput. Vis. Image Underst. | 3 |
| 2015 | Relatively-Paired Space Analysis: Learning a Latent Common Space From Relatively-Paired Observations
Zhanghui Kuang, Kwan-Yee Kenneth Wong |
Int. J. Comput. Vis. | 2 |
| 2014 | Pixel-Level Hand Detection with Shape-Aware Structured Forests
Xuhui Jia, Kwan-Yee Kenneth Wong |
ACCV (4) | 3 |
| 2013 | Relatively-Paired Space AnalysisabstractDiscovering a latent common space between different modalities plays an important role in cross-modality pattern recognition. Existing techniques often require absolutely-paired observations as training data, and are incapable of capturing more general seman-tic relationships between cross-modality observations. This greatly limits their appli-cations. In this paper, we propose a general framework for learning a latent common space from relatively-paired observations (i.e., two observations from different modali-ties are more-likely-paired than another two). Relative-pairing information is encoded using relative proximities of observations in the latent common space. By building a discriminative model and maximizing a distance margin, a projection function that maps observations into the latent common space is learned for each modality. Cross-modality pattern recognition can then be carried out in the latent common space. To evaluate its performance, the proposed framework has been applied to cross-pose face recognition and feature fusion. Experimental results demonstrate that the proposed framework out-performs other state-of-the-art approaches. 1 Zhanghui Kuang, Kwan-Yee Kenneth Wong |
BMVC | 2 |
| 2013 | Camera and light calibration from reflections on a sphere
Dirk Schnieders, Kwan-Yee Kenneth Wong |
Comput. Vis. Image Underst. | 2 |
| 2013 | Depth from Refraction Using a Transparent Medium with Unknown Pose and Refractive IndexabstractIn this paper, we introduce a novel method for depth acquisition based on refraction of light. A scene is captured directly by a camera and by placing a transparent medium between the scene and the camera. A depth map of the scene is then recovered from the displacements of scene points in the images. Unlike other existing depth from refraction methods, our method does not require prior knowledge of the pose and refractive index of the transparent medium, but instead can recover them directly from the input images. By analyzing the displacements of corresponding scene points in the images, we derive closed form solutions for recovering the pose of the transparent medium and develop an iterative method for estimating the refractive index of the medium. Experimental results on both synthetic and real-world data are presented, which demonstrate the effectiveness of the proposed method. Zhihu Chen, Kwan-Yee Kenneth Wong, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 2 |
| 2012 | Adaptive Background Defogging with Foreground Decremental Preconditioned Conjugate Gradient
Jacky S.-C. Yuk, Kwan-Yee Kenneth Wong |
ACCV (4) | 2 |
| 2012 | Self-calibration and Motion Recovery from Silhouettes with Two Mirrors
Hui Zhang 0062, Ling Shao 0001, Kwan-Yee Kenneth Wong |
ACCV (4) | 3 |
| 2012 | Learning image-specific parameters for interactive segmentationabstractIn this paper, we present a novel interactive image segmentation technique that automatically learns segmentation parameters tailored for each and every image. Unlike existing work, our method does not require any offline parameter tuning or training stage, and is capable of determining image-specific parameters according to some simple user interactions with the target image. We formulate the segmentation problem as an inference of a conditional random field (CRF) over a segmentation mask and the target image, and parametrize this CRF by different weights (e.g., color, texture and smoothing). The weight parameters are learned via an energy margin maximization, which is solved using a constraint approximation scheme and the cutting plane method. Experimental results show that our method, by learning image-specific parameters automatically, outperforms other state-of-the-art interactive image segmentation techniques. Zhanghui Kuang, Dirk Schnieders, Hao Zhou 0010, Kwan-Yee Kenneth Wong, Yizhou Yu |
CVPR | 4 |
| 2012 | Markov Weight Fields for face sketch synthesisabstractGreat progress has been made in face sketch synthesis in recent years. State-of-the-art methods commonly apply a Markov Random Fields (MRF) model to select local sketch patches from a set of training data. Such methods, however, have two major drawbacks. Firstly, the MRF model used cannot synthesize new sketch patches. Secondly, the optimization problem in solving the MRF is NP-hard. In this paper, we propose a novel Markov Weight Fields (MWF) model that is capable of synthesizing new sketch patches. We formulate our model into a convex quadratic programming (QP) problem to which the optimal solution is guaranteed. Based on the Markov property of our model, we further propose a cascade decomposition method (CDM) for solving such a large scale QP problem efficiently. Experimental results on the CUHK face sketch database and celebrity photos show that our model outperforms the common MRF model used in other state-of-the-art methods. Hao Zhou 0010, Zhanghui Kuang, Kwan-Yee Kenneth Wong |
CVPR | 3 |
| 2012 | Single-frame hand gesture recognition using color and depth kernel descriptors
Kwan-Yee Kenneth Wong |
ICPR | 2 |
| 2011 | An efficient pattern-less background modeling based on scale invariant local statesabstractA robust and efficient background modeling algorithm is crucial to the success of most of the intelligent video surveillance systems. Compared with intensity-based approaches, texture-based background modeling approaches have shown to be more robust against dynamic backgrounds and illumination changes, which are common in real life videos. However, many of the existing texture-based methods are too computationally expensive, which renders them useless in real-time applications. In this paper, a novel efficient texture-based background modeling algorithm is presented. Scale invariant local states (SILS) are introduced as pixel features for modeling a background pixel, and a pattern-less probabilistic measurement (PLPM) is derived to estimate the probability of a pixel being background from its SILS. An adaptive background modeling framework is also introduced for learning and representing a multi-modal background model. Experimental results show that the proposed method can run nearly 3 times faster than existing state-of-the-art texture-based method, without sacrificing the output quality. This allows more time for a real-time surveillance system to carry out other computationally intensive analysis on the detected foreground objects. Jacky S.-C. Yuk, Kwan-Yee Kenneth Wong |
AVSS | 2 |
| 2011 | Self-calibrating depth from refractionabstractIn this paper, we introduce a novel method for depth acquisition based on refraction of light. A scene is captured twice by a fixed perspective camera, with the first image captured directly by the camera and the second by placing a transparent medium between the scene and the camera. A depth map of the scene is then recovered from the displacements of scene points in the images. Unlike other existing depth from refraction methods, our method does not require the knowledge of the pose and refractive index of the transparent medium, but can recover them directly from the input images. We hence call our method self-calibrating depth from refraction. Experimental results on both synthetic and real-world data are presented, which demonstrate the effectiveness of the proposed method. Zhihu Chen, Kwan-Yee Kenneth Wong, Yasuyuki Matsushita, Miaomiao Liu 0001 |
ICCV | 2 |
| 2011 | Pose estimation from reflections for specular surface recoveryabstractThis paper addresses the problem of estimating the poses of a reference plane in specular shape recovery. Unlike existing methods which require an extra mirror or an extra reference plane and camera, our proposed method recovers the poses of the reference plane directly from its reflections on the specular surface. By establishing reflection correspondences on the reference plane in three distinct poses, our method estimates the poses of the reference plane in two steps. First, by applying a colinearity constraint to the reflection correspondences, a simple closed-form solution is derived for recovering the poses of the reference plane relative to its initial pose. Second, by applying a ray incidence constraint to the incident rays formed by the reflection correspondences and the visual rays cast from the image, a closed-form solution is derived for recovering the poses of the reference plane relative to the camera. The shape of the specular surface then follows. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and accuracy of our proposed method. Miaomiao Liu 0001, Kwan-Yee Kenneth Wong, Zhenwen Dai, Zhihu Chen |
ICCV | 2 |
| 2011 | Accurate Foreground Segmentation without Pre-learningabstractForeground segmentation has been widely used in many computer vision applications. However, most of the existing methods rely on a pre-learned motion or background model, which will increase the burden of users. In this paper, we present an automatic algorithm without pre-learning for segmenting foreground from background based on the fusion of motion, color and contrast information. Motion information is enhanced by a novel method called support edges diffusion (SED), which is built upon a key observation that edges of the difference image of two adjacent frames only appear in moving regions in most of the cases. Contrasts in background are attenuated while those in foreground are enhanced using gradient of the previous frame and that of the temporal difference. Experiments on many video sequences demonstrate the effectiveness and accuracy of the proposed algorithm. The segmentation results are comparable to those obtained by other state-of-the-art methods that depend on a pre-learned background or a stereo setup. Zhanghui Kuang, Hao Zhou 0010, Kwan-Yee Kenneth Wong |
ICIG | 3 |
| 2011 | Single-view reconstruction from an unknown spherical mirrorabstractThis paper introduces a novel method for 3D reconstruction from spherical mirrors in a single view. Traditional reconstruction algorithms either assume both intrinsic and extrinsic parameters of the cameras being known precisely and rely on multiple view information to recover 3D objects, or use reflections of the 3D objects on spherical mirrors with known radii in a single calibrated view. It will be shown in this paper that 3D objects can be recovered up to a scale using an unknown spherical mirror in a single view. Experimental results on both synthetic and real data show the practicality of the proposed reconstruction method. Zhihu Chen, Kwan-Yee Kenneth Wong, Miaomiao Liu 0001, Dirk Schnieders |
ICIP | 2 |
| 2011 | A Stratified Approach for Camera Calibration Using SpheresabstractThis paper proposes a stratified approach for camera calibration using spheres. Previous works have exploited epipolar tangents to locate frontier points on spheres for estimating the epipolar geometry. It is shown in this paper that other than the frontier points, two additional point features can be obtained by considering the bitangent envelopes of a pair of spheres. A simple method for locating the images of such point features and the sphere centers is presented. An algorithm for recovering the fundamental matrix in a plane plus parallax representation using these recovered image points and the epipolar tangents from three spheres is developed. A new formulation of the absolute dual quadric as a cone tangent to a dual sphere with the plane at infinity being its vertex is derived. This allows the recovery of the absolute dual quadric, which is used to upgrade the weak calibration to a full calibration. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and the high precision achieved by our proposed algorithm. Kwan-Yee Kenneth Wong, Guoqiang Zhang 0003, Zhihu Chen |
IEEE Trans. Image Process. | 1 |
| 2010 | Specular Surface Recovery from Reflections of a Planar Pattern Undergoing an Unknown Pure Translation
Miaomiao Liu 0001, Kwan-Yee Kenneth Wong, Zhenwen Dai, Zhihu Chen |
ACCV (2) | 2 |
| 2010 | A Multi-level Supporting Scheme for Face Recognition under Partial Occlusions and Disguise
Jacky S.-C. Yuk, Kwan-Yee Kenneth Wong, Ronald H. Y. Chung |
ACCV (4) | 2 |
| 2010 | Reconstruction of display and eyes from a single imageabstractThis paper introduces a novel method for reconstructing human eyes and visual display from reflections on the cornea. This problem is difficult because the camera is not directly facing the display, but instead captures the eyes of a person in front of the display. Reconstruction of eyes and display is useful for point-of-gaze estimation, which can be approximated from the 3D positions of the iris and display. It is shown that iris boundaries (limbus) and display reflections in a single intrinsically calibrated image provide enough information for such an estimation. The proposed method assumes a simplified geometric eyeball model with certain anatomical constants which are used to reconstruct the eye. A noise performance analysis shows the sensitivity of the proposed method to imperfect data. Experiments on various subjects show that it is possible to determine the approximate area of gaze on a display from a single image. Dirk Schnieders, Xingdou Fu, Kwan-Yee Kenneth Wong |
CVPR | 3 |
| 2010 | Local multiple orientations estimation using k-medoidsabstractEstimation of local multiple orientations plays an important role in many image processing and computer vision tasks. It has been shown that the detection of orientations in an image patch corresponds to fitting multiple axes to its Fourier transform. In this paper, k-medoids are introduced to detect local multiple orientations in the Fourier domain. Medoids are related to a well-known matrix eigenvector problem. A hierarchical schema with eigensystem and energy distribution analysis is employed to determine the number of orientations in an image patch. The proposed approach detects two types of orientation structure (ridges and edges) without difference. Experimental results on synthetic and real images show that the proposed method can detect multiple orientations with high accuracy and is robust against noise. Zhanghui Kuang, Guodong Pan, Kwan-Yee Kenneth Wong |
ICIP | 3 |
| 2010 | 3D reconstruction using silhouettes from unordered viewpoints
Chen Liang 0004, Kwan-Yee Kenneth Wong |
Image Vis. Comput. | 2 |
| 2009 | Polygonal Light Source Estimation
Dirk Schnieders, Kwan-Yee Kenneth Wong, Zhenwen Dai |
ACCV (3) | 2 |
| 2009 | Super-Resolution of Faces Using Texture Mapping on a Generic 3D ModelabstractThis paper proposes a novel face texture mapping framework to transform faces with different poses into a unique texture map. Under this framework, texture mapping can be realized by utilizing a generic 3D face model, standard Haar-like feature based detector, active appearance model and pose estimation algorithm. By this texture map, correspondence of every pixel at the face across multiple distinct input images can then be established, which enables super-resolution algorithms to be applied directly on registered texture map to render high resolution faces. This paper details the proposed framework, and illustrates how the proposed super-resolution algorithm works with the help of weighted average and median filters. Convincing experimental results are also presented to validate the effectiveness of the proposed framework and super-resolution algorithm. X. C. He, Jacky S.-C. Yuk, Kam-Pui Chow, Kwan-Yee Kenneth Wong, Ronald H. Y. Chung |
ICIG | 4 |
| 2009 | Self-Calibration of Turntable Sequences from SilhouettesabstractThis paper addresses the problem of recovering both the intrinsic and extrinsic parameters of a camera from the silhouettes of an object in a turntable sequence. Previous silhouette-based approaches have exploited correspondences induced by epipolar tangents to estimate the image invariants under turntable motion and achieved a weak calibration of the cameras. It is known that the fundamental matrix relating any two views in a turntable sequence can be expressed explicitly in terms of the image invariants, the rotation angle, and a fixed scalar. It will be shown that the imaged circular points for the turntable plane can also be formulated in terms of the same image invariants and fixed scalar. This allows the imaged circular points to be recovered directly from the estimated image invariants, and provide constraints for the estimation of the imaged absolute conic. The camera calibration matrix can thus be recovered. A robust method for estimating the fixed scalar from image triplets is introduced, and a method for recovering the rotation angles using the estimated imaged circular points and epipoles is presented. Using the estimated camera intrinsics and extrinsics, a Euclidean reconstruction can be obtained. Experimental results on real data sequences are presented, which demonstrate the high precision achieved by the proposed method. Hui Zhang 0062, Kwan-Yee Kenneth Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Motion Recovery for Uncalibrated Turntable Sequences Using Silhouettes and a Single Point
Hui Zhang 0062, Ling Shao 0001, Kwan-Yee Kenneth Wong |
ACIVS | 3 |
| 2008 | Using Illumination Estimated from Silhouettes to Carve Surface Details on Visual HullabstractThis paper deals with the problems of scene illumination estimation and shape recovery from an image sequence of a smooth textureless object. A novel method that exploits the surface points estimated from the silhouettes for recovering the scene illumination is introduced. Those surface points are acquired by a dual space approach and filtered according to their rank errors. Selected surface points allow a direct closed-form solution of illumination. In the mesh evolution step, an algorithm for optimizing the visual hull mesh is developed. It evolves the mesh by iteratively estimating both the surface normal and depth that maximize the photometric consistency across the sequence. Compared with previous work which optimizes the mesh by estimating the surface normal only, the proposed method shows better convergence and can recover better surface details, especially when concavities are deep and sharp. 1 Shuda Li, Kwan-Yee Kenneth Wong, Dirk Schnieders |
BMVC | 2 |
| 2008 | Recovering Light Directions and Camera Poses from a Single Sphere
Kwan-Yee Kenneth Wong, Dirk Schnieders, Shuda Li |
ECCV (1) | 1 |
| 2008 | 1D Camera Geometry and Its Application to the Self-Calibration of Circular Motion SequencesabstractThis paper proposes a novel method for robustly recovering the camera geometry of an uncalibrated image sequence taken under circular motion. Under circular motion, all the camera centers lie on a circle and the mapping from the plane containing this circle to the horizon line observed in the image can be modelled as a 1D projection. A 2 x 2 homography is introduced in this paper to relate the projections of the camera centers in two 1D views. It is shown that the two imaged circular points of the motion plane and the rotation angle between the two views can be derived directly from such a homography. This way of recovering the imaged circular points and rotation angles is intrinsically a multiple view approach, as all the sequence geometry embedded in the epipoles is exploited in the estimation of the homography for each view pair. This results in a more robust method compared to those computing the rotation angles using adjacent views only. The proposed method has been applied to self-calibrate turntable sequences using either point features or silhouettes, and highly accurate results have been achieved. Kwan-Yee Kenneth Wong, Guoqiang Zhang 0003, Chen Liang 0004, Hui Zhang 0062 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Robust Recovery of Shapes with Unknown Topology from the Dual SpaceabstractIn this paper, we address the problem of reconstructing an object surface from silhouettes. Previous works by other authors have shown that, based on the principle of duality, surface points can be recovered, theoretically, as the dual to the tangent plane space of the object. In practice, however, the identification of tangent basis in the tangent plane space is not trivial given a set of discretely sampled data. This problem is further complicated by the existence of bi-tangents to the object surface. The key contribution of this paper is the introduction of epipolar parameterization in identifying a well-defined local tangent basis. This extends the applicability of existing dual space reconstruction methods to fairly complicated shapes, without making any explicit assumption on the object topology. We verify our approach with both synthetic and real-world data, and compare it both qualitatively and quantitatively with other popular reconstruction algorithms. Experimental results demonstrate that our proposed approach produces more accurate estimation, whilst maintaining reasonable robustness towards shapes with complex topologies. Chen Liang 0004, Kwan-Yee Kenneth Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Camera Calibration from Images of SpheresabstractThis paper introduces a novel approach for solving the problem of camera calibration from spheres. By exploiting the relationship between the dual images of spheres and the dual image of the absolute conic (IAC), it is shown that the common pole and polar with regard to the conic images of two spheres are also the pole and polar with regard to the IAC. This provides two constraints for estimating the IAC and, hence, allows a camera to be calibrated from an image of at least three spheres. Experimental results show the feasibility of the proposed approach. Hui Zhang 0062, Kwan-Yee Kenneth Wong, Guoqiang Zhang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Design and Analysis of Optimization Methods for Subdivision Surface FittingabstractWe present a complete framework for computing a subdivision surface to approximate unorganized point sample data, which is a separable nonlinear least squares problem. We study the convergence and stability of three geometrically-motivated optimization schemes and reveal their intrinsic relations with standard methods for constrained nonlinear optimization. A commonly-used method in graphics, called point distance minimization, is shown to use a variant of the gradient descent step and thus has only linear convergence. The second method, called tangent distance minimization, which is well-known in computer vision, is shown to use the Gauss-Newton step, and thus demonstrates near quadratic convergence for zero residual problems but may not converge otherwise. Finally, we show that an optimization scheme called squared distance minimization, recently proposed by Pottmann et al., can be derived from the Newton method. Hence, with proper regularization, tangent distance minimization and squared distance minimization are more efficient than point distance minimization. We also investigate the effects of two step size control methods -- Levenberg-Marquardt regularization and the Armijo rule -- on the convergence stability and efficiency of the above optimization schemes. Kin-Shing D. Cheng, Wenping Wang 0001, Hong Qin 0001, Kwan-Yee Kenneth Wong, Huaiping Yang, Yang Liu 0014 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2006 | Extracting Surface Representations from Rim Curves
Hai Chen, Kwan-Yee Kenneth Wong, Chen Liang 0004 |
ACCV (2) | 2 |
| 2006 | 1D Camera Geometry and Its Application to Circular Motion EstimationabstractThis paper describes a new and robust method for estimating circular motion geometry from an uncalibrated image sequence. Under circular motion, all the camera centers lie on a circle, and the mapping of the plane containing this circle to the horizon line in the image can be modelled as a 1D projection. A 2×2 homography is introduced in this paper to relate the projections of the camera centers in two 1D views. It is shown that the two imaged circular points and the rotation angle between the two views can be derived directly from the eigenvectors and eigenvalues of such a homography respectively. The proposed 1D geometry can be nicely applied to circular motion estimation using either point correspondences or silhouettes. The method introduced here is intrinsically a multiple view approach as all the sequence geometry embedded in the epipoles is exploited in the computation of the homography for a view pair. This results in a robust method which gives accurate estimated rotation angles and imaged circular points. Experimental results are presented to demonstrate the simplicity and applicability of the new method. Guoqiang Zhang 0003, Hui Zhang 0062, Kwan-Yee Kenneth Wong |
BMVC | 3 |
| 2006 | Motion Estimation from SpheresabstractThis paper addresses the problem of recovering epipolar geometry from spheres. Previous works have exploited epipolar tangencies induced by frontier points on the spheres for motion recovery. It will be shown in this paper that besides epipolar tangencies, N^2 point features can be extracted from the apparent contours of the N spheres whenN \gt 2. An algorithm for recovering the fundamental matrices from such point features and the epipolar tangencies from 3 or more spheres is developed, with the point features providing a homography over the view pairs and the epipolar tangencies determining the epipoles. In general, there will be two solutions to the locations of the epipoles. One of the solutions corresponds to the true camera configuration, while the other corresponds to a mirrored configuration. Several methods are proposed to select the right solution. Experiments on using 3 and 4 spheres demonstrate that our algorithm can be carried out easily and can achieve a high precision. Guoqiang Zhang 0003, Kwan-Yee Kenneth Wong |
CVPR (1) | 2 |
| 2005 | Auto-Calibration and Motion Recovery from Silhouettes for Turntable SequencesabstractThis paper addresses the problem of structure and motion from silhou-ettes for turntable sequences. Previous works have exploited corresponding points induced by epipolar tangencies to estimate the image invariants un-der turntable motion and recover the epipolar geometry. In these approaches, however, camera intrinsics are needed in order to obtain Euclidean motion and reconstruction. This paper proposes a novel approach to precisely esti-mate the image invariants and the rotation angles in the absence of the camera intrinsics, and to perform auto-calibration. By exploiting a special parame-terization of the epipoles, it is shown that the imaged circular points can be formulated in terms of the image invariants. A fixed scalar κ, introduced to account for the different scales in the homogeneous representations of the image invariants used in the parameterizations, is found crucial in both cali-bration and motion estimation. Given the image invariants, namely the hori-zon, the imaged rotation axis and its orthogonal vanishing point, this scalar can be determined from the epipoles in an image triplet. A robust method for estimating κ is proposed and the rotation angles can be recovered using this estimated value of κ. All the estimated variables are then refined using bundle-adjustment and auto-calibration is performed using the imaged circu-lar points, the imaged rotation axis and the associated vanishing point. This allows the recovery of the full camera positions and orientations, and hence Euclidean reconstruction. Experimental results demonstrate the simplicity of this novel approach and the high precision in the estimated motion and reconstruction. 1 Hui Zhang 0062, Guoqiang Zhang 0003, Kwan-Yee Kenneth Wong |
BMVC | 3 |
| 2005 | Complex 3D Shape Recovery Using a Dual-Space ApproachabstractThis paper presents a novel method for reconstructing complex 3D objects with unknown topology using silhouettes extracted from image sequences. This method exploits the duality principle governing surface points and their corresponding tangent planes, and enables a direct estimation of points on the contour generators. A major problem in other related works concerns with the search for a tangent basis at singularities in the dual tangent space. This problem is addressed here by utilizing the epipolar parameterization for identifying a well-defined basis at each point, and thus avoids any form of search. For the degenerate cases where epipolar parameterization breaks up, a fast on-the-fly validation is performed for each computed surface point, which consequently leads to a significant improvement in robustness. As the resulting contour generator points are not suitable for direct triangulation, a topologically correct surface extracting method based on slicing plane is presented. Both experiments on synthetic and real world data show that the proposed method has comparable robustness as those existing volumetric methods regarding surface of complex topology, whilst producing more accurate estimation of surface points. Chen Liang 0004, Kwan-Yee Kenneth Wong |
CVPR (2) | 2 |
| 2005 | Explicit contour model for vehicle tracking with automatic hypothesis validationabstractThis paper addresses the problem of vehicle tracking under a single static, uncalibrated camera without any constraints on the scene or on the motion direction of vehicles. We introduce an explicit contour model, which not only provides a good approximation to the contours of all classes of vehicles but also embeds the contour dynamics in its parameterized template. We integrate the model into a Bayesian framework with multiple cues for vehicle tracking, and evaluate the correctness of a target hypothesis, with the information implied by the shape, by monitoring any conflicts within the hypothesis of every single target as well as between the hypotheses of all targets. We evaluated the proposed method using some real sequences, and demonstrated its effectiveness in tracking vehicles, which have their shape changed significantly while moving on curly, uphills roads. Boris Wai-Sing Yiu, Kwan-Yee Kenneth Wong, Francis Y. L. Chin, Ronald H. Y. Chung |
ICIP (2) | 2 |
| 2005 | Camera calibration with spheres: linear approachesabstractThis paper addresses the problem of camera calibration from spheres. By studying the relationship between the dual images of spheres and that of the absolute conic, a linear solution has been derived from a recently proposed non-linear semi-definite approach. However, experiments show that this approach is quite sensitive to noise. In order to overcome this problem, a second approach has been proposed, where the orthogonal calibration relationship is obtained by regarding any two spheres as a surface of revolution. This allows a camera to be fully calibrated from an image of three spheres. Besides, a conic homography is derived from the imaged spheres, and from its eigenvectors the orthogonal invariants can be computed directly. Experiments on synthetic and real data show the practicality of such an approach. Hui Zhang 0062, Guoqiang Zhang 0003, Kwan-Yee Kenneth Wong |
ICIP (2) | 3 |
| 2004 | Segmenting lumbar vertebrae in digital video fluoroscopic images through edge enhancementabstractVideo fluoroscopy provides a cost effective way for the diagnosis of low back pain. Backbones or vertebrae are usually segmented manually from fluoroscopic images of low quality during such a diagnosis. In this paper, we try to reduce human workload by performing automatic vertebrae detection and segmentation. Operators need to provide the rough location of landmarks only. The proposed algorithm would perform edge detection, which is based on pattern recognition of texture, along the snake formed from the landmarks. The snake would then attach to the edge detected. Experimental results show that the proposed system can segment vertebrae from video fluoroscopic image automatically and accurately. Shu-Fai Wong, Kwan-Yee Kenneth Wong |
ICARCV | 2 |
| 2004 | Robust Image Segmentation by Texture Sensitive Snake Under Low Contrast Environment
Shu-Fai Wong, Kwan-Yee Kenneth Wong |
ICINCO (2) | 2 |
| 2004 | Wavelet Network for Nonlinear Regression Using Probabilistic Framework
Shu-Fai Wong, Kwan-Yee Kenneth Wong |
ISNN (2) | 2 |
| 2004 | Fitting Subdivision Surfaces to Unorganized Point Data Using SDMabstractWe study the reconstruction of smooth surfaces from point clouds. We use a new squared distance error term in optimization to fit a subdivision surface to a set of unorganized points, which defines a closed target surface of arbitrary topology. The resulting method is based on the framework of squared distance minimization (SDM) proposed by Pottmann et al. Specifically, with an initial subdivision surface having a coarse control mesh as input, we adjust the control points by optimizing an objective function through iterative minimization of a quadratic approximant of the squared distance function of the target shape. Our experiments show that the new method (SDM) converges much faster than the commonly used optimization method using the point distance error function, which is known to have only linear convergence. This observation is further supported by our recent result that SDM can be derived from the Newton method with necessary modifications to make the Hessian positive definite and the fact that the Newton method has quadratic convergence. Kin-Shing D. Cheng, Wenping Wang 0001, Hong Qin 0001, Kwan-Yee Kenneth Wong, Huaiping Yang, Yang Liu 0014 |
PG | 4 |
| 2004 | Reconstruction of surfaces of revolution from single uncalibrated views
Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
Image Vis. Comput. | 1 |
| 2004 | Reconstruction of sculpture from its profiles with unknown camera positionsabstractProfiles of a sculpture provide rich information about its geometry, and can be used for shape recovery under known camera motion. By exploiting correspondences induced by epipolar tangents on the profiles, a successful solution to motion estimation from profiles has been developed in the special case of circular motion. The main drawbacks of using circular motion alone, namely the difficulty in adding new views and part of the object always being invisible, can be overcome by incorporating arbitrary general views of the object and registering its new profiles with the set of profiles resulted from the circular motion. In this paper, we describe a complete and practical system for producing a three-dimensional (3-D) model from uncalibrated images of an arbitrary object using its profiles alone. Experimental results on various objects are presented, demonstrating the quality of the reconstructions using the estimated motion. Kwan-Yee Kenneth Wong, Roberto Cipolla |
IEEE Trans. Image Process. | 1 |
| 2003 | Camera Calibration from Surfaces of RevolutionabstractThis paper addresses the problem of calibrating a pinhole camera from images of a surface of revolution. Camera calibration is the process of determining the intrinsic or internal parameters (i.e., aspect ratio, focal length, and principal point) of a camera, and it is important for both motion estimation and metric reconstruction of 3D models. In this paper, a novel and simple calibration technique is introduced, which is based on exploiting the symmetry of images of surfaces of revolution. Traditional techniques for camera calibration involve taking images of some precisely machined calibration pattern (such as a calibration grid). The use of surfaces of revolution, which are commonly found in daily life (e.g., bowls and vases), makes the process easier as a result of the reduced cost and increased accessibility of the calibration objects. In this paper, it is shown that two images of a surface of revolution will provide enough information for determining the aspect ratio, focal length, and principal point of a camera with fixed intrinsic parameters. The algorithms presented in this paper have been implemented and tested with both synthetic and real data. Experimental results show that the camera calibration method presented is both practical and accurate. Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Reconstruction of Surfaces of Revolution from Single Uncalibrated ViewsabstractThis paper addresses the problem of recovering the 3D shape of a surface of revolution from a single uncalibrated perspective view.The algorithm introduced here makes use of the invariant properties of a surface of revolution and its silhouette to locate the image of the revolution axis, and to calibrate the focal length of the camera.The image is then normalized and rectified such that the resulting silhouette exhibits bilateral symmetry.Such a rectification leads to a simpler differential analysis of the silhouette, and yields a simple equation for depth recovery.Ambiguities in the reconstruction are analyzed and experimental results on real images are presented, which demonstrate the quality of the reconstruction. Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
BMVC | 1 |
| 2002 | Structure and motion estimation from apparent contours under circular motion
Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
Image Vis. Comput. | 1 |
| 2001 | Structure and Motion from SilhouettesabstractThis paper addresses the problem of recovering structure and motion from silhouettes. Silhouettes are projections of contour generators which are viewpoint dependent, and hence do not readily provide point correspondences for exploitation in motion estimation. Previous works have exploited correspondences induced by epipolar tangencies, and a successful solution has been developed in the special case of circular motion (turnable sequences). However, the main drawbacks are (1) new views cannot be added easily at a later time, and (2) part of the structure will always remain invisible under circular motion. In this paper we overcome the above problems by incorporating arbitrary general views and estimating the camera poses using silhouettes alone. We present a complete and practical system which produces high quality 3D models from 2D uncalibrated silhouettes. The 3D models thus obtained can be refined incrementally by adding new arbitrary views and estimating their poses. Experimental results on various objects are presented, demonstrating the quality of the reconstructions. Kwan-Yee Kenneth Wong, Roberto Cipolla |
ICCV | 1 |
| 2001 | Reconstruction of sculpture from uncalibrated image profilesabstractProfiles of a sculpture provide rich information about its geometry, and can be used for model reconstruction under known camera motion. By exploiting correspondences induced by epipolar tangents on the profiles, a successful solution to motion estimation has been developed for the case of circular motion. Arbitrary general views can then be incorporated to refine the model built from circular motion. Kwan-Yee Kenneth Wong, Roberto Cipolla |
ICIP (1) | 1 |
| 2001 | Epipolar Geometry from Profiles under Circular MotionabstractAddresses the problem of motion estimation from profiles (apparent contours) of an object rotating on a turntable in front of a single camera. A practical and accurate technique for solving this problem from profiles alone is developed. It is precise enough to reconstruct the shape of the object. No correspondences between points or lines are necessary. Symmetry of the surface of revolution swept out by the rotating object is exploited to obtain the image of the rotation axis and the homography relating epipolar lines in two views robustly and elegantly. These, together with geometric constraints for images of rotating objects, are used to obtain first the image of the horizon, which is the projection of the plane that contains the camera centers, and then the epipoles, thus fully determining the epipolar geometry of the image sequence. The estimation of this geometry by this sequential approach avoids many of the problems found in other algorithms. The search for the epipoles, by far the most critical step, is carried out as a simple 1D optimization. Parameter initialization is trivial and completely automatic at all stages. After the estimation of the epipolar geometry, the Euclidean motion is recovered using the fixed intrinsic parameters of the camera obtained either from a calibration grid or from self-calibration techniques. Finally, the spinning object is reconstructed from its profiles using the motion estimated in the previous stage. Results from real data are presented, demonstrating the efficiency and usefulness of the proposed methods. Paulo R. S. Mendonça, Kwan-Yee Kenneth Wong, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | Camera Pose Estimation and Reconstruction from Image Profiles under Circular Motion
Paulo R. S. Mendonça, Kwan-Yee Kenneth Wong, Roberto Cipolla |
ECCV (2) | 2 |
| 1999 | Reconstruction and Motion Estimation from Apparent Contours under Circular MotionabstractIn this paper we address the problem of recovering structure and motion from the contours of a smooth-curved surface. A novel and simpler technique for computing the structure of an object from its profiles is introduced. Experiments with real data show encouraging results, which are comparable to those obtained from much more sophisticated techniques. Furthermore, a new method for motion estimation from sequences of profiles is proposed. Preliminary results demonstrate the feasibility of the algorithm. 1 Introduction The recovering of structure and motion from sequences of images is a central problem in computer vision, and its solution has generated a rich pool of algorithms [8, 1]. Most of these algorithms rely on correspondences of points or lines between images, and work well when the scene being viewed is composed of polyhedral parts. However, for smooth surfaces without noticeable texture, point and line correspondences may not be easily established. In this case the profile of... Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
BMVC | 1 |