VLDB 2026 Research / reviewers in the wild / expert
Tien-Tsin Wong
dblp:69/220
· DBLP profile ↗
163ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0002-7792-9307ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 143 · 6 first-author · 34 since 2021Artificial intelligence and machine learning · 39 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mamba-Driven Multi-View Discriminative Clustering via Global-Local Cross-View Sequence ModelingabstractMulti-view clustering (MVC) has recently garnered increasing attention for its ability to partition unlabeled samples into distinct clusters by leveraging complementary and consistent information from different views. Existing MVC methods primarily combine deep neural networks with contrastive learning for cross-view representation learning, yet often overlook the inherent global-local structural relationships among samples. While GNN-based methods capture local structures, they struggle to model global dependencies, leading to inferior inter-cluster separability. In contrast, Transformer-based methods excel at global aggregation but suffer from quadratic complexity, and their attention smoothing effect weakens fine-grained local structures, resulting in suboptimal intra-cluster compactness. To address these limitations, we propose a novel end-to-end MVC framework called Mamba-Driven Multi-View Discriminative Clustering via Global-Local Cross-View Sequence Modeling (MGLC). By flexibly constructing multi-view sequences, MGLC fully exploits the efficient sequence modeling capabilities of Mamba to jointly model cross-view dependencies and global-local structural relationships among samples. Furthermore, MGLC introduces a Cross-Mamba Fusion module to dynamically integrate cross-view and global-local structural representations. Additionally, MGLC incorporates a Dual Calibration Contrastive Learning module, guided by high-confidence pseudo-labels, that adaptively refines both feature and semantic representations while mitigating false negatives among semantically similar samples. Extensive comparative experiments and ablation studies demonstrate the effectiveness of MGLC. Yuanyang Zhang, Xinhang Wan, Jie Xu 0044, Cunjian Chen, Tien-Tsin Wong, Li Yao 0003, Yijie Lin 0001 |
AAAI | 6 |
| 2026 | DiVE: Decoupling Intra-layer Visual Evidence for Mitigating Hallucinations in Large Vision-Language ModelsabstractRecent Large Vision-Language Models (LVLMs) have achieved significant progress yet frequently suffer from visual hallucinations, often stemming from an over-reliance on language priors rather than visual evidence.Existing decoding-based approaches often rely on input perturbations to weaken language priors, but they do not explicitly decouple visual evidence from mixed vision-language representations.To address these limitations, we propose DiVE (Decoupling intra-layer Visual Evidence).DiVE dynamically identifies layers enriched with visual information and performs intra-layer decoupling to extract aggregated visual evidence.By suppressing this evidence to construct a language-priordominated reference distribution, DiVE employs contrastive decoding to calibrate the output logits, thereby mitigating hallucinations.Extensive experiments across diverse LVLM architectures demonstrate that DiVE achieves state-of-the-art performance among decodingbased methods on multiple benchmarks.Crucially, it eliminates the latency of an extra forward pass, offering a lightweight and efficient solution. Li Lin 0011, Hui Jiao, Li Yao 0003, Tien-Tsin Wong, Hanqian Wu |
ACL (1) | 5 |
| 2026 | GlassSplat: Geometric Consistency and Pruning for Reflection-Free 3D Scene ReconstructionabstractRendering high-fidelity 3D scenes is crucial for immersive applications like virtual reality and digital twins. However, standard 3D Gaussian Splatting (3DGS) relies heavily on multi-view consistency, making it fragile in real-world scenarios plagued by glass reflections. These reflections often manifest as geometric "floaters" or severe texture artifacts, obscuring the true background. Existing solutions, which typically employ single-image priors or NeRF-based in-painting, often lack explicit 3D constraints or rely on synthetic data, failing to generalize to complex environments. To address these challenges, we first present a novel benchmark dataset of 8 real-world scenes, capturing physically paired reflective and reflection-free images. Building on this, we propose GlassSplat, a robust framework designed to eliminate view-dependent artifacts and recover clean transmission geometry. Our method initializes with a reflection prior and introduces an Affine-Based Exposure Correction module to align global photometric inconsistencies. To distinguishing valid geometry from virtual outliers, we incorporate an Epipolar Consistency Loss and an uncertainty-weighted Depth Regularization. Finally, to physically purge residual noise, we devise a Visibility-Aware Pruning strategy that dynamically filters artifacts based on multi-view statistics. Extensive experiments demonstrate that GlassSplat significantly outperforms state-of-the-art approaches, effectively recovering a clean, artifact-free 3D scene representation. Jingjiao You, Yuanyang Zhang, Yining Xu 0001, Li Yao 0003, Cunjian Chen, Tien-Tsin Wong |
ICMR | 7 |
| 2026 | Consistent and Controllable Image Animation With Linear Motion Diffusion TransformersabstractImage animation has seen significant progress, driven by the powerful generative capabilities of diffusion models. However, maintaining appearance consistency with static input images and mitigating abrupt motion transitions in generated animations remain persistent challenges. While text-to-video (T2V) generation has demonstrated impressive performance with diffusion transformer models, the image animation field still largely relies on U-Net-based diffusion models, which lag behind the latest T2V approaches. Moreover, the quadratic complexity of vanilla self-attention mechanisms in Transformers imposes heavy computational demands, making image animation particularly resource-intensive. To address these issues, we propose MiraMo, a framework designed to enhance efficiency, appearance consistency, and motion smoothness in image animation. Specifically, MiraMo introduces three key elements: (1) A foundational text-to-video architecture replacing vanilla self-attention with efficient linear attention to reduce computational overhead while preserving generation quality; (2) A novel motion residual learning paradigm that focuses on modeling motion dynamics rather than directly predicting frames, improving temporal consistency; and (3) A DCT-based noise refinement strategy during inference to suppress sudden motion artifacts, complemented by a dynamics control module to balance motion smoothness and expressiveness. Extensive experiments against state-of-the-art methods validate the superiority of MiraMo in generating consistent, smooth, and controllable animations with accelerated inference speed. Additionally, we demonstrate the versatility of MiraMo through applications in motion transfer and video editing tasks. Xin Ma 0031, Yaohui Wang 0001, Genyun Jia, Tien-Tsin Wong, Cunjian Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Consistent and Controllable Image Animation with Motion Diffusion ModelsabstractDiffusion models have achieved significant progress in the task of image animation due to their powerful generative capabilities. However, preserving appearance consistency to the static input image, and avoiding abrupt motion change in the generated animation, remains challenging. In this paper, we introduce Cinemo, a novel image animation approach that aims at achieving better appearance consistency and motion smoothness. The core of Cinemo is to focus on learning the distribution of motion residuals, rather than directly predicting frames as in existing diffusion models. During the inference, we further mitigate the sudden motion changes in the generated video by introducing a novel DCT-based noise refinement strategy. To counteract the over-smoothing of motion, we introduce a dynamics degree control design for better control of the magnitude of motion. Altogether, these strategies enable Cinemo to produce highly consistent, smooth, and motion-controllable results. Extensive experiments compared with several state-of-the-art methods demonstrate the effectiveness and superiority of our proposed approach. In the end, we also demonstrate how our model can be applied for motion transfer or video editing of any given video. The project page is available at https://maxin-cn.github.io/cinemo_project/. Xin Ma 0031, Yaohui Wang 0001, Gengyun Jia, Tien-Tsin Wong, Yuan-Fang Li, Cunjian Chen |
CVPR | 5 |
| 2025 | Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 DatasetabstractManga, a popular form of multimodal artwork, has traditionally been overlooked in deep learning advancements due to the absence of a robust dataset and comprehensive annotation. Manga segmentation is the key to the digital migration of manga. There exists a significant domain gap between the manga and the natural images, that fails most existing learning-based methods. To address this gap, we introduce an augmented segmentation annotation for the Manga109 dataset, a collection of 109 manga volumes, that offers intricate artworks in a rich variety of styles. We introduce a detailed annotation that extends beyond the original simple bounding boxes to the segmentation masks with pixel-level precision. It provides object category, location, and instance information that can be used for semantic segmentation and instance segmentation. We also provide a comprehensive analysis of our annotation dataset from various aspects. We further measure the improvement of the state-of-the-art segmentation model after training it with our augmented dataset. The benefits of this augmented dataset are profound, with the potential to significantly enhance manga analysis algorithms and catalyze the novel development in digital art processing and cultural analytics. This annotation, named MangaSeg, is publicly available at https://huggingface.co/datasets/MS92/MangaSegmentation. Minshan Xie, Hanyuan Liu, Chengze Li, Tien-Tsin Wong |
CVPR | 5 |
| 2025 | BlueNeg: A 35MM Negative Film Dataset for Restoring Channel-Heterogeneous Deterioration
Hanyuan Liu, Chengze Li, Minshan Xie, Zhenni Wang, Jiawen Liang, Andrew Chi-Sing Leung, Tien-Tsin Wong |
ICCV | 7 |
| 2025 | VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical PriorabstractVideo diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos due to an inherent lack of understanding of physics, resulting in incorrect dynamics and event sequences. To address this limitation, we propose a novel two-stage image-to-video generation framework that explicitly incorporates physics with vision and language informed physical prior. In the first stage, we employ a Vision Language Model (VLM) as a coarse-grained motion planner, integrating chain-of-thought and physics-aware reasoning to predict a rough motion trajectories/changes that approximate real-world physical dynamics while ensuring the inter-frame consistency. In the second stage, we use the predicted motion trajectories/changes to guide the video generation of a VDM. As the predicted motion trajectories/changes are rough, noise is added during inference to provide freedom to the VDM in generating motion with more fine details. Extensive experimental results demonstrate that our framework can produce physically plausible motion, and comparative evaluations highlight the notable superiority of our approach over existing methods. More video results are available on our Project Page: https://madaoer.github.io/projects/physically_plausible_video_generation. Xindi Yang, Baolu Li 0001, Zhenfei Yin, Lei Bai 0001, Liqian Ma, Zhiyong Wang 0001, Jianfei Cai 0001, Tien-Tsin Wong, Huchuan Lu, Xu Jia 0012 |
ICCV | 9 |
| 2025 | D2Diff: A Dual-Domain Diffusion Model for Accurate Multi-Contrast MRI Synthesis
Sanuwani Dayarathna, Himashi Peiris, Kh Tohidul Islam, Tien-Tsin Wong, Zhaolin Chen |
MICCAI (2) | 4 |
| 2025 | ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion Models
Hanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li, Andrew Chi-Sing Leung, Tien-Tsin Wong |
ACM Multimedia | 6 |
| 2025 | Synchronized Multi-Frame Diffusion for Temporally Consistent Video StylizationabstractAbstract Text‐guided video‐to‐video stylization transforms the visual appearance of a source video to a different appearance guided on textual prompts. Existing text‐guided image diffusion models can be extended for stylized video synthesis. However, they struggle to generate videos with both highly detailed appearance and temporal consistency. In this paper, we propose a synchronized multi‐frame diffusion framework to maintain both the visual details and the temporal consistency. Frames are denoised in a synchronous fashion, and more importantly, information of different frames is shared since the beginning of the denoising process. Such information sharing ensures that a consensus, in terms of the overall structure and color distribution, among frames can be reached in the early stage of the denoising process before it is too late. The optical flow from the original video serves as the connection, and hence the venue for information sharing, among frames. We demonstrate the effectiveness of our method in generating high‐quality and diverse results in extensive experiments. Our method shows superior qualitative and quantitative results compared to state‐of‐the‐art video editing methods. Minshan Xie, Hanyuan Liu, Chengze Li, Tien-Tsin Wong |
Comput. Graph. Forum | 4 |
| 2025 | Screentone-Preserved Manga RetargetingabstractAbstract As a popular comic style, manga offers a unique impression by utilizing a rich set ofbitonal patterns, or screentones, for illustration. However, screentones can easily be degraded when manga is resized in terms of aspect ratio and resolution for manga re‐layout and e‐manga migration applications. To tackle this problem, we propose the first automatic manga retargeting method that synthesizes a retargeted manga image while preserving the prominent structure and fine screentone intended by the manga artist. While modern natural photo retargeting methods can achieve prominent structure preservation, preserving screentones within arbitrarily shaped regions is very challenging due to two properties of manga: (i) pattern constancy under translation, and (ii) non‐compatibility with interpolation. To circumvent this barrier, we propose learning a quantized representation of screentones that is translation‐invariant and pointwisely representable through a tailored manga reconstruction network with a screentone‐anchored codebook. Thanks to these merits, we can perform the re‐synthesis operation using existing photo retargeting methods and achieve the desired manga retargeting results. We conducted extensive qualitative and quantitative experiments to validate the effectiveness of our method, and we achieved notably compelling results compared to alternative methods. Minshan Xie, Menghan Xia, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
Comput. Graph. Forum | 5 |
| 2025 | Make-Your-Video: Customized Video Generation Using Textual and Structural GuidanceabstractCreating a vivid video from the event or scenario in our imagination is a truly fascinating experience. Recent advancements in text-to-video synthesis have unveiled the potential to achieve this with prompts only. While text is convenient in conveying the overall scene context, it may be insufficient to control precisely. In this paper, we explore customized video generation by utilizing text as context description and motion structure (e.g., frame-wise depth) as concrete guidance. Our method, dubbed Make-Your-Video, involves joint-conditional video generation using a Latent Diffusion Model that is pre-trained for still image synthesis and then promoted for video generation with the introduction of temporal modules. This two-stage learning scheme not only reduces the computing resources required, but also improves the performance by transferring the rich concepts available in image datasets solely into video generation. Moreover, we use a simple yet effective causal attention mask strategy to enable longer video synthesis, which mitigates the potential quality degradation effectively. Experimental results show the superiority of our method over existing baselines, particularly in terms of temporal coherence and fidelity to users' guidance. In addition, our model enables several intriguing applications that demonstrate potential for practical usage. Jinbo Xing, Menghan Xia, Yuechen Zhang, Yong Zhang 0034, Yingqing He, Hanyuan Liu, Haoxin Chen, Xiaodong Cun, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 12 |
| 2024 | DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors
Jinbo Xing, Menghan Xia, Yong Zhang 0034, Hao Chen 0011, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
ECCV (46) | 10 |
| 2024 | SKETCH2MANGA: Shaded Manga Screening from Sketch with Diffusion ModelsabstractWhile manga is a popular entertainment form, creating manga is tedious, especially adding screentones to the created sketch, namely manga screening. Unfortunately, there is no existing method that tailors for automatic manga screening, probably due to the difficulty in generating shaded high-frequency screentones of high-quality. Classic manga screening approaches generally require user input to provide screentone exemplars or a reference manga image. Recent deep learning models enable automatic generation by learning from a large-scale dataset. However, the state-of-the-art models still fail to generate high-quality shaded screentones due to the lack of a tailored model and high-quality manga training data. In this paper, we propose a novel sketch-to-manga framework that first generates a color illustration from the sketch and then generates a screentoned manga based on the intensity guidance. Our method significantly outperforms existing methods in generating high-quality manga with shaded high-frequency screentones. Xueting Liu 0001, Chengze Li, Minshan Xie, Tien-Tsin Wong |
ICIP | 5 |
| 2024 | Text-Guided Texturing by Synchronized Multi-View DiffusionabstractThis paper introduces a novel approach to synthesize texture to dress up a 3D object, given a text prompt. Based on the pre-trained text-to-image (T2I) diffusion model, existing methods usually employ a project-and-inpaint approach, in which a view of the given object is first generated and warped to another view for inpainting. But it tends to generate inconsistent texture due to the asynchronous diffusion of multiple views. We believe that such asynchronous diffusion and insufficient information sharing among views are the root causes of the inconsistent artifacts. In this paper, we propose a synchronized multi-view diffusion approach that allows the diffusion processes from different views to reach a consensus on the generated content early in the process, and hence ensures the texture consistency. To synchronize the diffusion, we share the denoised content among different views in each denoising step, specifically by blending the latent content in the texture domain from overlapping views. Our method demonstrates superior performance in generating consistent, seamless and highly detailed textures, comparing to state-of-the-art methods. © 2024 Copyright held by the owner/author(s). Minshan Xie, Hanyuan Liu, Tien-Tsin Wong |
SIGGRAPH Asia | 4 |
| 2024 | ToonCrafter: Generative Cartoon InterpolationabstractWe introduce ToonCrafter, a novel approach that transcends traditional correspondence-based cartoon video interpolation, paving the way for generative interpolation. Traditional methods, that implicitly assume linear motion and the absence of complicated phenomena like dis-occlusion, often struggle with the exaggerated non-linear and large motions with occlusion commonly found in cartoons, resulting in implausible or even failed interpolation results. To overcome these limitations, we explore the potential of adapting live-action video priors to better suit cartoon interpolation within a generative framework. ToonCrafter effectively addresses the challenges faced when applying live-action video motion priors to generative cartoon interpolation. First, we design a toon rectification learning strategy that seamlessly adapts live-action video priors to the cartoon domain, resolving the domain gap and content leakage issues. Next, we introduce a dual-reference-based 3D decoder to compensate for lost details due to the highly compressed latent prior spaces, ensuring the preservation of fine details in interpolation results. Finally, we design a flexible sketch encoder that empowers users with interactive control over the interpolation results. Experimental results demonstrate that our proposed method not only produces visually convincing and more natural dynamics, but also effectively handles dis-occlusion. The comparative evaluation demonstrates the notable superiority of our approach over existing competitors. Code and model weights are available at https://doubiiu.github.io/projects/ToonCrafter Jinbo Xing, Hanyuan Liu, Menghan Xia, Yong Zhang 0034, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
ACM Trans. Graph. | 7 |
| 2024 | Taming Reversible Halftoning Via Predictive LuminanceabstractTraditional halftoning usually drops colors when dithering images with binary dots, which makes it difficult to recover the original color information. We proposed a novel halftoning technique that converts a color image into a binary halftone with full restorability to its original version. Our novel base halftoning technique consists of two convolutional neural networks (CNNs) to produce the reversible halftone patterns, and a noise incentive block (NIB) to mitigate the flatness degradation issue of CNNs. Furthermore, to tackle the conflicts between the blue-noise quality and restoration accuracy in our novel base method, we proposed a predictor-embedded approach to offload predictable information from the network, which in our case is the luminance information resembling from the halftone pattern. Such an approach allows the network to gain more flexibility to produce halftones with better blue-noise quality without compromising the restoration quality. Detailed studies on the multiple-stage training method and loss weightings have been conducted. We have compared our predictor-embedded method and our novel method regarding spectrum analysis on halftone, halftone accuracy, restoration accuracy, and the data embedding studies. Our entropy evaluation evidences our halftone contains less encoding information than our novel base method. The experiments show our predictor-embedded method gains more flexibility to improve the blue-noise quality of halftones and maintains a comparable restoration quality with a higher tolerance for disturbances. Cheuk-Kit Lau, Menghan Xia, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | LF2MV: Learning an Editable Meta-View Towards Light Field RepresentationabstractLight fields are 4D scene representations that are typically structured as arrays of views or several directional samples per pixel in a single view. However, this highly correlated structure is not very efficient to transmit and manipulate, especially for editing. To tackle this issue, we propose a novel representation learning framework that can encode the light field into a single meta-view that is both compact and editable. Specifically, the meta-view composes of three visual channels and a complementary meta channel that is embedded with geometric and residual appearance information. The visual channels can be edited using existing 2D image editing tools, before reconstructing the whole edited light field. To facilitate edit propagation against occlusion, we design a special editing-aware decoding network that consistently propagates the visual edits to the whole light field upon reconstruction. Extensive experiments show that our proposed method achieves competitive representation accuracy and meanwhile enables consistent edit propagation. Menghan Xia, Jose Echevarria, Minshan Xie, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorabstractSpeech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formulate the cross-modal mapping into a regression task, which suffers from the regression-to-mean problem leading to over-smoothed facial motions. In this paper, we propose to cast speech-driven facial animation as a code query task in a finite proxy space of the learned codebook, which effectively promotes the vividness of the generated motions by reducing the cross-modal mapping uncertainty. The codebook is learned by self-reconstruction over real facial motions and thus embedded with realistic facial motion priors. Over the discrete motion space, a temporal autoregressive model is employed to sequentially synthesize facial motions from the input speech signal, which guarantees lip-sync as well as plausible facial expressions. We demonstrate that our approach outperforms current state-of-the-art methods both qualitatively and quantitatively. Also, a user study further justifies our superiority in perceptual quality. Code and video demo are available at https://doubiiu.github.io/projects/codetalker. Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang 0001, Tien-Tsin Wong |
CVPR | 6 |
| 2023 | Scale-Arbitrary Invertible Image DownscalingabstractConventional social media platforms usually downscale high-resolution (HR) images to restrict their resolution to a specific size for saving transmission/storage cost, which makes those visual details inaccessible to other users. To bypass this obstacle, recent invertible image downscaling methods jointly model the downscaling/upscaling problems and achieve impressive performance. However, they only consider fixed integer scale factors and may be inapplicable to generic downscaling tasks towards resolution restriction as posed by social media platforms. In this paper, we propose an effective and universal Scale-Arbitrary Invertible Image Downscaling Network (AIDN), to downscale HR images with arbitrary scale factors in an invertible manner. Particularly, the HR information is embedded in the downscaled low-resolution (LR) counterparts in a nearly imperceptible form such that our AIDN can further restore the original HR images solely from the LR images. The key to supporting arbitrary scale factors is our proposed Conditional Resampling Module (CRM) that conditions the downscaling/upscaling kernels and sampling locations on both scale factors and image content. Extensive experimental results demonstrate that our AIDN achieves top performance for invertible downscaling with both arbitrary integer and non-integer scale factors. Also, both quantitative and qualitative evaluations show our AIDN is robust to the lossy image compression standard. The source code and trained models are publicly available at https://github.com/Doubiiu/AIDN. Jinbo Xing, Wenbo Hu 0002, Menghan Xia, Tien-Tsin Wong |
IEEE Trans. Image Process. | 4 |
| 2023 | Point Set Self-EmbeddingabstractThis work presents an innovative method for point set self-embedding, that encodes the structural information of a dense point set into its sparser version in a visual but imperceptible form. The self-embedded point set can function as the ordinary downsampled one and be visualized efficiently on mobile devices. Particularly, we can leverage the self-embedded information to fully restore the original point set for detailed analysis on remote servers. This task is challenging, since both the self-embedded point set and the restored point set should resemble the original one. To achieve a learnable self-embedding scheme, we design a novel framework with two jointly-trained networks: one to encode the input point set into its self-embedded sparse point set and the other to leverage the embedded information for inverting the original point set back. Further, we develop a pair of up-shuffle and down-shuffle units in the two networks, and formulate loss terms to encourage the shape similarity and point distribution in the results. Extensive qualitative and quantitative results demonstrate the effectiveness of our method on both synthetic and real-scanned datasets. The source code and trained models will be publicly available at https://github.com/liruihui/Self-Embedding. Ruihui Li, Xianzhi Li 0001, Tien-Tsin Wong, Chi-Wing Fu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | End-to-End Line Drawing VectorizationabstractVector graphics is broadly used in a variety of forms, such as illustrations, logos, posters, billboards, and printed ads. Despite its broad use, many artists still prefer to draw with pen and paper, which leads to a high demand of converting raster designs into the vector form. In particular, line drawing is a primary art and attracts many research efforts in automatically converting raster line drawings to vector form. However, the existing methods generally adopt a two-step approach, stroke segmentation and vectorization. Without vector guidance, the raster-based stroke segmentation frequently obtains unsatisfying segmentation results, such as over-grouped strokes and broken strokes. In this paper, we make an attempt in proposing an end-to-end vectorization method which directly generates vectorized stroke primitives from raster line drawing in one step. We propose a Transformer-based framework to perform stroke tracing like human does in an automatic stroke-by-stroke way with a novel stroke feature representation and multi-modal supervision to achieve vectorization with high quality and fidelity. Qualitative and quantitative evaluations show that our method achieves state of the art performance. Hanyuan Liu, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
AAAI | 4 |
| 2022 | Neural Recognition of Dashed Curves with Gestalt Law of ContinuityabstractDashed curve is a frequently used curve form and is widely used in various drawing and illustration applications. While humans can intuitively recognize dashed curves from disjoint curve segments based on the law of continuity in Gestalt psychology, it is extremely difficult for computers to model the Gestalt law of continuity and recognize the dashed curves since high-level semantic understanding is needed for this task. The various appear-ances and styles of the dashed curves posed on a potentially noisy background further complicate the task. In this paper, we propose an innovative Transformer-based framework to recognize dashed curves based on both high-level features and low-level clues. The framework manages to learn the computational analogy of the Gestalt Law in various do-mains to locate and extract instances of dashed curves in both raster and vector representations. Qualitative and quantitative evaluations demonstrate the efficiency and ro-bustness of our framework over all existing solutions. Hanyuan Liu, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
CVPR | 4 |
| 2022 | Pseudo Bias-Balanced Learning for Debiased Chest X-Ray Classification
Luyang Luo, Dunyuan Xu, Hao Chen 0011, Tien-Tsin Wong, Pheng-Ann Heng |
MICCAI (8) | 4 |
| 2022 | Personalized Image Recoloring for Color Vision Deficiency CompensationabstractSeveral image recoloring methods have been proposed to compensate for the loss of contrast caused by color vision deficiency (CVD). However, these methods only work for dichromacy (a case in which one of the three types of cone cells loses its function completely), while the majority of CVD is anomalous trichromacy (another case in which one of the three types of cone cells partially loses its function). In this paper, a novel degree-adaptable recoloring algorithm is presented, which recolors images by minimizing an objective function constrained by contrast enhancement and naturalness preservation. To assess the effectiveness of the proposed method, a quantitative evaluation using common metrics and subjective studies involving 14 volunteers with varying degrees of CVD are conducted. The results of the evaluation experiment show that the proposed personalized recoloring method outperforms the state-of-the-art methods, achieving desirable contrast enhancement adapted to different degrees of CVD while preserving naturalness as much as possible. Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Kenji Kashiwagi, Issei Fujishiro, Tien-Tsin Wong, Xiaoyang Mao |
IEEE Trans. Multim. | 6 |
| 2022 | Disentangled Image Colorization via Global AnchorsabstractColorization is multimodal by nature and challenges existing frameworks to achieve colorful and structurally consistent results. Even the sophisticated autoregressive model struggles to maintain long-distance color consistency due to the fragility of sequential dependence. To overcome this challenge, we propose a novel colorization framework that disentangles color multimodality and structure consistency through global color anchors, so that both aspects could be learned effectively. Our key insight is that several carefully located anchors could approximately represent the color distribution of an image, and conditioned on the anchor colors, we can predict the image color in a deterministic manner by utilizing internal correlation. To this end, we construct a colorization model with dual branches, where the color modeler predicts the color distribution for anchor color representation, and the color generator predicts the pixel colors by referring the sampled anchor colors. Importantly, the anchors are located under two principles: color independence and global coverage, which is realized with clustering analysis on the deep color features. To simplify the computation, we creatively adopt soft superpixel segmentation to reduce the image primitives, which still nicely reserves the reversibility to pixel-wise representation. Extensive experiments show that our method achieves notable superiority over various mainstream frameworks in perceptual quality. Thanks to anchor-based color representation, our model has the flexibility to support diverse and controllable colorization as well. Menghan Xia, Wenbo Hu 0005, Tien-Tsin Wong, Jue Wang 0001 |
ACM Trans. Graph. | 3 |
| 2022 | Sprite-from-Sprite: Cartoon Animation Decomposition with Self-supervised Sprite EstimationabstractWe present an approach to decompose cartoon animation videos into a set of "sprites" --- the basic units of digital cartoons that depict the contents and transforms of each animated object. The sprites in real-world cartoons are unique: artists may draw arbitrary sprite animations for expressiveness, where the animated content is often complicated, irregular, and challenging; alternatively, artists may also reduce their workload by tweening and adjusting sprites, or even reuse static sprites, in which case the transformations are relatively regular and simple. Based on these observations, we propose a sprite decomposition framework using Pixel Multilayer Perceptrons (Pixel MLPs) where the estimation of each sprite is conditioned on and guided by all other sprites. In this way, once those relatively regular and simple sprites are resolved, the decomposition of the remaining "challenging" sprites can simplified and eased with the guidance of other sprites. We call this method "sprite-from-sprite" cartoon decomposition. We study ablative architectures of our framework, and the user study demonstrates that our results are the most preferred ones in 19/20 cases. Lvmin Zhang, Tien-Tsin Wong |
ACM Trans. Graph. | 2 |
| 2021 | Bidirectional Projection Network for Cross Dimension Scene Understandingabstract2D image representations are in regular grids and can be processed efficiently, whereas 3D point clouds are unordered and scattered in 3D space. The information inside these two visual domains is well complementary, e.g., 2D images have fine-grained texture while 3D point clouds contain plentiful geometry information. However, most current visual recognition systems process them individually. In this paper, we present a bidirectional projection network (BPNet) for joint 2D and 3D reasoning in an end-to-end manner. It contains 2D and 3D sub-networks with symmetric architectures, that are connected by our proposed bidirectional projection module (BPM). Via the BPM, complementary 2D and 3D information can interact with each other in multiple architectural levels, such that advantages in these two visual domains can be combined for better scene recognition. Extensive quantitative and qualitative experimental evaluations show that joint reasoning over 2D and 3D visual domains can benefit both 2D and 3D scene understanding simultaneously. Our BPNet achieves top performance on the ScanNetV2 benchmark for both 2D and 3D semantic segmentation. Code is available at https://github.com/wbhu/BPNet. Wenbo Hu 0002, Hengshuang Zhao, Li Jiang 0009, Jiaya Jia, Tien-Tsin Wong |
CVPR | 5 |
| 2021 | Exploiting Aliasing for Manga RestorationabstractAs a popular entertainment art form, manga enriches the line drawings details with bitonal screentones. However, manga resources over the Internet usually show screen-tone artifacts because of inappropriate scanning/rescaling resolution. In this paper, we propose an innovative two-stage method to restore quality bitonal manga from de-graded ones. Our key observation is that the aliasing induced by downsampling bitonal screentones can be utilized as informative clues to infer the original resolution and screentones. First, we predict the target resolution from the degraded manga via the Scale Estimation Network (SE-Net) with spatial voting scheme. Then, at the target resolution, we restore the region-wise bitonal screentones via the Manga Restoration Network (MR-Net) discriminatively, depending on the degradation degree. Specifically, the original screentones are directly restored in pattern-identifiable regions, and visually plausible screentones are synthesized in pattern-agnostic regions. Quantitative evaluation on synthetic data and visual assessment on real-world cases illustrate the effectiveness of our method. Minshan Xie, Menghan Xia, Tien-Tsin Wong |
CVPR | 3 |
| 2021 | User-Guided Line Art Flat Filling With Split Filling MechanismabstractFlat filling is a critical step in digital artistic content creation with the objective of filling line arts with flat colors. We present a deep learning framework for user-guided line art flat filling that can compute the "influence areas" of the user color scribbles, i.e., the areas where the user scribbles should propagate and influence. This framework explicitly controls such scribble influence areas for artists to manipulate the colors of image details and avoid color leakage/contamination between scribbles, and simultaneously, leverages data-driven color generation to facilitate content creation. This framework is based on a Split Filling Mechanism (SFM), which first splits the user scribbles into individual groups and then independently processes the colors and influence areas of each group with a Convolutional Neural Network (CNN). Learned from more than a million illustrations, the framework can estimate the scribble influence areas in a content-aware manner, and can smartly generate visually pleasing colors to assist the daily works of artists. We show that our proposed framework is easy to use, allowing even amateurs to obtain professional-quality results on a wide variety of line arts. Lvmin Zhang, Chengze Li, Edgar Simo-Serra, Yi Ji 0001, Tien-Tsin Wong, Chunping Liu |
CVPR | 5 |
| 2021 | Deep Halftoning with Reversible Binary PatternabstractExisting halftoning algorithms usually drop colors and fine details when dithering color images with binary dot patterns, which makes it extremely difficult to recover the original information. To dispense the recovery trouble in future, we propose a novel halftoning technique that converts a color image into binary halftone with full restorability to the original version. The key idea is to implicitly embed those previously dropped information into the halftone patterns. So, the halftone pattern not only serves to reproduce the image tone, maintain the blue-noise randomness, but also represents the color information and fine details. To this end, we exploit two collaborative convolutional neural networks (CNNs) to learn the dithering scheme, under a nontrivial self-supervision formulation. To tackle the flatness degradation issue of CNNs, we propose a novel noise incentive block (NIB) that can serve as a generic CNN plug-in for performance promotion. At last, we tailor a guiding-aware training scheme that secures the convergence direction as regulated. We evaluate the invertible halftones in multiple aspects, which evidences the effectiveness of our method. Menghan Xia, Wenbo Hu 0002, Xueting Liu 0001, Tien-Tsin Wong |
ICCV | 4 |
| 2021 | Conditional Directed Graph Convolution for 3D Human Pose EstimationabstractGraph convolutional networks have significantly improved 3D human pose estimation by representing the human skeleton as an undirected graph. However, this representation fails to reflect the articulated characteristic of human skeletons as the hierarchical orders among the joints are not explicitly presented. In this paper, we propose to represent the human skeleton as a directed graph with the joints as nodes and bones as edges that are directed from parent joints to child joints. By so doing, the directions of edges can explicitly reflect the hierarchical relationships among the nodes. Based on this representation, we further propose a spatial-temporal conditional directed graph convolution to leverage varying non-local dependence for different poses by conditioning the graph topology on input poses. Altogether, we form a U-shaped network, named U-shaped Conditional Directed Graph Convolutional Network, for 3D human pose estimation from monocular videos. To evaluate the effectiveness of our method, we conducted extensive experiments on two challenging large-scale benchmarks: Human3.6M and MPI-INF-3DHP. Both quantitative and qualitative results show that our method achieves top performance. Also, ablation studies show that directed graphs can better exploit the hierarchy of articulated human skeletons than undirected graphs, and the conditional connections can yield adaptive graph topologies for different poses. Wenbo Hu 0002, Changgong Zhang, Fangneng Zhan, Lei Zhang 0006, Tien-Tsin Wong |
ACM Multimedia | 5 |
| 2021 | Flow-aware synthesis: A generic motion model for video frame interpolationabstractA popular and challenging task in video research, frame interpolation aims to increase the frame rate of video. Most existing methods employ a fixed motion model, e.g., linear, quadratic, or cubic, to estimate the intermediate warping field. However, such fixed motion models cannot well represent the complicated non-linear motions in the real world or rendered animations. Instead, we present an adaptive flow prediction module to better approximate the complex motions in video. Furthermore, interpolating just one intermediate frame between consecutive input frames may be insufficient for complicated non-linear motions. To enable multi-frame interpolation, we introduce the time as a control variable when interpolating frames between original ones in our generic adaptive flow prediction module. Qualitative and quantitative experimental results show that our method can produce high-quality results and outperforms the existing state-of-the-art methods on popular public datasets. Jinbo Xing, Wenbo Hu 0002, Yuechen Zhang, Tien-Tsin Wong |
Comput. Vis. Media | 4 |
| 2021 | Video Snapshot: Single Image Motion Expansion via Invertible Motion EmbeddingabstractUnlike images, finding the desired video content in a large pool of videos is not easy due to the time cost of loading and watching. Most video streaming and sharing services provide the video preview function for a better browsing experience. In this paper, we aim to generate a video preview from a single image. To this end, we propose two cascaded networks, the motion embedding network and the motion expansion network. The motion embedding network aims to embed the spatio-temporal information into an embedded image, called video snapshot. On the other end, the motion expansion network is proposed to invert the video back from the input video snapshot. To hold the invertibility of motion embedding and expansion during training, we design four tailor-made losses and a motion attention module to make the network focus on the temporal information. In order to enhance the viewing experience, our expansion network involves an interpolation module to produce a longer video preview with a smooth transition. Extensive experiments demonstrate that our method can successfully embed the spatio-temporal information of a video into one "live" image, which can be converted back to a video preview. Quantitative and qualitative evaluations are conducted on a large number of videos to prove the effectiveness of our proposed method. In particular, statistics of PSNR and SSIM on a large number of videos show the proposed method is general, and it can generate a high-quality video from a single image. Qianshu Zhu, Chu Han, Guoqiang Han 0002, Tien-Tsin Wong, Shengfeng He |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Seamless manga inpainting with semantics awarenessabstractManga inpainting fills up the disoccluded pixels due to the removal of dialogue balloons or "sound effect" text. This process is long needed by the industry for the language localization and the conversion to animated manga. It is mostly done manually, as existing methods (mostly for natural image inpainting) cannot produce satisfying results. Manga inpainting is more tricky than natural image inpainting because its highly abstract illustration using structural lines and screentone patterns, which confuses the semantic interpretation and visual content synthesis. In this paper, we present the first manga inpainting method, a deep learning model, that generates high-quality results. Instead of direct inpainting, we propose to separate the complicated inpainting into two major phases, semantic inpainting and appearance synthesis. This separation eases both the feature understanding and hence the training of the learning model. A key idea is to disentangle the structural line and screentone, that helps the network to better distinguish the structural line and the screentone features for semantic interpretation. Both the visual comparison and the quantitative experiments evidence the effectiveness of our method and justify its superiority over existing state-of-the-art methods in the application of manga inpainting. Minshan Xie, Menghan Xia, Xueting Liu 0001, Chengze Li, Tien-Tsin Wong |
ACM Trans. Graph. | 5 |
| 2021 | Perceptual-Aware Sketch Simplification Based on Integrated VGG LayersabstractDeep learning has been recently demonstrated as an effective tool for raster-based sketch simplification. Nevertheless, it remains challenging to simplify extremely rough sketches. We found that a simplification network trained with a simple loss, such as pixel loss or discriminator loss, may fail to retain the semantically meaningful details when simplifying a very sketchy and complicated drawing. In this paper, we show that, with a well-designed multi-layer perceptual loss, we are able to obtain aesthetic and neat simplification results preserving semantically important global structures as well as fine details without blurriness and excessive emphasis on local structures. To do so, we design a multi-layer discriminator by fusing all VGG feature layers to differentiate sketches and clean lines. The weights used in layer fusing are automatically learned via an intelligent adjustment mechanism. Furthermore, to evaluate our method, we compare our method to state-of-the-art methods through multiple experiments, including visual comparison and intensive user study. Xuemiao Xu, Minshan Xie, Peiqi Miao, Wenpeng Xiao, Huaidong Zhang, Xueting Liu 0001, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2020 | Erasing Appearance Preservation in Optimization-Based Smoothing
Lvmin Zhang, Chengze Li, Yi Ji 0001, Chunping Liu, Tien-Tsin Wong |
ECCV (6) | 5 |
| 2020 | Example-Based Colourization Via Dense Encoding PyramidsabstractAbstract We propose a novel deep example‐based image colourization method called dense encoding pyramid network. In our study, we define the colourization as a multinomial classification problem. Given a greyscale image and a reference image, the proposed network leverages large‐scale data and then predicts colours by analysing the colour distribution of the reference image. We design the network as a pyramid structure in order to exploit the inherent multi‐scale, pyramidal hierarchy of colour representations. Between two adjacent levels, we propose a hierarchical decoder–encoder filter to pass the colour distributions from the lower level to higher level in order to take both semantic information and fine details into account during the colourization process. Within the network, a novel parallel residual dense block is proposed to effectively extract the local–global context of the colour representations by widening the network. Several experiments, as well as a user study, are conducted to evaluate the performance of our network against state‐of‐the‐art colourization methods. Experimental results show that our network is able to generate colourful, semantically correct and visually pleasant colour images. In addition, unlike fully automatic colourization that produces fixed colour images, the reference image of our network is flexible; both natural images and simple colour palettes can be used to guide the colourization. Chu-Feng Xiao 0001, Chu Han, Zhuming Zhang, Harry Qin, Tien-Tsin Wong, Guoqiang Han 0002, Shengfeng He |
Comput. Graph. Forum | 5 |
| 2020 | Mononizing binocular videosabstractThis paper presents the idea of mono-nizing binocular videos and a framework to effectively realize it. Mono-nize means we purposely convert a binocular video into a regular monocular video with the stereo information implicitly encoded in a visual but nearly-imperceptible form. Hence, we can impartially distribute and show the mononized video as an ordinary monocular video. Unlike ordinary monocular videos, we can restore from it the original binocular video and show it on a stereoscopic display. To start, we formulate an encoding-and-decoding framework with the pyramidal deformable fusion module to exploit long-range correspondences between the left and right views, a quantization layer to suppress the restoring artifacts, and the compression noise simulation module to resist the compression noise introduced by modern video codecs. Our framework is self-supervised, as we articulate our objective function with loss terms defined on the input: a monocular term for creating the mononized video, an invertibility term for restoring the original video, and a temporal term for frame-to-frame coherence. Further, we conducted extensive experiments to evaluate our generated mononized videos and restored binocular videos for diverse types of images and 3D movies. Quantitative results on both standard metrics and user perception studies show the effectiveness of our method. Wenbo Hu 0002, Menghan Xia, Chi-Wing Fu, Tien-Tsin Wong |
ACM Trans. Graph. | 4 |
| 2020 | Manga filling style conversion with screentone variational autoencoderabstractWestern color comics and Japanese-style screened manga are two popular comic styles. They mainly differ in the style of region-filling. However, the conversion between the two region-filling styles is very challenging, and manually done currently. In this paper, we identify that the major obstacle in the conversion between the two filling styles stems from the difference between the fundamental properties of screened region-filling and colored region-filling. To resolve this obstacle, we propose a screentone variational autoencoder, ScreenVAE, to map the screened manga to an intermediate domain. This intermediate domain can summarize local texture characteristics and is interpolative. With this domain, we effectively unify the properties of screening and color-filling, and ease the learning for bidirectional translation between screened manga and color comics. To carry out the bidirectional translation, we further propose a network to learn the translation between the intermediate domain and color comics. Our model can generate quality screened manga given a color comic, and generate color comic that retains the original screening intention by the bitonal manga artist. Several results are shown to demonstrate the effectiveness and convenience of the proposed method. We also demonstrate how the intermediate domain can assist other applications such as manga inpainting and photo-to-comic conversion. Minshan Xie, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
ACM Trans. Graph. | 4 |
| 2019 | Colorblind-Shareable VideosabstractThe two distinctive visual experiences of binocular display, with and without stereoscopic glasses, have been recently utilized for visual sharing between the colorblind and the normal-vision audiences. However, all existing methods only work for still images, and lack of temporal consistency for video application. In this paper, we propose the first synthesis method for colorblind-sharable videos that possess the temporal consistency for both visual experiences of colorblind and normal-vision, and retains all other crucial characteristics for visual sharing with colorblind. We formulate this challenging multi-constraint problem as a global optimization and minimize an objective function consisting of temporal term, color preservation term, color distinguishability term, and binocular fusibility term. Qualitative and quantitative experiments are conducted to evaluate the effectiveness of the proposed method comparing to existing methods. Xinghong Hu, Xueting Liu 0001, Tien-Tsin Wong |
CW | 4 |
| 2019 | Deep Line Drawing Vectorization via Line Subdivision and Topology ReconstructionabstractAbstract Vectorizing line drawing is necessary for the digital workflows of 2D animation and engineering design. But it is challenging due to the ambiguity of topology, especially at junctions. Existing vectorization methods either suffer from low accuracy or cannot deal with high‐resolution images. To deal with a variety of challenging containing different kinds of complex junctions, we propose a two‐phase line drawing vectorization method that analyzes the global and local topology. In the first phase, we subdivide the lines into partial curves, and in the second phase, we reconstruct the topology at junctions. With the overall topology estimated in the two phases, we can trace and vectorize the curves. To qualitatively and quantitatively evaluate our method and compare it with the existing methods, we conduct extensive experiments on not only existing datasets but also our newly synthesized dataset which contains different types of complex and ambiguous junctions. Experimental statistics show that our method greatly outperforms existing methods in terms of computational speed and achieves visually better topology reconstruction accuracy. Zhuming Zhang, Chu Han, Chengze Li, Tien-Tsin Wong |
Comput. Graph. Forum | 6 |
| 2019 | Deep residual learning for denoising Monte Carlo renderingsabstractLearning-based techniques have recently been shown to be effective for denoising Monte Carlo rendering methods. However, there remains a quality gap to state-of-the-art handcrafted denoisers. In this paper, we propose a deep residual learning based method that outperforms both state-of-the-art handcrafted denoisers and learning-based denoisers. Unlike the indirect nature of existing learning-based methods (which e.g., estimate the parameters and kernel weights of an explicit feature based filter), we directly map the noisy input pixels to the smoothed output. Using this direct mapping formulation, we demonstrate that even a simple-and-standard ResNet and three common auxiliary features (depth, normal, and albedo) are sufficient to achieve high-quality denoising. This minimal requirement on auxiliary data simplifies both training and integration of our method into most production rendering pipelines. We have evaluated our method on unseen images created by a different renderer. Consistently superior quality denoising is obtained in all cases. Kin-Ming Wong, Tien-Tsin Wong |
Comput. Vis. Media | 2 |
| 2019 | Colorblind-shareable videos by synthesizing temporal-coherent polynomial coefficientsabstractTo share the same visual content between color vision deficiencies (CVD) and normal-vision people, attempts have been made to allocate the two visual experiences of a binocular display (wearing and not wearing glasses) to CVD and normal-vision audiences. However, existing approaches only work for still images. Although state-of-the-art temporal filtering techniques can be applied to smooth the per-frame generated content, they may fail to maintain the multiple binocular constraints needed in our applications, and even worse, sometimes introduce color inconsistency (same color regions map to different colors). In this paper, we propose to train a neural network to predict the temporal coherent polynomial coefficients in the domain of global color decomposition. This indirect formulation solves the color inconsistency problem. Our key challenge is to design a neural network to predict the temporal coherent coefficients, while maintaining all required binocular constraints. Our method is evaluated on various videos and all metrics confirm that it outperforms all existing solutions. Xinghong Hu, Xueting Liu 0001, Zhuming Zhang, Menghan Xia, Chengze Li, Tien-Tsin Wong |
ACM Trans. Graph. | 6 |
| 2019 | Deep binocular tone mapping
Zhuming Zhang, Chu Han, Shengfeng He, Xueting Liu 0001, Xinghong Hu, Tien-Tsin Wong |
Vis. Comput. | 7 |
| 2018 | Deep Inverse Halftoning via Progressively Residual Learning
Menghan Xia, Tien-Tsin Wong |
ACCV (6) | 2 |
| 2018 | Binocular Tone Mapping with Improved Overall Contrast and Local DetailsabstractAbstract Tone mapping is a commonly used technique that maps the set of colors in high‐dynamic‐range (HDR) images to another set of colors in low‐dynamic‐range (LDR) images, to fit the need for print‐outs, LCD monitors and projectors. Unfortunately, during the compression of dynamic range, the overall contrast and local details generally cannot be preserved simultaneously. Recently, with the increased use of stereoscopic devices, the notion of binocular tone mapping has been proposed in the existing research study. However, the existing research lacks the binocular perception study and is unable to generate the optimal binocular pair that presents the most visual content. In this paper, we propose a novel perception‐based binocular tone mapping method, that can generate an optimal binocular image pair (generating left and right images simultaneously) from an HDR image that presents the most visual content by designing a binocular perception metric. Our method outperforms the existing method in terms of both visual and time performance. Zhuming Zhang, Xinghong Hu, Xueting Liu 0001, Tien-Tsin Wong |
Comput. Graph. Forum | 4 |
| 2018 | TransHist: Occlusion-robust shape detection in cluttered imagesabstractShape matching plays an important role in various computer vision and graphics applications such as shape retrieval, object detection, image editing, image retrieval, etc. However, detecting shapes in cluttered images is still quite challenging due to the incomplete edges and changing perspective. In this paper, we propose a novel approach that can efficiently identify a queried shape in a cluttered image. The core idea is to acquire the transformation from the queried shape to the cluttered image by summarising all point-to-point transformations between the queried shape and the image. To do so, we adopt a point-based shape descriptor, the pyramid of arc-length descriptor (PAD), to identify point pairs between the queried shape and the image having similar local shapes. We further calculate the transformations between the identified point pairs based on PAD. Finally, we summarise all transformations in a 4D transformation histogram and search for the main cluster. Our method can handle both closed shapes and open curves, and is resistant to partial occlusions. Experiments show that our method can robustly detect shapes in images in the presence of partial occlusions, fragile edges, and cluttered backgrounds. Chu Han, Xueting Liu 0001, Lok Tsun Sinn, Tien-Tsin Wong |
Comput. Vis. Media | 4 |
| 2018 | Deep unsupervised pixelizationabstractIn this paper, we present a novel unsupervised learning method for pixelization. Due to the difficulty in creating pixel art, preparing the paired training data for supervised learning is impractical. Instead, we propose an unsupervised learning framework to circumvent such difficulty. We leverage the dual nature of the pixelization and depixelization, and model these two tasks in the same network in a bi-directional manner with the input itself as training supervision. These two tasks are modeled as a cascaded network which consists of three stages for different purposes. GridNet transfers the input image into multi-scale grid-structured images with different aliasing effects. PixelNet associated with GridNet to synthesize pixel arts with sharp edges and perceptually optimal local structures. DepixelNet connects the previous network and aims to recover the pixelized result to the original image. For the sake of unsupervised learning, the mirror loss is proposed to hold the reversibility of feature representations in the process. In addition, adversarial, L1, and gradient losses are involved in the network to obtain pixel arts by retaining color correctness and smoothness. We show that our technique can synthesize crisper and perceptually more appropriate pixel arts than state-of-the-art image downscaling methods. We evaluate the proposed method with extensive experiments on many images. The proposed method outperforms state-of-the-art methods in terms of visual quality and user preference. Chu Han, Shengfeng He, Qianshu Zhu, Yinjie Tan, Guoqiang Han 0002, Tien-Tsin Wong |
ACM Trans. Graph. | 7 |
| 2018 | Invertible grayscaleabstractOnce a color image is converted to grayscale, it is a common belief that the original color cannot be fully restored, even with the state-of-the-art colorization methods. In this paper, we propose an innovative method to synthesize invertible grayscale. It is a grayscale image that can fully restore its original color. The key idea here is to encode the original color information into the synthesized grayscale, in a way that users cannot recognize any anomalies. We propose to learn and embed the color-encoding scheme via a convolutional neural network (CNN). It consists of an encoding network to convert a color image to grayscale, and a decoding network to invert the grayscale to color. We then design a loss function to ensure the trained network possesses three required properties: (a) color invertibility, (b) grayscale conformity, and (c) resistance to quantization error. We have conducted intensive quantitative experiments and user studies over a large amount of color images to validate the proposed method. Regardless of the genre and content of the color input, convincing results are obtained in all cases. Menghan Xia, Xueting Liu 0001, Tien-Tsin Wong |
ACM Trans. Graph. | 3 |
| 2018 | Two-stage sketch colorizationabstractSketch or line art colorization is a research field with significant market demand. Different from photo colorization which strongly relies on texture information, sketch colorization is more challenging as sketches may not have texture. Even worse, color, texture, and gradient have to be generated from the abstract sketch lines. In this paper, we propose a semi-automatic learning-based framework to colorize sketches with proper color, texture as well as gradient. Our framework consists of two stages. In the first drafting stage, our model guesses color regions and splashes a rich variety of colors over the sketch to obtain a color draft. In the second refinement stage, it detects the unnatural colors and artifacts, and try to fix and refine the result. Comparing to existing approaches, this two-stage design effectively divides the complex colorization task into two simpler and goal-clearer subtasks. This eases the learning and raises the quality of colorization. Our model resolves the artifacts such as water-color blurring, color distortion, and dull textures. We build an interactive software based on our model for evaluation. Users can iteratively edit and refine the colorization. We evaluate our learning model and the interactive system through an extensive user study. Statistics shows that our method outperforms the state-of-art techniques and industrial applications in several aspects including, the visual quality, the ability of user control, user experience, and other metrics. Lvmin Zhang, Chengze Li, Tien-Tsin Wong, Yi Ji 0001, Chunping Liu |
ACM Trans. Graph. | 3 |
| 2018 | Globally Consistent Wrinkle-Aware Shading of Line DrawingsabstractShading is a tedious process for artists involved in 2D cartoon and manga production given the volume of contents that the artists have to prepare regularly over tight schedule. While we can automate shading production with the presence of geometry, it is impractical for artists to model the geometry for every single drawing. In this work, we aim to automate shading generation by analyzing the local shapes, connections, and spatial arrangement of wrinkle strokes in a clean line drawing. By this, artists can focus more on the design rather than the tedious manual editing work, and experiment with different shading effects under different conditions. To achieve this, we have made three key technical contributions. First, we model five perceptual cues by exploring relevant psychological principles to estimate the local depth profile around strokes. Second, we formulate stroke interpretation as a global optimization model that simultaneously balances different interpretations suggested by the perceptual cues and minimizes the interpretation discrepancy. Lastly, we develop a wrinkle-aware inflation method to generate a height field for the surface to support the shading region computation. In particular, we enable the generation of two commonly-used shading styles: 3D-like soft shading and manga-style flat shading. Pradeep Kumar Jayaraman, Chi-Wing Fu, Jianmin Zheng, Xueting Liu 0001, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Packing Vertex Data into Hardware-Decompressible TexturesabstractMost graphics hardware features memory to store textures and vertex data for rendering. However, because of the irreversible trend of increasing complexity of scenes, rendering a scene can easily reach the limit of memory resources. Thus, vertex data are preferably compressed, with a requirement that they can be decompressed during rendering. In this paper, we present a novel method to exploit existing hardware texture compression circuits to facilitate the decompression of vertex data in graphics processing unit (GPUs). This built-in hardware allows real-time, random-order decoding of data. However, vertex data must be packed into textures, and careless packing arrangements can easily disrupt data coherence. Hence, we propose an optimization approach for the best vertex data permutation that minimizes compression error. All of these result in fast and high-quality vertex data decompression for real-time rendering. To further improve the visual quality, we introduce vertex clustering to reduce the dynamic range of data during quantization. Our experiments demonstrate the effectiveness of our method for various vertex data of 3D models during rendering with the advantages of a minimized memory footprint and high frame rate. Kin Chung Kwan, Xuemiao Xu, Tien-Tsin Wong, Wai-Man Pang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Boundary-aware texture region segmentation from mangaabstractDue to the lack of color in manga (Japanese comics), black-and-white textures are often used to enrich visual experience. With the rising need to digitize manga, segmenting texture regions from manga has become an indispensable basis for almost all manga processing, from vectorization to colorization. Unfortunately, such texture segmentation is not easy since textures in manga are composed of lines and exhibit similar features to structural lines (contour lines). So currently, texture segmentation is still manually performed, which is labor-intensive and time-consuming. To extract a texture region, various texture features have been proposed for measuring texture similarity, but precise boundaries cannot be achieved since boundary pixels exhibit different features from inner pixels. In this paper, we propose a novel method which also adopts texture features to estimate texture regions. Unlike existing methods, the estimated texture region is only regarded an initial, imprecise texture region. We expand the initial texture region to the precise boundary based on local smoothness via a graph-cut formulation. This allows our method to extract texture regions with precise boundaries. We have applied our method to various manga images and satisfactory results were achieved in all cases. Xueting Liu 0001, Chengze Li, Tien-Tsin Wong |
Comput. Vis. Media | 3 |
| 2017 | Deep extraction of manga structural linesabstractExtraction of structural lines from pattern-rich manga is a crucial step for migrating legacy manga to digital domain. Unfortunately, it is very challenging to distinguish structural lines from arbitrary, highly-structured, and black-and-white screen patterns. In this paper, we present a novel data-driven approach to identify structural lines out of pattern-rich manga, with no assumption on the patterns. The method is based on convolutional neural networks. To suit our purpose, we propose a deep network model to handle the large variety of screen patterns and raise output accuracy. We also develop an efficient and effective way to generate a rich set of training data pairs. Our method suppresses arbitrary screen patterns no matter whether these patterns are regular, irregular, tone-varying, or even pictorial, and regardless of their scales. It outputs clear and smooth structural lines even if these lines are contaminated by and immersed in complex patterns. We have evaluated our method on a large number of mangas of various drawing styles. Our method substantially outperforms state-of-the-art methods in terms of visual quality. We also demonstrate its potential in various manga applications, including manga colorization, manga retargeting, and 2.5D manga generation. Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
ACM Trans. Graph. | 3 |
| 2017 | ASCII Art Synthesis from Natural PhotographsabstractWhile ASCII art is a worldwide popular art form, automatic generating structure-based ASCII art from natural photographs remains challenging. The major challenge lies on extracting the perception-sensitive structure from the natural photographs so that a more concise ASCII art reproduction can be produced based on the structure. However, due to excessive amount of texture in natural photos, extracting perception-sensitive structure is not easy, especially when the structure may be weak and within the texture region. Besides, to fit different target text resolutions, the amount of the extracted structure should also be controllable. To tackle these challenges, we introduce a visual perception mechanism of non-classical receptive field modulation (non-CRF modulation) from physiological findings to this ASCII art application, and propose a new model of non-CRF modulation which can better separate the weak structure from the crowded texture, and also better control the scale of texture suppression. Thanks to our non-CRF model, more sensible ASCII art reproduction can be obtained. In addition, to produce more visually appealing ASCII arts, we propose a novel optimization scheme to obtain the optimal placement of proportional-font characters. We apply our method on a rich variety of images, and visually appealing ASCII art can be obtained in all cases. Xuemiao Xu, Linyuan Zhong, Minshan Xie, Xueting Liu 0001, Harry Qin, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2017 | Blue noise sampling using an N-body simulation-based method
Kin-Ming Wong, Tien-Tsin Wong |
Vis. Comput. | 2 |
| 2016 | Efficient Non-Consecutive Feature Tracking for Robust Structure-From-MotionabstractStructure-from-motion (SfM) largely relies on feature tracking. In image sequences, if disjointed tracks caused by objects moving in and out of the field of view, occasional occlusion, or image noise are not handled well, corresponding SfM could be affected. This problem becomes severer for large-scale scenes, which typically requires to capture multiple sequences to cover the whole scene. In this paper, we propose an efficient non-consecutive feature tracking framework to match interrupted tracks distributed in different subsequences or even in different videos. Our framework consists of steps of solving the feature "dropout" problem when indistinctive structures, noise or large image distortion exists, and of rapidly recognizing and joining common features located in different subsequences. In addition, we contribute an effective segment-based coarse-to-fine SfM algorithm for robustly handling large data sets. Experimental results on challenging video data demonstrate the effectiveness of the proposed system. Guofeng Zhang 0001, Haomin Liu, Zilong Dong, Jiaya Jia, Tien-Tsin Wong, Hujun Bao |
IEEE Trans. Image Process. | 5 |
| 2016 | Pyramid of arclength descriptor for generating collage of shapesabstractThis paper tackles a challenging 2D collage generation problem, focusing on shapes: we aim to fill a given region by packing irregular and reasonably-sized shapes with minimized gaps and overlaps. To achieve this nontrivial problem, we first have to analyze the boundary of individual shapes and then couple the shapes with partially-matched boundary to reduce gaps and overlaps in the collages. Second, the search space in identifying a good coupling of shapes is highly enormous, since arranging a shape in a collage involves a position, an orientation, and a scale factor. Yet, this matching step needs to be performed for every single shape when we pack it into a collage. Existing shape descriptors are simply infeasible for computation in a reasonable amount of time. To overcome this, we present a brand new, scale- and rotation-invariant 2D shape descriptor, namely pyramid of arclength descriptor (PAD). Its formulation is locally supported, scalable, and yet simple to construct and compute. These properties make PAD efficient for performing the partial-shape matching. Hence, we can prune away most search space with simple calculation, and efficiently identify candidate shapes. We evaluate our method using a large variety of shapes with different types and contours. Convincing collage results in terms of visual quality and time performance are obtained. Kin Chung Kwan, Lok Tsun Sinn, Chu Han, Tien-Tsin Wong, Chi-Wing Fu |
ACM Trans. Graph. | 4 |
| 2016 | Seamless visual sharing with color vision deficienciesabstractApproximately 250 million people suffer from color vision deficiency (CVD). They can hardly share the same visual content with normal-vision audiences. In this paper, we propose the first system that allows CVD and normal-vision audiences to share the same visual content simultaneously. The key that we can achieve this is because the ordinary stereoscopic display (non-autostereoscopic ones) offers users two visual experiences (with and without wearing stereoscopic glasses). By allocating one experience to CVD audiences and one to normal-vision audiences, we allow them to share. The core problem is to synthesize an image pair, that when they are presented binocularly, CVD audiences can distinguish the originally indistinguishable colors; and when it is in monocular presentation, normal-vision audiences cannot distinguish its difference from the original image. We solve the image-pair recoloring problem by optimizing an objective function that minimizes the color deviation for normal-vision audiences, and maximizes the color distinguishability and binocular fusibility for CVD audiences. Our method is extensively evaluated via multiple quantitative experiments and user studies. Convincing results are obtained in all our test cases. Wuyao Shen, Xinghong Hu, Tien-Tsin Wong |
ACM Trans. Graph. | 4 |
| 2016 | Globally optimal toon trackingabstractThe ability to identify objects or region correspondences between consecutive frames of a given hand-drawn animation sequence is an indispensable tool for automating animation modification tasks such as sequence-wide recoloring or shape-editing of a specific animated character. Existing correspondence identification methods heavily rely on appearance features, but these features alone are insufficient to reliably identify region correspondences when there exist occlusions or when two or more objects share similar appearances. To resolve the above problems, manual assistance is often required. In this paper, we propose a new correspondence identification method which considers both appearance features and motions of regions in a global manner. We formulate correspondence likelihoods between temporal region pairs as a network flow graph problem which can be solved by a well-established optimization algorithm. We have evaluated our method with various animation sequences and results show that our method consistently outperforms the state-of-the-art methods without any user guidance. Xueting Liu 0001, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 3 |
| 2016 | Text-aware balloon extraction from manga
Xueting Liu 0001, Chengze Li, Tien-Tsin Wong, Xuemiao Xu |
Vis. Comput. | 4 |
| 2015 | Optimization-Based Gradient Mesh Colour TransferabstractAbstract In vector graphics, gradient meshes represent an image object by one or more regularly connected grids. Every grid point has attributes as the position, colour and gradients of these quantities specified. Editing the attributes of an existing gradient mesh (such as the colour gradients) is not only non‐intuitive but also time‐consuming. To facilitate user‐friendly colour editing, we develop an optimization‐based colour transfer method for gradient meshes. The key idea is built on the fact that we can approximate a colour transfer operation on gradient meshes with a linear transfer function. In this paper, we formulate the approximation as an optimization problem, which aims to minimize the colour distribution of the example image and the transferred gradient mesh. By adding proper constraints, i.e. image gradients, to the optimization problem, the details of the gradient meshes can be better preserved. With the linear transfer function, we are able to edit the colours and colour gradients of the mesh points automatically, while preserving the structure of the gradient mesh. The experimental results show that our method can generate pleasing recoloured gradient meshes. Yi Xiao 0004, Andrew Chi-Sing Leung, Yukun Lai, Tien-Tsin Wong |
Comput. Graph. Forum | 5 |
| 2015 | Region-based structure line detection for cartoonsabstractCartoons are a worldwide popular visual entertainment medium with a long history. Nowadays, with the boom of electronic devices, there is an increasing need to digitize old classic cartoons as a basis for further editing, including deformation, colorization, etc. To perform such editing, it is essential to extract the structure lines within cartoon images. Traditional edge detection methods are mainly based on gradients. These methods perform poorly in the face of compression artifacts and spatially-varying line colors, which cause gradient values to become unreliable. This paper presents the first approach to extract structure lines in cartoons based on regions. Our method starts by segmenting an image into regions, and then classifies them as edge regions and non-edge regions. Our second main contribution comprises three measures to estimate the likelihood of a region being a non-edge region. These measure darkness, local contrast, and shape. Since the likelihoods become unreliable as regions become smaller, we further classify regions using both likelihoods and the relationships to neighboring regions via a graph-cut formulation. Our method has been evaluated on a wide variety of cartoon images, and convincing results are obtained in all cases. Xueting Liu 0001, Tien-Tsin Wong, Xuemiao Xu |
Comput. Vis. Media | 3 |
| 2015 | Closure-aware sketch simplificationabstractIn this paper, we propose a novel approach to simplify sketch drawings. The core problem is how to group sketchy strokes meaningfully, and this depends on how humans understand the sketches. The existing methods mainly rely on thresholding low-level geometric properties among the strokes, such as proximity, continuity and parallelism. However, it is not uncommon to have strokes with equal geometric properties but different semantics. The lack of semantic analysis will lead to the inability in differentiating the above semantically different scenarios. In this paper, we point out that, due to the gestalt phenomenon of closure , the grouping of strokes is actually highly influenced by the interpretation of regions. On the other hand, the interpretation of regions is also influenced by the interpretation of strokes since regions are formed and depicted by strokes. This is actually a chicken-or-the-egg dilemma and we solve it by an iterative cyclic refinement approach. Once the formed stroke groups are stabilized, we can simplify the sketchy strokes by replacing each stroke group with a smooth curve. We evaluate our method on a wide range of different sketch styles and semantically meaningful simplification results can be obtained in all test cases. Xueting Liu 0001, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 2 |
| 2015 | All-Frequency Direct Illumination with Vectorized VisibilityabstractMany existing pre-computed radiance transfer (PRT) approaches for all-frequency lighting store the information of a 3D object in the pre-vertex manner. To preserve the fidelity of high frequency effects, the 3D object must be tessellated densely. Otherwise, rendering artifacts due to interpolation may appear. This paper presents an all-frequency lighting algorithm for direct illumination based on a new visibility representation which approximates a visibility function using a sequence of 3D vectors. The algorithm is able to construct the visibility function of an on-screen pixel on-the-fly. Hence even though the 3D object is not tessellated densely, the rendering artifacts can be suppressed greatly. Besides, a summed area table based rendering algorithm, which is able to handle the integration over a non-axis aligned polygon, is developed. Using our approach, we can rotate lighting environment, change view point, and adjust the shininess of the 3D object in a real-time manner. Experimental results show that our approach can render plausible all-frequency lighting effects for direct illumination in real-time, especially for specular shadows, which are difficult for other methods to obtain. Tze-Yui Ho, Yi Xiao 0004, Ruibin Feng, Andrew Chi-Sing Leung, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | Concentric Spherical Representation for Omnidirectional Soft ShadowabstractAbstract Soft shadows play an important role in photo‐realistic rendering. Although there are many efficient soft shadow algorithms, most of them focus on the one‐side light source situation, where a planar light source is on the outside of the scene. In fact, in many situations, such as games, light sources are omnidirectional. They may be surrounded by a number of 3D objects. This paper proposes a soft shadow algorithm for the omnidirectional situation. We develop a concentric spherical representation to model the behaviour of omnidirectional light sources. To provide better rendering results, a novel summed area table based filtering scheme for spherical functions is proposed. In addition, we utilize unicube mapping, which samples the spherical space more uniformly, to further improve the filtering quality. Yi Xiao 0004, Andrew Chi-Sing Leung, Tze-Yui Ho, Tien-Tsin Wong |
Comput. Graph. Forum | 5 |
| 2013 | An improved parallel contrast-aware halftoningabstractDigital image halftoning is a widely used technique. However, achieving high fidelity tone reproduction and structural preservation with low computational time cost remains a challenging problem. This paper presents a highly parallel algorithm to boost real-time application of serial structure-preserving error diffusion. The contrast-aware halftoning approach is one such technique with superior structure preservation, but it offers only a limited opportunity for graphics processing unit (GPU) acceleration. Our method integrates contrast-aware halftoning into a new parallelizable error-diffusion halftoning framework. To eliminate visually disturbing artifacts resulting from parallelization, we propose a novel multiple quantization model and space-filling curve to maintain tone consistency, blue-noise property, and structure consistency. Our GPU implementation on a commodity personal computer achieves a real-time performance for a moderately sized image. We demonstrate the high quality and performance of the proposed approach with a variety of examples, and provide comparisons with state-of-the-art methods. Ling-yue Liu, Wei Chen 0001, Tien-Tsin Wong, Wenting Zheng, Wei-dong Geng |
J. Zhejiang Univ. Sci. C | 3 |
| 2013 | Parallel structure-aware halftoning
Huisi Wu, Tien-Tsin Wong, Pheng-Ann Heng |
Multim. Tools Appl. | 2 |
| 2013 | Example-Based Color Transfer for Gradient MeshesabstractEditing a photo-realistic gradient mesh is a tough task. Even only editing the colors of an existing gradient mesh can be exhaustive and time-consuming. To facilitate user-friendly color editing, we develop an example-based color transfer method for gradient meshes, which borrows the color characteristics of an example image to a gradient mesh. We start by exploiting the constraints of the gradient mesh, and accordingly propose a linear-operator-based color transfer framework. Our framework operates only on colors and color gradients of the mesh points and preserves the topological structure of the gradient mesh. Bearing the framework in mind, we build our approach on PCA-based color transfer. After relieving the color range problem, we incorporate a fusion-based optimization scheme to improve color similarity between the reference image and the recolored gradient mesh. Finally, a multi-swatch transfer scheme is provided to enable more user control. Our approach is simple, effective, and much faster than color transferring the rastered gradient mesh directly. The experimental results also show that our method can generate pleasing recolored gradient meshes. Yi Xiao 0004, Andrew Chi-Sing Leung, Yukun Lai, Tien-Tsin Wong |
IEEE Trans. Multim. | 5 |
| 2013 | Cube2Video: Navigate Between Cubic Panoramas in Real-TimeabstractOnline virtual navigation systems enable users to hop from one 360° panorama to another, which belong to a sparse point-to-point collection, resulting in a less pleasant viewing experience. In this paper, we present a novel method, namely Cube2Video, to support navigating between cubic panoramas in a video-viewing mode. Our method circumvents the intrinsic challenge of cubic panoramas, i.e., the discontinuities between cube faces, in an efficient way. The proposed method extends the matching-triangulation-interpolation procedure with special considerations of the spherical domain. A triangle-to-triangle homography-based warping is developed to achieve physically plausible and visually pleasant interpolation results. The temporal smoothness of the synthesized video sequence is improved by means of a compensation transformation. As experimental results demonstrate, our method can synthesize pleasant video sequences in real time, thus mimicking walking or driving navigation. Qiang Zhao 0005, Wei Feng 0005, Jiawan Zhang, Tien-Tsin Wong |
IEEE Trans. Multim. | 5 |
| 2013 | Stereoscopizing cel animationsabstractWhile hand-drawn cel animation is a world-wide popular form of art and entertainment, introducing stereoscopic effect into it remains difficult and costly, due to the lack of physical clues. In this paper, we propose a method to synthesize convincing stereoscopic cel animations from ordinary 2D inputs, without labor-intensive manual depth assignment nor 3D geometry reconstruction. It is mainly automatic due to the need of producing lengthy animation sequences, but with the option of allowing users to adjust or constrain all intermediate results. The system fits nicely into the existing production flow of cel animation. By utilizing the T-junction cue available in cartoons, we first infer the initial, but not reliable, ordering of regions. One of our major contributions is to resolve the temporal inconsistency of ordering by formulating it as a graph-cut problem. However, the resultant ordering remains insufficient for generating convincing stereoscopic effect, as ordering cannot be directly used for depth assignment due to its discontinuous nature. We further propose to synthesize the depth through an optimization process with the ordering formulated as constraints. This is our second major contribution. The optimized result is the spatiotemporally smooth depth for synthesizing stereoscopic effect. Our method has been evaluated on a wide range of cel animations and convincing stereoscopic effect is obtained in all cases. Xueting Liu 0001, Xuan S. Yang, Linling Zhang, Tien-Tsin Wong |
ACM Trans. Graph. | 5 |
| 2013 | Change Blindness ImagesabstractChange blindness refers to human inability to recognize large visual changes between images. In this paper, we present the first computational model of change blindness to quantify the degree of blindness between an image pair. It comprises a novel context-dependent saliency model and a measure of change, the former dependent on the site of the change, and the latter describing the amount of change. This saliency model in particular addresses the influence of background complexity, which plays an important role in the phenomenon of change blindness. Using the proposed computational model, we are able to synthesize changed images with desired degrees of blindness. User studies and comparisons to state-of-the-art saliency models demonstrate the effectiveness of our model. Li-Qian Ma, Kun Xu 0003, Tien-Tsin Wong, Bi-Ye Jiang, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2012 | Binocular tone mappingabstractBy extending from monocular displays to binocular displays, one additional image domain is introduced. Existing binocular display systems only utilize this additional image domain for stereopsis. Our human vision is not only able to fuse two displaced images, but also two images with difference in detail, contrast and luminance, up to a certain limit. This phenomenon is known as binocular single vision . Humans can perceive more visual content via binocular fusion than just a linear blending of two views. In this paper, we make a first attempt in computer graphics to utilize this human vision phenomenon, and propose a binocular tone mapping framework. The proposed framework generates a binocular low-dynamic range (LDR) image pair that preserves more human-perceivable visual content than a single LDR image using the additional image domain. Given a tone-mapped LDR image (left, without loss of generality), our framework optimally synthesizes its counterpart (right) in the image pair from the same source HDR image. The two LDR images are different, so that they can aggregately present more human-perceivable visual richness than a single arbitrary LDR image, without triggering visual discomfort . To achieve this goal, a novel binocular viewing comfort predictor (BVCP) is also proposed to prevent such visual discomfort. The design of BVCP is based on the findings in vision science. Through our user studies, we demonstrate the increase of human-perceivable visual richness and the effectiveness of the proposed BVCP in conservatively predicting the visual discomfort threshold of human observers. Xuan S. Yang, Linling Zhang, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 3 |
| 2012 | Statistical Invariance for Texture SynthesisabstractEstimating illumination and deformation fields on textures is essential for both analysis and application purposes. Traditional methods for such estimation usually require complicated and sometimes labor-intensive processing. In this paper, we propose a new perspective for this problem and suggest a novel statistical approach which is much simpler and more efficient. Our experiments show that many textures in daily life are statistically invariant in terms of colors and gradients. Variations of such statistics can be assumed to be influenced by illumination and deformation. This implies that we can inversely estimate the spatially varying illumination and deformation according to the variation of the texture statistics. This enables us to decompose a texture photo into an illumination field, a deformation field, and an implicit texture which are illumination- and deformation-free, within a short period of time, and with minimal user input. By processing and recombining these components, a variety of synthesis effects, such as exemplar preparation, texture replacement, surface relighting, as well as geometry modification, can be well achieved. Finally, convincing results are shown to demonstrate the effectiveness of the proposed method. Xiaopei Liu, Tien-Tsin Wong, Chi-Wing Fu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Conjoining Gestalt rules for abstraction of architectural drawingsabstractWe present a method for structural summarization and abstraction of complex spatial arrangements found in architectural drawings. The method is based on the well-known Gestalt rules, which summarize how forms, patterns, and semantics are perceived by humans from bits and pieces of geometric information. Although defining a computational model for each rule alone has been extensively studied, modeling a conjoint of Gestalt rules remains a challenge. In this work, we develop a computational framework which models Gestalt rules and more importantly, their complex interactions. We apply conjoining rules to line drawings, to detect groups of objects and repetitions that conform to Gestalt principles. We summarize and abstract such groups in ways that maintain structural semantics by displaying only a reduced number of repeated elements, or by replacing them with simpler shapes. We show an application of our method to line drawings of architectural models of various styles, and the potential of extending the technique to other computer-generated illustrations, and three-dimensional models. Liangliang Nan, Andrei Sharf, Ke Xie 0001, Tien-Tsin Wong, Oliver Deussen, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2011 | Making burr puzzles from 3D modelsabstractA 3D burr puzzle is a 3D model that consists of interlocking pieces with a single-key property. That is, when the puzzle is assembled, all the pieces are notched except one single key component which remains mobile. The intriguing property of the assembled burr puzzle is that it is stable, perfectly interlocked, without glue or screws, etc. Moreover, a burr puzzle consisting of a small number of pieces is still rather difficult to solve since the assembly must follow certain orders while the combinatorial complexity of the puzzle's piece arrangements is extremely high. In this paper, we generalize the 6-piece orthogonal burr puzzle (a knot) to design and model burr puzzles from 3D models. Given a 3D input model, we first interactively embed a network of knots into the 3D shape. Our method automatically optimizes and arranges the orientation of each knot, and modifies pieces of adjacent knots with an appropriate connection type. Then, following the geometry of the embedded pieces, the entire 3D model is partitioned by splitting the solid while respecting the assembly motion of embedded pieces. The main technical challenge is to enforce the single-key property and ensure the assembly/disassembly remains feasible, as the puzzle pieces in a network of knots are highly interlocked. Lastly, we also present an automated approach to generate the visualizations of the puzzle assembly process. Shi-Qing Xin, Chi-Fu William Lai, Chi-Wing Fu, Tien-Tsin Wong, Ying He 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 4 |
| 2011 | Unicube for Dynamic Environment MappingabstractCube mapping is widely used in many graphics applications due to the availability of hardware support. However, it does not sample the spherical surface evenly. Recently, a uniform spherical mapping, isocube mapping, was proposed. It exploits the six-face structure used in cube mapping and samples the spherical surface evenly. Unfortunately, some texels in isocube mapping are not rectilinear. This nonrectilinear property may degrade the filtering quality. This paper proposes a novel spherical mapping, namely unicube mapping. It has the advantages of cube mapping (exploitation of hardware and rectilinear structure) and isocube mapping (evenly sampling pattern). In the implementation, unicube mapping uses a simple function to modify the lookup vector before the conventional cube map lookup process. Hence, unicube mapping fully exploits the cube map hardware for real-time filtering and lookup. More importantly, its rectilinear partition structure allows a direct and real-time acquisition of the texture environment. This property facilitates dynamic environment mapping in a real time manner. Tze-Yui Ho, Andrew Chi-Sing Leung, Ping-Man Lam, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2011 | Spatiotemporal Sampling of Dynamic Environment SequencesabstractEnvironment sampling is a popular technique for rendering scenes with distant environment illumination. However, the temporal consistency of animations synthesized under dynamic environment sequences has not been fully studied. This paper addresses this problem and proposes a novel method, namely spatiotemporal sampling, to fully exploit both the temporal and spatial coherence of environment sequences. Our method treats an environment sequence as a spatiotemporal volume and samples the sequence by stratifying the volume adaptively. For this purpose, we first present a new metric to measure the importance of each stratified volume. A stratification algorithm is then proposed to adaptively suppress the abrupt temporal and spatial changes in the generated sampling patterns. The proposed method is able to automatically adjust the number of samples for each environment frame and produce temporally coherent sampling patterns. Comparative experiments demonstrate the capability of our method to produce smooth and consistent animations under dynamic environment sequences. Shue Kwan Mak, Tien-Tsin Wong, Andrew Chi-Sing Leung |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Motion Imitation with a Handheld CameraabstractIn this paper, we present a novel method to extract motion of a dynamic object from a video that is captured by a handheld camera, and apply it to a 3D character. Unlike the motion capture techniques, neither special sensors/trackers nor a controllable environment is required. Our system significantly automates motion imitation which is traditionally conducted by professional animators via manual keyframing. Given the input video sequence, we track the dynamic reference object to obtain trajectories of both 2D and 3D tracking points. With them as constraints, we then transfer the motion to the target 3D character by solving an optimization problem to maintain the motion gradients. We also provide a user-friendly editing environment for users to fine tune the motion details. As casual videos can be used, our system, therefore, greatly increases the supply source of motion data. Examples of imitating various types of animal motion are shown. Guofeng Zhang 0001, Hanqing Jiang, Jin Huang 0001, Jiaya Jia, Tien-Tsin Wong, Kun Zhou 0001, Hujun Bao |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2010 | Efficient Non-consecutive Feature Tracking for Structure-from-Motion
Guofeng Zhang 0001, Zilong Dong, Jiaya Jia, Tien-Tsin Wong, Hujun Bao |
ECCV (5) | 4 |
| 2010 | Structure-preserving multiscale vessel enhancing diffusion filterabstractEnhancement of vessels in medical images is still an unsolved problem. Multiscale approaches were proposed to improve the vessel enhancement effect based on the structure size and image resolution. Vessel enhancing diffusion (VED) filter is one of the multiscale approaches, which was based on the scale space theory. VED performs well on enhancing vessel structures but cannot preserve complex structures such as the vessel junctions. In this paper, a structure-preserving diffusion tensor is defined in the diffusion equation, which brings a structure-preserving vessel enhancing diffusion filter. Through the multiscale framework, the proposed method enhances the vessel structures especially the complex structure such as junctions. Experimental evaluation performed on various vessel data sets demonstrated the effectiveness of the proposed method. Yiping Chen 0002, Liansheng Wang 0002, Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Xiang Li 0014 |
ICIP | 6 |
| 2010 | Uniformly sampling multi-resolution analysis for image-based relighting
Ping-Man Lam, Andrew Chi-Sing Leung, Tien-Tsin Wong, Chi-Wing Fu |
J. Vis. Commun. Image Represent. | 3 |
| 2010 | Camouflage imagesabstractCamouflage images contain one or more hidden figures that remain imperceptible or unnoticed for a while. In one possible explanation, the ability to delay the perception of the hidden figures is attributed to the theory that human perception works in two main phases: feature search and conjunction search. Effective camouflage images make feature based recognition difficult, and thus force the recognition process to employ conjunction search, which takes considerable effort and time. In this paper, we present a technique for creating camouflage images. To foil the feature search, we remove the original subtle texture details of the hidden figures and replace them by that of the surrounding apparent image. To leave an appropriate degree of clues for the conjunction search, we compute and assign new tones to regions in the embedded figures by performing an optimization between two conflicting terms, which we call immersion and standout , corresponding to hiding and leaving clues, respectively. We show a large number of camouflage images generated by our technique, with or without user guidance. We have tested the quality of the images in an extensive user study, showing a good control of the difficulty levels. Hung-Kuo Chu, Wei-Hsin Hsu, Niloy J. Mitra, Daniel Cohen-Or, Tien-Tsin Wong, Tong-Yee Lee |
ACM Trans. Graph. | 5 |
| 2010 | Data-driven image color theme enhancementabstractIt is often important for designers and photographers to convey or enhance desired color themes in their work. A color theme is typically defined as a template of colors and an associated verbal description. This paper presents a data-driven method for enhancing a desired color theme in an image. We formulate our goal as a unified optimization that simultaneously considers a desired color theme, texture-color relationships as well as automatic or user-specified color constraints. Quantifying the difference between an image and a color theme is made possible by color mood spaces and a generalization of an additivity relationship for two-color combinations. We incorporate prior knowledge, such as texture-color relationships, extracted from a database of photographs to maintain a natural look of the edited images. Experiments and a user study have confirmed the effectiveness of our method. Baoyuan Wang, Yizhou Yu, Tien-Tsin Wong, Chun Chen 0001, Ying-Qing Xu |
ACM Trans. Graph. | 3 |
| 2010 | Resizing by symmetry-summarizationabstractImage resizing can be achieved more effectively if we have a better understanding of the image semantics. In this paper, we analyze the translational symmetry , which exists in many real-world images. By detecting the symmetric lattice in an image, we can summarize , instead of only distorting or cropping, the image content. This opens a new space for image resizing that allows us to manipulate, not only image pixels, but also the semantic cells in the lattice. As a general image contains both symmetry & non-symmetry regions and their natures are different, we propose to resize symmetry regions by summarization and non-symmetry region by warping. The difference in resizing strategy induces discontinuity at their shared boundary. We demonstrate how to reduce the artifact. To achieve practical resizing applications for general images, we developed a fast symmetry detection method that can detect multiple disjoint symmetry regions, even when the lattices are curved and perspectively viewed. Comparisons to state-of-the-art resizing techniques and a user study were conducted to validate the proposed method. Convincing visual results are shown to demonstrate its effectiveness. Huisi Wu, Yu-Shuen Wang, Kun-Chuan Feng, Tien-Tsin Wong, Tong-Yee Lee, Pheng-Ann Heng |
ACM Trans. Graph. | 4 |
| 2010 | Structure-based ASCII artabstractThe wide availability and popularity of text-based communication channels encourage the usage of ASCII art in representing images. Existing tone-based ASCII art generation methods lead to halftone-like results and require high text resolution for display, as higher text resolution offers more tone variety. This paper presents a novel method to generate structure-based ASCII art that is currently mostly created by hand. It approximates the major line structure of the reference image content with the shape of characters. Representing the unlimited image content with the extremely limited shapes and restrictive placement of characters makes this problem challenging. Most existing shape similarity metrics either fail to address the misalignment in real-world scenarios, or are unable to account for the differences in position, orientation and scaling. Our key contribution is a novel alignment-insensitive shape similarity (AISS) metric that tolerates misalignment of shapes while accounting for the differences in position, orientation and scaling. Together with the constrained deformation approach, we formulate the ASCII art generation as an optimization that minimizes shape dissimilarity and deformation . Convincing results and user study are shown to demonstrate its effectiveness. Xuemiao Xu, Linling Zhang, Tien-Tsin Wong |
ACM Trans. Graph. | 3 |
| 2010 | All-Frequency Lighting with Multiscale Spherical Radial Basis FunctionsabstractThis paper proposes a novel multiscale spherical radial basis function (MSRBF) representation for all-frequency lighting. It supports the illumination of distant environment as well as the local illumination commonly used in practical applications, such as games. The key is to define a multiscale and hierarchical structure of spherical radial basis functions (SRBFs) with basis functions uniformly distributed over the sphere. The basis functions are divided into multiple levels according to their coverage (widths). Within the same level, SRBFs have the same width. Larger width SRBFs are responsible for lower frequency lighting while the smaller width ones are responsible for the higher frequency lighting. Hence, our approach can achieve the true all-frequency lighting that is not achievable by the single-scale SRBF approach. Besides, the MSRBF approach is scalable as coarser rendering quality can be achieved without reestimating the coefficients from the raw data. With the homogeneous form of basis functions, the rendering is highly efficient. The practicability of the proposed method is demonstrated with real-time rendering and effective compression for tractable storage. Ping-Man Lam, Tze-Yui Ho, Andrew Chi-Sing Leung, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2010 | Evolving Mazes from ImagesabstractWe propose a novel reaction diffusion (RD) simulator to evolve image-resembling mazes. The evolved mazes faithfully preserve the salient interior structures in the source images. Since it is difficult to control the generation of desired patterns with traditional reaction diffusion, we develop our RD simulator on a different computational platform, cellular neural networks. Based on the proposed simulator, we can generate the mazes that exhibit both regular and organic appearance, with uniform and/or spatially varying passage spacing. Our simulator also provides high controllability of maze appearance. Users can directly and intuitively "paint" to modify the appearance of mazes in a spatially varying manner via a set of brushes. In addition, the evolutionary nature of our method naturally generates maze without any obvious seam even though the input image is a composite of multiple sources. The final maze is obtained by determining a solution path that follows the user-specified guiding curve. We validate our method by evolving several interesting mazes from different source images. Xiaopei Liu, Tien-Tsin Wong, Andrew Chi-Sing Leung |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2009 | Consistent Depth Maps Recovery from a Video SequenceabstractThis paper presents a novel method for recovering consistent depth maps from a video sequence. We propose a bundle optimization framework to address the major difficulties in stereo reconstruction, such as dealing with image noise, occlusions, and outliers. Different from the typical multi-view stereo methods, our approach not only imposes the photo-consistency constraint, but also explicitly associates the geometric coherence with multiple frames in a statistical way. It thus can naturally maintain the temporal coherence of the recovered dense depth maps without over-smoothing. To make the inference tractable, we introduce an iterative optimization scheme by first initializing the disparity maps using a segmentation prior and then refining the disparities by means of bundle optimization. Instead of defining the visibility parameters, our method implicitly models the reconstruction noise as well as the probabilistic visibility. After bundle optimization, we introduce an efficient space-time fusion algorithm to further reduce the reconstruction noise. Our automatic depth recovery is evaluated using a variety of challenging video examples. Guofeng Zhang 0001, Jiaya Jia, Tien-Tsin Wong, Hujun Bao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | The Rhombic Dodecahedron Map: An Efficient Scheme for Encoding Panoramic VideoabstractOmnidirectional videos are usually mapped to planar domain for encoding with off-the-shelf video compression standards. However, existing work typically neglects the effect of the sphere-to-plane mapping. In this paper, we show that by carefully designing the mapping, we can improve the visual quality, stability and compression efficiency of encoding omnidirectional videos. Here we propose a novel mapping scheme, known as the rhombic dodecahedron map (RD map) to represent data over the spherical domain. By using a family of skew great circles as the subdivision kernel, the RD map not only produces a sampling pattern with very low discrepancy, it can also support a highly efficient data indexing mechanism over the spherical domain. Since the proposed map is quad-based, geodesic-aligned, and of very low area and shape distortion, we can reliably apply 2-D wavelet-based and DCT-based encoding methods that are originally designated to planar perspective videos. At the end, we perform a series of analysis and experiments to investigate and verify the effectiveness of the proposed method; with its ultra-fast data indexing capability, we show that we can playback omnidirectional videos with very high frame rates on conventional PCs with GPU support. Chi-Wing Fu, Tien-Tsin Wong, Andrew Chi-Sing Leung |
IEEE Trans. Multim. | 3 |
| 2009 | Efficient Relighting of RBF-Based Illumination Adjustable ImagesabstractAn illumination adjustable image (IAI) contains a large number of prerecorded images under various light directions. Relighting a scene under complicated lighting conditions can be achieved from the IAI. Using the radial basis function (RBF) approach to represent an IAI is proven to be more efficient than using the spherical harmonic approach. However, to represent high-frequency lighting effects, we need to use many RBFs. Hence, the relighting speed could be very slow. This brief investigates a partial reconstruction scheme for relighting an IAI based on the locality of RBFs. Compared with the conventional RBF and spherical harmonics (SH) approaches, the proposed scheme has a much faster relighting speed under the similar distortion performance. Tze-Yui Ho, Andrew Chi-Sing Leung, Ping-Man Lam, Tien-Tsin Wong |
IEEE Trans. Neural Networks | 4 |
| 2009 | Refilming with Depth-Inferred VideosabstractCompared to still image editing, content-based video editing faces the additional challenges of maintaining the spatiotemporal consistency with respect to geometry. This brings up difficulties of seamlessly modifying video content, for instance, inserting or removing an object. In this paper, we present a new video editing system for creating spatiotemporally consistent and visually appealing refilming effects. Unlike the typical filming practice, our system requires no labor-intensive construction of 3D models/surfaces mimicking the real scene. Instead, it is based on an unsupervised inference of view-dependent depth maps for all video frames. We provide interactive tools requiring only a small amount of user input to perform elementary video content editing, such as separating video layers, completing background scene, and extracting moving objects. These tools can be utilized to produce a variety of visual effects in our system, including but not limited to video composition, "predator" effect, bullet-time, depth-of-field, and fog synthesis. Some of the effects can be achieved in real time. Guofeng Zhang 0001, Zilong Dong, Jiaya Jia, Tien-Tsin Wong, Hujun Bao |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2008 | Generating massive high-quality random numbers using GPUabstractPseudo-random number generators (PRNG) have been intensively used in many stochastic algorithms in artificial intelligence, computer graphics and other scientific computing. However, the current commodity GPU design does not facilitate the efficient implementation of high-quality PRNGs that require high-precision integer arithmetics and bitwise operations. In this paper, we propose a framework to generate a high-quality PRNG shader for all kinds of GPUs. We adopt the cellular automata (CA) PRNG to facilitate high speed and parallel random number generation. The configuration of the CA PRNG is completed automatically by optimizing an objective function that accounts for quality of generated random sequences. To visually evaluate the result, we apply the best PRNG shader to photon mapping. Timing statistics show that our GPU parallelized PRNG is much faster than a pure CPU implementation. Wai-Man Pang, Tien-Tsin Wong, Pheng-Ann Heng |
IEEE Congress on Evolutionary Computation | 2 |
| 2008 | Recovering consistent video depth maps via bundle optimizationabstractThis paper presents a novel method for reconstructing high-quality video depth maps. A bundle optimization model is proposed to address the key issues, including image noise and occlusions, in stereo reconstruction. Our method not only uses the color constancy constraint, but also explicitly incorporates the geometric coherence constraint associating multiple frames in a video, thus can naturally maintain the temporal coherence of the recovered video depths without introducing over-smoothing artifact. To make the inference problem tractable, we introduce an iterative optimization scheme by first initializing disparity maps using segmentation prior and then refining the disparities by means of bundle optimization. Unlike previous work estimating complex visibility parameters, our approach implicitly models the probabilistic visibility in a statistical way. The effectiveness of our automatic method is demonstrated using challenging video examples. Guofeng Zhang 0001, Jiaya Jia, Tien-Tsin Wong, Hujun Bao |
CVPR | 3 |
| 2008 | Volumetric Ultrasound Panorama Based on 3D SIFT
Dong Ni 0001, Yingge Qu, Xuan S. Yang, Yim-Pan Chui, Tien-Tsin Wong, Simon S. M. Ho, Pheng-Ann Heng |
MICCAI (2) | 5 |
| 2008 | Perceptual image preview
Wei Feng 0005, Zhouchen Lin, Tien-Tsin Wong |
Multim. Syst. | 4 |
| 2008 | Discriminative analysis of skull morphology in adolescent idiopathic scoliosis patients: Comparative study with normal controls
Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Winnie Chiu-Wing Chu, Benson H. Y. Yeung, Jack Chun-Yiu Cheng |
Pattern Recognit. | 4 |
| 2008 | Erratum to: "Discriminative analysis of skull morphology in adolescent idiopathic scoliosis patients: Comparative study with normal controls" [Pattern Recognition 41 (9) 2800-2811]
Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Winnie Chiu-Wing Chu, Benson H. Y. Yeung, Jack Chun-Yiu Cheng |
Pattern Recognit. | 4 |
| 2008 | A GPU favor representation method for plenoptic-illumination function based on an efficient spherical partition scheme
Andrew Chi-Sing Leung, Ping-Man Lam, Tien-Tsin Wong |
Signal Process. Image Commun. | 3 |
| 2008 | Self-animating images: illusory motion using repeated asymmetric patternsabstractIllusory motion in a still image is a fascinating research topic in the study of human motion perception. Physiologists and psychologists have attempted to understand this phenomenon by constructing simple, color repeated asymmetric patterns (RAP) and have found several useful rules to enhance the strength of illusory motion. Based on their knowledge, we propose a computational method to generate self-animating images. First, we present an optimized RAP placement on streamlines to generate illusory motion for a given static vector field. Next, a general coloring scheme for RAP is proposed to render streamlines. Furthermore, to enhance the strength of illusion and respect the shape of the region, a smooth vector field with opposite directional flow is automatically generated given an input image. Examples generated by our method are shown as evidence of the illusory effect and the potential applications for entertainment and design purposes. Ming-Te Chi, Tong-Yee Lee, Yingge Qu, Tien-Tsin Wong |
ACM Trans. Graph. | 4 |
| 2008 | Intrinsic colorizationabstractIn this paper, we present an example-based colorization technique robust to illumination differences between grayscale target and color reference images. To achieve this goal, our method performs color transfer in an illumination-independent domain that is relatively free of shadows and highlights. It first recovers an illumination-independent intrinsic reflectance image of the target scene from multiple color references obtained by web search. The reference images from the web search may be taken from different vantage points, under different illumination conditions, and with different cameras. Grayscale versions of these reference images are then used in decomposing the grayscale target image into its intrinsic reflectance and illumination components. We transfer color from the color reflectance image to the grayscale reflectance image, and obtain the final result by relighting with the illumination component of the target image. We demonstrate via several examples that our method generates results with excellent color consistency. Xiaopei Liu, Yingge Qu, Tien-Tsin Wong, Stephen Lin 0001, Andrew Chi-Sing Leung, Pheng-Ann Heng |
ACM Trans. Graph. | 4 |
| 2008 | Structure-aware halftoningabstractThis paper presents an optimization-based halftoning technique that preserves the structure and tone similarities between the original and the halftone images. By optimizing an objective function consisting of both the structure and the tone metrics, the generated halftone images preserve visually sensitive texture details as well as the local tone. It possesses the blue-noise property and does not introduce annoying patterns. Unlike the existing edge-enhancement halftoning, the proposed method does not suffer from the deficiencies of edge detector. Our method is tested on various types of images. In multiple experiments and the user study, our method consistently obtains the best scores among all tested methods. Wai-Man Pang, Yingge Qu, Tien-Tsin Wong, Daniel Cohen-Or, Pheng-Ann Heng |
ACM Trans. Graph. | 3 |
| 2008 | Richness-preserving manga screeningabstractDue to the tediousness and labor intensive cost, some manga artists have already employed computer-assisted methods for converting color photographs to manga backgrounds. However, existing bitonal image generation methods usually produce unsatisfactory uniform screening results that are not consistent with traditional mangas, in which the artist employs a rich set of screens. In this paper, we propose a novel method for generating bitonal manga backgrounds from color photographs. Our goal is to preserve the visual richness in the original photograph by utilizing not only screen density, but also the variety of screen patterns. To achieve the goal, we select screens for different regions in order to preserve the tone similarity, texture similarity, and chromaticity distinguishability. The multi-dimensional scaling technique is employed in such a color-to-pattern matching for maintaining pattern dissimilarity of the screens. Users can control the mapping by a few parameters and interactively fine-tune the result. Several results are presented to demonstrate the effectiveness and convenience of the proposed method. Yingge Qu, Wai-Man Pang, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 3 |
| 2008 | Animating animal motion from stillabstractEven though the temporal information is lost, a still picture of moving animals hints at their motion. In this paper, we infer motion cycle of animals from the "motion snapshots" (snapshots of different individuals) captured in a still picture. By finding the motion path in the graph connecting motion snapshots, we can infer the order of motion snapshots with respect to time, and hence the motion cycle. Both "half-cycle" and "full-cycle" motions can be inferred in a unified manner. Therefore, we can animate a still picture of a moving animal group by morphing among the ordered snapshots. By refining the pose, morphology, and appearance consistencies, smooth and realistic animal motion can be synthesized. Our results demonstrate the applicability of the proposed method to a wide range of species, including birds, fishes, mammals, and reptiles. Xuemiao Xu, Xiaopei Liu, Tien-Tsin Wong, Andrew Chi-Sing Leung |
ACM Trans. Graph. | 4 |
| 2007 | Robust Metric Reconstruction from Challenging Video SequencesabstractAlthough camera self-calibration and metric reconstruction have been extensively studied during the past decades, automatic metric reconstruction from long video sequences with varying focal length is still very challenging. Several critical issues in practical implementations are not adequately addressed. For example, how to select the initial frames for initializing the projective reconstruction? What criteria should be used? How to handle the large zooming problem? How to choose an appropriate moment for upgrading the projective reconstruction to a metric one? This paper gives a careful investigation of all these issues. Practical and effective approaches are proposed. In particular, we show that existing image-based distance is not an adequate measurement for selecting the initial frames. We propose a novel measurement to take into account the zoom degree, the self-calibration quality, as well as image-based distance. We then introduce a new strategy to decide when to upgrade the projective reconstruction to a metric one. Finally, to alleviate the heavy computational cost in the bundle adjustment, a local on-demand approach is proposed. Our method is also extensively compared with the state-of-the-art commercial software to evidence its robustness and stability. Guofeng Zhang 0001, Xueying Qin, Wei Hua 0002, Tien-Tsin Wong, Pheng-Ann Heng, Hujun Bao |
CVPR | 4 |
| 2007 | Moving Object Extraction with a Hand-held CameraabstractThis paper presents a new method to detect and accurately extract the moving object from a video sequence taken by a hand-held camera. In order to extract the high quality moving foreground, previous approaches usually assume that the background is static or through only planar-perspective transformation. In our method, based on the robust motion estimation, we are capable of handling challenging videos where the background contains complex depth and the camera undergoes unknown motions. We propose the appearance and structure consistency constraint in 3D warping to robustly model the background, which greatly improves the foreground separation even on the object boundary. The estimated dense motion field and the bi- layer segmentation result are iteratively refined where continuous and discrete optimizations are alternatively used. Experimental results of high quality moving object extraction from challenging videos demonstrate the effectiveness of our method. Guofeng Zhang 0001, Jiaya Jia, Tien-Tsin Wong, Pheng-Ann Heng, Hujun Bao |
ICCV | 4 |
| 2007 | Analysis on Bidirectional Associative Memories with Multiplicative Weight Noise
Andrew Chi-Sing Leung, John Sum, Tien-Tsin Wong |
ICONIP (1) | 3 |
| 2007 | Orthopedics Surgery Trainer with PPU-Accelerated Blood and Tissue Simulation
Wai-Man Pang, Harry Qin, Yim-Pan Chui, Tien-Tsin Wong, Kwok-Sui Leung, Pheng-Ann Heng |
MICCAI (2) | 4 |
| 2007 | Landmark Correspondence Optimization for Coupled Surfaces
Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Winnie Chiu-Wing Chu, Benson H. Y. Yeung, Jack Chun-Yiu Cheng |
MICCAI (2) | 4 |
| 2007 | Interactive Reaction-Diffusion on Surface TilesabstractThis paper proposes to perform reaction-diffusion on surface tiles. The square tiles fit nicely and cost-effectively in GPU memory, whereas we also apply distortion minimization on tiles so as to precisely reduce the unbalanced scale and resolution problem of chemicals in the reaction- diffusion. The interconnection nature of tiles accounts for the surface topology, and thus allows the chemicals to flow naturally over surfaces of arbitrary genus. Furthermore, by taking advantage of the tile structure, we can efficiently perform localized reaction-diffusion, and adjust the pattern formulation in an interactive manner on the GPU. To demonstrate its performance, we develop an interactive system that allows texture designers to fine-tune and alter the reaction-diffusion process by directly painting chemicals onto the object surface. Finally, we also develop several non-trivial applications of reaction-diffusion, including the geometry-dependent reaction-diffusion and deformation-aware reaction-diffusion. Kui-Yip Lo, Hongwei Li 0004, Chi-Wing Fu, Tien-Tsin Wong |
PG | 4 |
| 2007 | Noise-proofing the doubly SH-projected coefficients for synthesizing images under environment lighting
Ping-Man Lam, Tze-Yui Ho, Andrew Chi-Sing Leung, Tien-Tsin Wong |
Signal Process. Image Commun. | 4 |
| 2007 | Discrete Wavelet Transform on Consumer-Level Graphics HardwareabstractDiscrete wavelet transform (DWT) has been heavily studied and developed in various scientific and engineering fields. Its multiresolution and locality nature facilitates applications requiring progressiveness and capturing high-frequency details. However, when dealing with enormous data volume, its performance may drastically reduce. On the other hand, with the recent advances in consumer-level graphics hardware, personal computers nowadays usually equip with a graphics processing unit (GPU) based graphics accelerator which offers SIMD-based parallel processing power. This paper presents a SIMD algorithm that performs the convolution-based DWT completely on a GPU, which brings us significant performance gain on a normal PC without extra cost. Although the forward and inverse wavelet transforms are mathematically different, the proposed algorithm unifies them to an almost identical process that can be efficiently implemented on GPU. Different wavelet kernels and boundary extension schemes can be easily incorporated by simply modifying input parameters. To demonstrate its applicability and performance, we apply it to wavelet-based geometric design, stylized image processing, texture-illuminance decoupling, and JPEG2000 image encoding Tien-Tsin Wong, Andrew Chi-Sing Leung, Pheng-Ann Heng, Jianqing Wang |
IEEE Trans. Multim. | 1 |
| 2007 | Solid texture synthesis from 2D exemplarsabstractWe present a novel method for synthesizing solid textures from 2D texture exemplars. First, we extend 2D texture optimization techniques to synthesize 3D texture solids. Next, the non-parametric texture optimization approach is integrated with histogram matching, which forces the global statistics of the synthesized solid to match those of the exemplar. This improves the convergence of the synthesis process and enables using smaller neighborhoods. In addition to producing compelling texture mapped surfaces, our method also effectively models the material in the interior of solid objects. We also demonstrate that our method is well-suited for synthesizing textures with a large number of channels per texel. Johannes Kopf 0001, Chi-Wing Fu, Daniel Cohen-Or, Oliver Deussen, Dani Lischinski, Tien-Tsin Wong |
ACM Trans. Graph. | 6 |
| 2007 | Tileable BTFabstractThis paper presents a modular framework to efficiently apply the bidirectional texture functions (BTF) onto object surfaces. The basic building blocks are the BTF tiles. By constructing one set of BTF tiles, a wide variety of objects can be textured seamlessly without re-synthesizing the BTF. The proposed framework nicely decouples the surface appearance from the geometry. With this appearance-geometry decoupling, one can build a library of BTF tile sets to instantaneously dress and render various objects under variable lighting and viewing conditions. The core of our framework is a novel method for synthesizing seamless high-dimensional BTF tiles, that are difficult for existing synthesis techniques. Its key is to shorten the cutting paths and broaden the choices of samples so as to increase the chance of synthesizing seamless BTF tiles. To tackle the enormous data, the tile synthesis process is performed in compressed domain. This not just allows the handling of large BTF data during the synthesis, but also facilitates compact storage of the BTF in GPU memory during the rendering. Man-Kang Leung, Wai-Man Pang, Chi-Wing Fu, Tien-Tsin Wong, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2007 | Isocube: Exploiting the Cubemap HardwareabstractThis paper proposes a novel six-face spherical map, isocube, that fully utilizes the cubemap hardware built in most GPUs. Unlike the cubemap, the proposed isocube uniformly samples the unit sphere (uniformly distributed), and all samples span the same solid angle (equally important). Its mapping computation contains only a small overhead. By feeding the cubemap hardware with the six-face isocube map, the isocube can exploit all built-in texturing operators tailored for the cubemap and achieve a very high frame rate. In addition, we develop an anisotropic filtering that compensates aliasing artifacts due to texture magnification. This filtering technique extends the existing hardware anisotropic filtering and can be applied not only to the proposed isocube, but also to other texture mapping applications. Tien-Tsin Wong, Andrew Chi-Sing Leung |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2007 | Stereoscopic Video Synthesis from a Monocular VideoabstractThis paper presents an automatic and robust approach to synthesize stereoscopic videos from ordinary monocular videos acquired by commodity video cameras. Instead of recovering the depth map, the proposed method synthesizes the binocular parallax in stereoscopic video directly from the motion parallax in monocular video. The synthesis is formulated as an optimization problem via introducing a cost function of the stereoscopic effects, the similarity, and the smoothness constraints. The optimization selects the most suitable frames in the input video for generating the stereoscopic video frames. With the optimized selection, convincing and smooth stereoscopic video can be synthesized even by simple constant-depth warping. No user interaction is required. We demonstrate the visually plausible results obtained given the input clips acquired by ordinary handheld video camera. Guofeng Zhang 0001, Wei Hua 0002, Xueying Qin, Tien-Tsin Wong, Hujun Bao |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2006 | Parallel Hybrid Genetic Algorithms on Consumer-Level Graphics HardwareabstractIn this paper, we report a parallel Hybrid Genetic Algorithm (HGA) on consumer-level graphics cards. HGA extends the classical genetic algorithm by incorporating the Cauchy mutation operator from evolutionary programming. In our parallel HGA, all steps except the random number generation procedure are performed in Graphics Processing Unit (GPU) and thus our parallel HGA can be executed effectively and efficiently. We propose the pseudo-deterministic selection method which is comparable to the traditional global selection approach with significant execution time performance advantages. We perform experiments to compare our parallel HGA with our previous parallel FEP (Fast Evolutionary programming) and demonstrate that the former is much more effective and efficient than the latter. The parallel and sequential implementations of HGA are compared in a number of experiments, it is observed that the former outperforms the latter significantly. The effectiveness and efficiency of the pseudo-deterministic selection method is also studied. Man Leung Wong, Tien-Tsin Wong |
IEEE Congress on Evolutionary Computation | 2 |
| 2006 | Morphometric Analysis for Pathological Abnormality Detection in the Skull Vaults of Adolescent Idiopathic Scoliosis Girls
Lin Shi 0001, Pheng-Ann Heng, Tien-Tsin Wong, Winnie Chiu-Wing Chu, Benson H. Y. Yeung, Jack Chun-Yiu Cheng |
MICCAI (1) | 3 |
| 2006 | GPU-friendly warped display for scope-maintained video surveillance
Tien-Tsin Wong, Pheng-Ann Heng |
Multim. Syst. | 2 |
| 2006 | Dense Photometric Stereo: A Markov Random Field ApproachabstractWe address the problem of robust normal reconstruction by dense photometric stereo, in the presence of complex geometry, shadows, highlight, transparencies, variable attenuation in light intensities, and inaccurate estimation in light directions. The input is a dense set of noisy photometric images, conveniently captured by using a very simple set-up consisting of a digital video camera, a reflective mirror sphere, and a handheld spotlight. We formulate the dense photometric stereo problem as a Markov network and investigate two important inference algorithms for Markov Random Fields (MRFs)--graph cuts and belief propagation--to optimize for the most likely setting for each node in the network. In the graph cut algorithm, the MRF formulation is translated into one of energy minimization. A discontinuity-preserving metric is introduced as the compatibility function, which allows alpha-expansion to efficiently perform the maximum a posteriori (MAP) estimation. Using the identical dense input and the same MRF formulation, our tensor belief propagation algorithm recovers faithful normal directions, preserves underlying discontinuities, improves the normal estimation from one of discrete to continuous, and drastically reduces the storage requirement and running time. Both algorithms produce comparable and very faithful normals for complex scenes. Although the discontinuity-preserving metric in graph cuts permits efficient inference of optimal discrete labels with a theoretical guarantee, our estimation algorithm using tensor belief propagation converges to comparable results, but runs faster because very compact messages are passed and combined. We present very encouraging results on normal reconstruction. A simple algorithm is proposed to reconstruct a surface from a normal map recovered by our method. With the reconstructed surface, an inverse process, known as relighting in computer graphics, is proposed to synthesize novel images of the given scene under user-specified light source and direction. The synthesis is made to run in real time by exploiting the state-of-the-art graphics processing unit (GPU). Our method offers many unique advantages over previous relighting methods and can handle a wide range of novel light sources and directions. Tai-Pang Wu, Kam-Lun Tang, Chi-Keung Tang, Tien-Tsin Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | GPU-friendly rendering for illumination adjustable images
Tze-Yui Ho, Ping-Man Lam, Andrew Chi-Sing Leung, Tien-Tsin Wong |
Signal Process. Image Commun. | 4 |
| 2006 | An RBF-based compression method for image-based relightingabstractIn image-based relighting, a pixel is associated with a number of sampled radiance values. This paper presents a two-level compression method. In the first level, the plenoptic property of a pixel is approximated by a spherical radial basis function (SRBF) network. That means that the spherical plenoptic function of each pixel is represented by a number of SRBF weights. In the second level, we apply a wavelet-based method to compress these SRBF weights. To reduce the visual artifact due to quantization noise, we develop a constrained method for estimating the SRBF weights. Our proposed approach is superior to JPEG, JPEG2000, and MPEG. Compared with the spherical harmonics approach, our approach has a lower complexity, while the visual quality is comparable. The real-time rendering method for our SRBF representation is also discussed. Andrew Chi-Sing Leung, Tien-Tsin Wong, Ping-Man Lam, Kwok-Hung Choy |
IEEE Trans. Image Process. | 2 |
| 2006 | Intelligent Inferencing and Haptic Simulation for Chinese Acupuncture Learning and TrainingabstractThis paper presents an intelligent virtual environment for Chinese acupuncture learning and training using state-of-the-art virtual reality technology. It is the first step toward developing a comprehensive virtual human model for studying Chinese medicine. Students can learn and practice acupuncture in the proposed 3-D interactive virtual environment that supports a force feedback interface for needle insertion. Thus, students not only "see" but also "touch" the virtual patient. With high performance computers, highly informative and flexible visualization of acupuncture points of various related meridian and collateral can be highlighted to guide the students during training. A computer-based expert system using our newly proposed intelligent fuzzy petri net is designed and implemented to train the students to treat different diseases using acupuncture. Such an intelligent virtual reality system can provide an interesting and effective learning environment for Chinese acupuncture. Pheng-Ann Heng, Tien-Tsin Wong, Rong Yang 0006, Yim-Pan Chui, Yongming Xie, Kwong-Sak Leung, P.-C. Leung |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2006 | Manga colorizationabstractThis paper proposes a novel colorization technique that propagates color over regions exhibiting pattern-continuity as well as intensity-continuity. The proposed method works effectively on colorizing black-and-white manga which contains intensive amount of strokes, hatching, halftoning and screening. Such fine details and discontinuities in intensity introduce many difficulties to intensity-based colorization methods. Once the user scribbles on the drawing, a local, statistical based pattern feature obtained with Gabor wavelet filters is applied to measure the pattern-continuity. The boundary is then propagated by the level set method that monitors the pattern-continuity. Regions with open boundaries or multiple disjointed regions with similar patterns can be sensibly segmented by a single scribble. With the segmented regions, various colorization techniques can be applied to replace colors, colorize with stroke preservation, or even convert pattern to shading. Several results are shown to demonstrate the effectiveness and convenience of the proposed method. Yingge Qu, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 2 |
| 2006 | Deringing cartoons by image analogiesabstractIn this article, we propose a novel method to reduce ringing artifacts in BDCT-encoded cartoon images using image analogies. The quantization procedure of BDCT compression (such as JPEG and MPEG) introduces annoying visual artifacts. Our main focus is on the removal of ringing artifacts that is seldom addressed by existing methods. In the proposed method, the contaminated image is modeled as a Markov random field (MRF). We “learn” the behavior of contamination by extracting massive numbers of artifact patterns from a training set, and organizing them using tree-structured vector quantization (TSVQ). Instead of postfiltering the input contaminated image, we synthesize an artifact-reduced image. Our method is noniterative and hence, can remove artifacts within a very short period of time. We show that substantial improvement is achieved using the proposed method in terms of visual quality and statistics. Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 2 |
| 2006 | Noise-Resistant Fitting for Spherical HarmonicsabstractSpherical harmonic (SH) basis functions have been widely used for representing spherical functions in modeling various illumination properties. They can compactly represent low-frequency spherical functions. However, when the unconstrained least square method is used for estimating the SH coefficients of a hemispherical function, the magnitude of these SH coefficients could be very large. Hence, the rendering result is very sensitive to quantization noise (introduced by modern texture compression like S3TC, IEEE half float data type on GPU, or other lossy compression methods) in these SH coefficients. Our experiments show that, as the precision of SH coefficients is reduced, the rendered images may exhibit annoying visual artifacts. To reduce the noise sensitivity of the SH coefficients, this paper first discusses how the magnitude of SH coefficients affects the rendering result when there is quantization noise. Then, two fast fitting methods for estimating the noise-resistant SH coefficients are proposed. They can effectively control the magnitude of the estimated SH coefficients and, hence, suppress the rendering artifacts. Both statistical and visual results confirm our theory. Ping-Man Lam, Andrew Chi-Sing Leung, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2005 | Parallel evolutionary algorithms on graphics processing unitabstractEvolutionary algorithms (EAs) are effective and robust methods for solving many practical problems such as feature selection, electrical circuit synthesis, and data mining. However, they may execute for a long time for some difficult problems, because several fitness evaluations must be performed. A promising approach to overcome this limitation is to parallelize these algorithms. In this paper, we propose to implement a parallel EA on consumer-level graphics cards. We perform experiments to compare our parallel EA with an ordinary EA and demonstrate that the former is much more effective than the latter. Since consumer-level graphics cards are available in ubiquitous personal computers and these computers are easy to use and manage, more people are able to use our parallel algorithm to solve their problems encountered in real-world applications. Man Leung Wong, Tien-Tsin Wong, Ka-Ling Fok |
Congress on Evolutionary Computation | 2 |
| 2005 | Dense Photometric Stereo Using Tensorial Belief PropagationabstractWe address the normal reconstruction problem by photometric stereo using a uniform and dense set of photometric images captured at fixed viewpoint. Our method is robust to spurious noises caused by highlight and shadows and non-Lambertian reflections. To simultaneously recover normal orientations and preserve discontinuities, we model the dense photometric stereo problem into two coupled Markov random fields (MRFs): a smooth field for normal orientations, and a spatial line process for normal orientation discontinuities. We propose a very fast tensorial belief propagation method to approximate the maximum a posteriori (MAP) solution of the Markov network. Our tensor-based message passing scheme not only improves the normal orientation estimation from one of discrete to continuous, but also reduces storage and running time drastically. A convenient handheld device was built to collect a scattered set of photometric samples, from which a dense and uniform set on the lighting direction sphere is obtained. We present very encouraging results on a wide range of difficult objects to show the efficacy of our approach. Kam-Lun Tang, Chi-Keung Tang, Tien-Tsin Wong |
CVPR (1) | 3 |
| 2005 | A Markov Random Field Approach for Dense Photometric StereoabstractWe present a surprisingly simple system that allows for robust normal reconstruction by photometric stereo using a uniform and dense set of photometric images captured at fixed viewpoint, in the presense of spurious noises caused by highlight, shadows and non-Lambertian reflections. Our system consists of a mirror sphere, a spotlight and a DV camera only. Using this, a dense set of unbiased but noisy photometric data that roughly distributed uniformly on the light direction sphere is produced. To simultaneously recover normal orientations and preserve discontinuities, we model the dense photometric stereo problem into two coupled Markov random fields (MRFs): a smooth field for normal orientations, and a spatial line process for normal orientation discontinuities. A very fast tensorial belief propagation method is used to approximate the maximum a posteriori (MAP) solution of the Markov network. We present very encouraging results on a wide range of difficult objects to show the efficacy of our approach. Kam-Lun Tang, Chi-Keung Tang, Tien-Tsin Wong |
CVPR (2) | 3 |
| 2005 | Support Vector Clustering for Brain Activation Detection
Defeng Wang, Lin Shi 0001, Daniel S. Yeung, Pheng-Ann Heng, Tien-Tsin Wong, Eric C. C. Tsang |
MICCAI | 5 |
| 2005 | Spherical Q2-tree for Sampling Dynamic Environment Sequences
Tien-Tsin Wong, Andrew Chi-Sing Leung |
Rendering Techniques | 2 |
| 2005 | Compressing the illumination-adjustable images with principal component analysisabstractThe ability to change illumination is a crucial factor in image-based modeling and rendering. Image-based relighting offers such capability. However, the tradeoff is the enormous increase of storage requirement. In this paper, we propose a compression scheme that effectively reduces the data volume while maintaining the real-time relighting capability. The proposed method is based on principal component analysis (PCA). A block-wise PCA is used to practically process the huge input data. The output of PCA is a set of eigenimages and the corresponding relighting coefficients. By dropping those low-energy eigenimages, the data size is drastically reduced. To further compress the data, eigenimages left are compressed using transform coding and quantization while the relighting coefficients are compressed using uniform quantization. We also suggest the suitable target bit rate for each phase of the compression method in order to preserve the visual quality. Finally, we propose a real-time engine that relights images from the compressed data. Pun-Mo Ho, Tien-Tsin Wong, Andrew Chi-Sing Leung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | A system for real-time panorama generation and display in tele-immersive applicationsabstractWide field-of-view (FOV) is necessary for many industrial applications, such as air traffic control, large vehicle driving and navigation. Unfortunately, the supporting structure/frame in most systems usually blocks part of the view, results in "blind spot" and raises the risk to the pilot. In this paper, we introduce a video-based tele-immersive system, called the immersive cockpit. It captures live videos from the working site and recreates an immersive environment at the remote site where the pilot situates. It immerses the pilot at the remote site with a panoramic view of the environment, and hence improves interactivity and safety. The design goals of our system are real-time, live, low-cost, and scalable. We stitch multiple video streams captured from ordinary charged couple device cameras to generate a panoramic video. To avoid being blocked by the supporting frame, we allow a flexible placement of cameras. This approach trades the accuracy of the generated panoramic image for a larger FOV. To reduce the computation, parameters for stitching are determined once during the system initialization. The panoramic video is presented on an immersive display which covers the FOV of the viewer. We discuss how to correctly present the panoramic video on this nonplanar immersive display screen by sweet spot relocation. We also present the result and the performance evaluation of the system. Wai-Kwan Tang, Tien-Tsin Wong, Pheng-Ann Heng |
IEEE Trans. Multim. | 2 |
| 2005 | Visual simulation of weathering by gamma-ton tracingabstractWeathering modeling introduces blemishes such as dirt, rust, cracks and scratches to virtual scenery. In this paper we present a visual stimulation technique that works well for a wide variety of weathering phenomena. Our technique, called γ-ton tracing, is based on a type of aging-inducing particles called γ-tons. Modeling a weathering effect with γ-ton tracing involves tracing a large number of γ-tons through the scene in a way similar to photon tracing and then generating the weathering effect using the recorded γ-ton transport information. With this technique, we can produce weathering effects that are customized to the scene geometry and tailored to the weathering sources. Several effects that are challenging for existing techniques can be readily captured by γ-ton tracing. These include global transport effects. or "stainbleeding". γ-ton tracing also enables visual simulations of complex multi-weathering effects. Lastly γ-ton tracing can generate weathering effects that not only involve texture changes but also large-scale geometry changes. We demonstrate our technique with a variety of examples. Yanyun Chen, Lin Xia, Tien-Tsin Wong, Xin Tong 0001, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2004 | A real-time Cantonese text-to-audiovisual speech synthesizerabstractThis paper describes the design and development of a Cantonese TTVS synthesizer, which can generate highly natural synthetic speech that is precisely time-synchronized with a real-time 3D face rendering. Our Cantonese TTVS synthesizer utilizes a homegrown Cantonese syllable-based concatenative text-to-speech system named CU VOCAL. This paper describes the extension of CU VOCAL to output syllable labels and durations that correspond to the output acoustic wave file. The syllables are decomposed and their initials/finals are mapped to the nearest IPA symbols that correspond to static viseme models. We have authored sixteen static viseme models together with two emotion-based face models. In order to achieve 3D face rendering, we have designed and implemented a blending technique that computes the linear combinations of the static face models to effect smooth transitions in between models. We demonstrate that this design and implementation of a TTVS synthesizer can achieve real-time performance in generation. Jianqing Wang, Ka-Ho Wong, Pheng-Ann Heng, Helen M. Meng, Tien-Tsin Wong |
ICASSP (1) | 5 |
| 2004 | Segmentation of Left Ventricle via Level Set Method Based on Enriched Speed Term
Yingge Qu, Qiang Chen 0004, Pheng-Ann Heng, Tien-Tsin Wong |
MICCAI (1) | 4 |
| 2004 | Digital photo similarity analysis in frequency domain and photo album compressionabstractWith the increasing popularity of digital camera, organizing and managing the large collection of digital photos effectively are therefore required. In this paper, we study the techniques of photo album sorting, clustering and compression in DCT frequency domain without having to decompress JPEG photos into spatial domain firstly. We utilize the first several non-zero DCT coefficients to build our feature set and calculate the energy histograms in frequency domain directly. We then calculate the similarity distances of every two photos, and perform photo album sorting and adaptive clustering algorithms to group the most similar photos together. We further compress those clustered photos by a MPEG-like algorithm with variable IBP frames and adaptive search windows. Our methods provide a compact and reasonable format for people to store and transmit their large number of digital photos. Experiments prove that our algorithm is efficient and effective for digital photo processing. Tien-Tsin Wong, Pheng-Ann Heng |
MUM | 2 |
| 2004 | Data compression with spherical wavelets and wavelets for the image-based relighting
Ze Wang 0018, Andrew Chi-Sing Leung, Yisheng Zhu, Tien-Tsin Wong |
Comput. Vis. Image Underst. | 4 |
| 2004 | Eigen-image based compression for the image-based relighting with cascade recursive least squared networks
Ze Wang 0018, Andrew Chi-Sing Leung, Tien-Tsin Wong, Yisheng Zhu |
Pattern Recognit. | 3 |
| 2004 | A compression method for a massive image data set in image-based rendering
Ping-Man Lam, Andrew Chi-Sing Leung, Tien-Tsin Wong |
Signal Process. Image Commun. | 3 |
| 2004 | Data compression on the illumination adjustable images by PCA and ICA
Ze Wang 0018, Andrew Chi-Sing Leung, Yisheng Zhu, Tien-Tsin Wong |
Signal Process. Image Commun. | 4 |
| 2004 | Binary-Space-Partitioned Images for Resolving Image-Based VisibilityabstractWe propose a novel 2D representation for 3D visibility sorting, the Binary-Space-Partitioned Image (BSPI), to accelerate real-time image-based rendering. BSPI is an efficient 2D realization of a 3D BSP tree, which is commonly used in computer graphics for time-critical visibility sorting. Since the overall structure of a BSP tree is encoded in a BSPI, traversing a BSPI is comparable to traversing the corresponding BSP tree. BSPI performs visibility sorting efficiently and accurately in the 2D image space by warping the reference image triangle-by-triangle instead of pixel-by-pixel. Multiple BSPIs can be combined to solve "disocclusion," when an occluded portion of the scene becomes visible at a novel viewpoint. Our method is highly automatic, including a tensor voting preprocessing step that generates candidate image partition lines for BSPIs, filters the noisy input data by rejecting outliers, and interpolates missing information. Our system has been applied to a variety of real data, including stereo, motion, and range images. Chi-Wing Fu, Tien-Tsin Wong, Wai-Shun Tong, Chi-Keung Tang, Andrew J. Hanson |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2003 | Data Compression for a Massive Image Data Set in IBMRabstractIn the area of image-based modeling and rendering (IBMR), a new scene under a novel illumination condition is synthesized by interpolating reference images. However, for good quality rendering tremendous reference images under various illumination conditions are required. We present a two-level compression method for reference images. In the first level, the spherical harmonic transform is used to approximate the plenoptic function of each pixel In the second level, we use the embedded zerotrees wavelet (EZW) method for removing the spatial redundancy in spherical coefficients produced from the first level process. Simulation results show our approach is much superior than compressing reference images with JPEG. Andrew Chi-Sing Leung, Ping-Man Lam, Tien-Tsin Wong |
Computer Graphics International | 3 |
| 2003 | Practical Super-Resolution from Dynamic Video SequencesabstractThis paper introduces a practical approach for superresolution, the process of reconstructing a high-resolution image from the low-resolution input ones. The emphasis of our work is to super-resolve frames from dynamic video sequences, which may contain significant object occlusion or scene changes. As the quality of super-resolved images highly relies on the correctness of image alignment between consecutive frames, we employ the robust optical flow method to accurately estimate motion between the image pair. An efficient and reliable scheme is designed to detect and discard incorrect matchings, which may degrade the output quality. We also introduce the usage of elliptical weighted average (EWA) filter to model the spatially variant point spread function (PSF) of acquisition system in order to improve accuracy of the model. A number of complex and dynamic video sequences are tested to demonstrate the applicability and reliability of our algorithm. Zhongding Jiang, Tien-Tsin Wong, Hujun Bao |
CVPR (2) | 2 |
| 2003 | Image-based relighting as the sampling and reconstruction of the plenoptic illumination functionabstractImage-based modeling and rendering has been demonstrated as a cost-effective and efficient approach to real-time graphics systems. In this paper, we describe an extended formulation of the plenoptic function, called the plenoptic illumination function, which explicitly specifies the illumination component. Techniques based on it can be extended to support relighting as well as view interpolation. Based on the linearity of illumination, image-based relighting can be performed with complex lighting configuration. The core of this framework is compression, and we therefore show how to exploit two types of data correlation, intra-pixel and inter-pixel correlations, in order to achieve a manageable storage size. Tien-Tsin Wong |
ICASSP (4) | 1 |
| 2003 | PCA-based compression for image-based relightingabstractThe ability to change illumination is a crucial factor in image-based modeling and rendering. Image-based relighting offers such capability. However, the trade-off is the enormous increase of storage requirement. In this paper, we propose a compression scheme that effectively reduces the data volume while maintaining the real-time relighting capability. The proposed method is based on principal component analysis (PCA). A block-wise PCA is used to practically process the huge input data. The output of PCA is a set of eigenimages and the corresponding relighting coefficients. By dropping those low-energy eigenimages, the data size is drastically reduced. To further compress the data, eigenimages left are compressed using transform coding and quantization while the relighting coefficients are compressed using uniform quantization. We also suggest the suitable target bit rate for each phase of the compression method in order to preserve the visual quality. Finally, we propose real-time engine that relights images from the compressed data. Pun-Mo Ho, Tien-Tsin Wong, Kwok-Hung Choy, Andrew Chi-Sing Leung |
ICME | 2 |
| 2003 | An improved optimal bit allocation method for sub-band coding
Ze Wang 0018, Yin Lee, Andrew Chi-Sing Leung, Tien-Tsin Wong, Yisheng Zhu |
Pattern Recognit. Lett. | 4 |
| 2003 | Compression of illumination-adjustable imagesabstractThe image-based modeling and rendering (IBMR) approaches allow the time complexity of synthesizing novel images to be independent of scene complexity. Unfortunately, illumination control (relighting) is no longer trivial under the image-based framework. To relight the image-based scenery, the scene must be captured under various illumination conditions. This drastically increases the data volume. Hence, data compression is a must. In this paper, we describe a compression algorithm for an illumination-adjustable image representation. We focus on compressing constant-viewpoint images. A divide-and-conquer approach is proposed. The compression algorithm consists of three parts which exploit the intrapixel, interpixel, and interchannel data correlations. Experimental result shows that the proposed method not just effectively compresses the data but also outperforms standard image and video coding methods. Tien-Tsin Wong, Andrew Chi-Sing Leung |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | The immersive cockpitabstractWide field-of-view (FOV) is necessary for many industrial applications, such as air traffic control, large vehicle driving and navigation. Unfortunately, the supporting structure/frame in most systems usually blocks part of the view, results in "blind spot" and raises the risk. In some cases, the working site is hazardous to the pilot. In this video demonstration, we introduce a video-based tele-immersive system, called the Immersive Cockpit. It captures live videos from the working site and recreates an immersive environment at the remote site where the pilot situates. It immerses the pilot at the remote site with a panoramic view of the environment, hence improves interactivity and safety. The design goals of our system are real-time, live, low-cost and scalable.We stitch multiple video streams captured from ordinary CCD cameras to generate a panoramic video. To avoid being blocked by the supporting frame, we allow a flexible placement of cameras. This approach trades the accuracy of the generated panorama image for a larger field-of-view. The panoramic video is presented on an immersive display which covers the field-of-view of the viewer. Wai-Kwan Tang, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Multimedia | 2 |
| 2002 | Relighting with the Reflected Irradiance Field: Representation, Sampling and Reconstruction
Zhouchen Lin, Tien-Tsin Wong, Harry Shum |
Int. J. Comput. Vis. | 2 |
| 2002 | The plenoptic illumination functionabstractImage-based modeling and rendering has been demonstrated as a cost-effective and efficient approach to virtual reality applications. The computational model that most image-based techniques are based on is the plenoptic function. Since the original formulation of the plenoptic function does not include illumination, most previous image-based virtual reality applications simply assume that the illumination is fixed. We propose a formulation of the plenoptic function, called the plenoptic illumination function, which explicitly specifies the illumination component. Techniques based on this new formulation can be extended to support relighting as well as view interpolation. To relight images with various illumination configurations, we also propose a local illumination model, which utilizes the rules of image superposition. We demonstrate how this new formulation can be applied to extend two existing image-based representations, panorama representation such as QuickTime VR and two-plane parameterization, to support relighting with trivial modifications. The core of this framework is compression, and we therefore show how to exploit two types of data correlation, the intra-pixel and the inter-pixel correlations, in order to achieve a manageable storage size. Tien-Tsin Wong, Chi-Wing Fu, Pheng-Ann Heng, Andrew Chi-Sing Leung |
IEEE Trans. Multim. | 1 |
| 2001 | Relighting with the Reflected Irradiance Field: Representation, Sampling and ReconstructionabstractImage-based relighting (IBL) is a technique to change the illumination of an image-based object/scene. In this paper, we define a representation called the reflected irradiance field which records the reflection from an object surface irradiated by a point light source that moves on a plane. This representation is dual to that of the light field. It synthesizes a novel image under a different illumination by interpolating and superimposing appropriate recorded samples. Furthermore, we study the minimum sampling problem of the reflected irradiance field, i.e., how many point light sources are needed during sampling. We find that there exists a geometry-independent bound for the sampling interval whenever the second-order derivatives of the surface BRDF and the minimum depth of the scene are bounded. This bound ensures that the error in the reconstructed image is controlled by a given tolerance, regardless of the geometry. Experiments on both synthetic and real surfaces are conducted to verify our analysis. Zhouchen Lin, Tien-Tsin Wong, Harry Shum |
CVPR (1) | 2 |
| 2001 | Continuous field based free-form surface modeling and morphing
Hujun Bao, Pheng-Ann Heng, Tien-Tsin Wong, Qunsheng Peng 0001 |
Comput. Graph. | 4 |
| 2001 | Constrained Fairing for MeshesabstractIn this paper, we present a novel fairing algorithm for the removal of noise from uniform triangular meshes without shrinkage and serious distortion. The key feature of this algorithm is to keep all triangle centers invariant at each smoothing step by including some constraints in the energy minimization functional. The constrained functional is then minimized efficiently using an iterative method. Further, we apply this smoothing technique to a multiresolution representation to remove arbitrary levels of detail. A volume‐preserving decimation algorithm is presented to generate the multiresolution representation. The experimental results demonstrate the combined algorithm's stability and efficiency. Xinguo Liu, Hujun Bao, Qunsheng Peng 0001, Pheng-Ann Heng, Tien-Tsin Wong |
Comput. Graph. Forum | 5 |
| 2000 | Progressive Geometry Compression for MeshesabstractA novel progressive geometry compression scheme is presented in this paper. In this scheme, a mesh is represented as a base mesh followed by some groups of vertex split operations using an improved simplification method in which each level of the mesh can be refined into the next level by carrying out a group of vertex split operations in any order. Consequently, the progressive mesh (PM) representation can be effectively encoded by permuting the vertex split operations in each group. Meanwhile, a geometry predictor using the Laplacian operator is designed to predict each new vertex position using its neighbours. The correction is quantized and encoded using a Huffman coding scheme. Experimental results show that our algorithm obtains higher compression ratios than previous work. It is very suitable for the progressive transmission of geometric models over the Internet. Xinguo Liu, Hujun Bao, Qunsheng Peng 0001, Pheng-Ann Heng, Tien-Tsin Wong, Hanqiu Sun |
PG | 5 |
| 1999 | A Panoramic-Based Walkthrough System Using Real PhotosabstractRecent advances in image-based rendering techniques allow us to have interactive panoramic viewing of real world environment using real photos. With panorama visualization tools like QuickTime VR or Realspace Viewer, we are able to have fast navigation in an image-based environment. However, the navigation control is restricted to panning, zooming and tilting of a single panoramic view. Smooth walkthrough between panoramic nodes is not yet supported. In this paper, we first describe how we apply the epipolar geometry between cylindrical projection manifolds to find point correspondences between real-world panoramic photos. A correspondence matching program is developed to allow users to identify correspondence interactively. The correspondence information is identified by defining patches on the panoramic image. Then, the system can generate a triangular mesh for the panorama based on the given patches. After that, we are able to provide free navigation between panoramic nodes by means of the triangle-based image warping technique. Chan Yan Fai, Fok Man Hong, Chi-Wing Fu, Pheng-Ann Heng, Tien-Tsin Wong |
PG | 5 |
| 1998 | Multiresolution Isosurface Extraction with Adaptive Skeleton ClimbingabstractAn isosurface extraction algorithm which can directly generate multiresolution isosurfaces from volume data is introduced. It generates low resolution isosurfaces, with 4 to 25 times fewer triangles than that generated by marching cubes algorithm, in comparable running times. By climbing from vertices (0‐skeleton) to edges (1‐skeleton) to faces (2‐skeleton), the algorithm constructs boxes which adapt to the geometry of the true isosurface. Unlike previous adaptive marching cubes algorithms, the algorithm does not suffer from the gap‐filling problem. Although the triangles in the meshes may not be optimally reduced, it is much faster than postprocessing triangle reduction algorithms. Hence the coarse meshes it produces can be used as the initial starts for the mesh optimization, if mesh optimality is the main concern. Tim Poston, Tien-Tsin Wong, Pheng-Ann Heng |
Comput. Graph. Forum | 2 |
| 1998 | Illumination of image-based objectsabstractA new data representation of image-based objects is presented. With this representation, the user can change the illumination as well as the viewpoint of an image-based scene. Physically correct imagery can be generated without knowing any geometrical information (e.g. depth or surface normal) of the scene. By treating each pixel on the image plane as a surface element, we can measure its apparent BRDF (bidirectional reflectance distribution function) by collecting information in the sampled images. These BRDFs allow us to calculate the correct pixel colour under a new illumination set-up by fitting the intensity, direction and number of the light sources. We demonstrate that the proposed representation allows re-rendering of the scene illuminated by different types of light sources. Moreover, two compression schemes, spherical harmonics and discrete cosine transform, are proposed to compress the huge amount of tabular BRDF data. © 1998 John Wiley & Sons, Ltd. Tien-Tsin Wong, Pheng-Ann Heng, Siu-Hang Or, Wai-Yin Ng |
Comput. Animat. Virtual Worlds | 1 |
| 1997 | Two New Quorum Based Algorithms for Distributed Mutual ExclusionabstractTwo novel suboptimal algorithms for mutual exclusion in distributed systems are presented. One is based on the modification of Maekawa's (1985) grid based quorum scheme. The size of quorums is approximately /spl radic/2/spl radic/N where N is the number of sites in a network, as compared to 2/spl radic/N of the original method. The method is simple and geometrically evident. The second one is based on the idea of difference sets in combinatorial theory. The resulting scheme is very close to optimal in terms of quorum size. Wai-Shing Luk, Tien-Tsin Wong |
ICDCS | 2 |
| 1997 | "Skeleton climbing": fast isosurfaces with fewer trianglesabstractSkeleton climbing is an algorithm that builds triangulated isosurfaces in 3D grid data, more economically than marching cubes, and without the time penalty of current mesh decimation algorithms. Building the surface from its intersections with grid edges (1-skeleton), then faces (2-skeleton), then cubes (3-skeleton), treats the data in a uniform way; this allows a 25% reduction in the number of triangles produced, while still creating a true separating surface at similar speed. Tim Poston, Pheng-Ann Heng, Tien-Tsin Wong |
PG | 4 |
| 1997 | Illuminating image-based objectsabstractWe present a new scheme of data representation for image-based objects. It allows the illumination to be changed interactively without knowing any geometrical information (e.g. depth or surface normal) of the scene, but the resulting images are physically correct. The scene is first sampled from different view points and under different illuminations. By treating each pixel on the image plane as a surface element, the sampled images are used to measure the apparent BRDF of each surface element. Two compression schemes, spherical harmonics and discrete cosine transform, are proposed to compress the tabular BRDF data. Whenever the user changes the illumination a certain number of views are reconstructed. The correct user perspective view is then displayed using the standard texture mapping hardware. Hence, the intensity, the type and the number of the light sources can be manipulated interactively. Tien-Tsin Wong, Pheng-Ann Heng, Siu-Hang Or, Wai-Yin Ng |
PG | 1 |