EDBT 2026 Demo / reviewers in the wild / expert
Tat-Jen Cham
dblp:29/3808 · also Cham Tat Jen
· DBLP profile ↗
100ranked-venue papers
13as first author
34since 2021 · last 2025
0000-0001-5264-2572ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 76 · 11 first-author · 20 since 2021Artificial intelligence and machine learning · 69 · 13 first-author · 26 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual ForagingabstractImagine searching a collection of coins for quarters (0.25), dimes (0.10), nickels (0.05), and pennies (0.01)—a hybrid foraging task where observers look for multiple instances of multiple target types. In such tasks, how do target values and their prevalence influence foraging and eye movement behaviors (e.g., should you prioritize rare quarters or common nickels)? To explore this, we conducted human psychophysics experiments, revealing that humans are proficient reward foragers. Their eye fixations are drawn to regions with higher average rewards, fixation durations are longer on more valuable targets, and their cumulative rewards exceed chance, approaching the upper bound of optimal foragers. To probe these decision-making processes of humans, we developed a transformer-based Visual Forager (VF) model trained via reinforcement learning. Our VF model takes a series of targets, their corresponding values, and the search image as inputs, processes the images using foveated vision, and produces a sequence of eye movements along with decisions on whether to collect each fixated item. Our model outperforms all baselines, achieves cumulative rewards comparable to those of humans, and approximates human foraging behavior in eye movements and foraging biases within time-limited environments. Furthermore, stress tests on out-of-distribution tasks with novel targets, unseen values, and varying set sizes demonstrate the VF model’s effective generalization. Our work offers valuable insights into the relationship between eye movements and decision-making, with our model serving as a powerful tool for further exploration of this connection. All data, code, and models are available at https://github.com/ZhangLab-DeepNeuroCogLab/visual-forager. Dingwei Tan, Yen-Ling Kuo, Zhaowei Sun, Jeremy M. Wolfe, Tat-Jen Cham, Mengmi Zhang |
CVPR | 6 |
| 2025 | Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
Tianhao Wu 0015, Chuanxia Zheng, Frank Guan, Andrea Vedaldi, Tat-Jen Cham |
ICCV | 5 |
| 2025 | Diffusion Pretraining for Gait Recognition in the WildabstractRecently, diffusion models have garnered much attention for their remarkable generative capabilities. Yet, their application for representation learning remains largely unexplored. In this paper, we explore the potential of diffusion models to pretrain the backbone of a deep learning model for a specific application—gait recognition in the wild. To do so, we condition a latent diffusion model on the output of a gait recognition model backbone. Our pretraining experiments on the Gait3D and GREW datasets reveal an interesting phenomenon: diffusion pretraining causes the gait recognition backbone to separate gait sequences belonging to different subjects further apart than those belonging to the same subjects. Subsequently, our transfer learning experiments on Gait3D and GREW show that the pretrained backbone can serve as an effective initialization for the downstream gait recognition task, improving gait recognition accuracies by as much as 7.9% on Gait3D and 4.2% on GREW. Wei Ming Neo, Koichi Shinoda, Tat-Jen Cham |
ICIP | 3 |
| 2025 | Semantix: An Energy-guided Sampler for Semantic Style TransferabstractRecent advances in style and appearance transfer are impressive, but most methods isolate global style and local appearance transfer, neglecting semantic correspondence. Additionally, image and video tasks are typically handled in isolation, with little focus on integrating them for video transfer. To address these limitations, we introduce a novel task, *Semantic Style Transfer*, which involves transferring style and appearance features from a reference image to a target visual content based on semantic correspondence. We subsequently propose a training-free method, *Semantix*, an energy-guided sampler designed for Semantic Style Transfer that simultaneously guides both style and appearance transfer based on semantic understanding capacity of pre-trained diffusion models. Additionally, as a sampler, *Semantix* can be seamlessly applied to both image and video models, enabling semantic style transfer to be generic across various visual media. Specifically, once inverting both reference and context images or videos to noise space by SDEs, *Semantix* utilizes a meticulously crafted energy function to guide the sampling process, including three key components: *Style Feature Guidance*, *Spatial Feature Guidance* and *Semantic Distance* as a regularisation term. Experimental results demonstrate that *Semantix* not only effectively accomplishes the task of semantic style transfer across images and videos, but also surpasses existing state-of-the-art solutions in both fields. Huiang He, Minghui Hu 0001, Chuanxia Zheng, Tat-Jen Cham |
ICLR | 5 |
| 2025 | TSC-Net: Prediction of Pedestrian Trajectories by Trajectory-Scene-Cell ClassificationabstractTo predict future trajectories of pedestrians, scene is as important as the history trajectory since i) scene reflects the position of possible goals of the pedestrian ii) trajectories are affected by the semantic information of the scene. It requires the model to capture scene information and learn the relation between scenes and trajectories. However, existing methods either apply Convolutional Neural Networks (CNNs) to summarize the scene to a feature vector, which raises the feature misalignment issue, or convert trajectory to heatmaps to align with the scene map, which ignores the interactions among different pedestrians. In this work, we introduce the trajectory-scene-cell feature to represent both trajectories and scenes in one feature space. By decoupling the trajectory in temporal domain and the scene in spatial domain, trajectory feature and scene feature are re-organized in different types of cell feature, which well aligns trajectory and scene, and allows the framework to model both human-human and human-scene interactions. Moreover, the Trajectory-Scene-Cell Network (TSC-Net) with new trajectory prediction manner is proposed, where both goal and intermediate positions of the trajectory are predict by cell classification and offset regression. Comparative experiments show that TSC-Net achieves the SOTA performance on several datasets with most of the metrics. Especially for the goal estimation, TSC-Net is demonstrated better on predicting goals for trajectories with irregular speed. Tat-Jen Cham |
ICLR | 2 |
| 2025 | Explicit Correspondence Matching for Generalizable Neural Radiance FieldsabstractWe present a new generalizable NeRF method that is able to directly generalize to new unseen scenarios and perform novel view synthesis with as few as two source views. The key to our approach lies in the explicitly modeled correspondence matching information, so as to provide the geometry prior to the prediction of NeRF color and density for volume rendering. The explicit correspondence matching is quantified with the cosine similarity between image features sampled at the 2D projections of a 3D point on different views, which is able to provide reliable cues about the surface geometry. Unlike previous methods where image features are extracted independently for each view, we consider modeling the cross-view interactions via Transformer cross-attention, which greatly improves the feature matching quality. Our method achieves state-of-the-art results on different evaluation settings, with the experiments showing a strong correlation between our learned cosine feature similarity and volume density, demonstrating the effectiveness and superiority of our proposed method. Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | One-Shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural TexturingabstractHuman motion transfer aims at animating a static source image with a driving video. While recent advances in one-shot human motion transfer have led to significant improvement in results, it remains challenging for methods with 2D body landmarks, skeleton and semantic mask to accurately capture correspondences between source and driving poses due to the large variation in motion and articulation complexity. In addition, the accuracy and precision of DensePose degrade the image quality for neural-rendering-based methods. To address the limitations and by both considering the importance of appearance and geometry for motion transfer, in this work, we proposed a unified framework that combines multi-scale feature warping and neural texture mapping to recover better 2D appearance and 2.5D geometry, partly by exploiting the information from DensePose, yet adapting to its inherent limited accuracy. Our model takes advantage of multiple modalities by jointly training and fusing them, which allows it to robust neural texture features that cope with geometric errors as well as multi-scale dense motion flow that better preserves appearance. Experimental results with full and half-view body video datasets demonstrate that our model can generalize well and achieve competitive results, and that it is particularly effective in handling challenging cases such as those with substantial self-occlusions. Yuzhu Ji, Chuanxia Zheng, Tat-Jen Cham |
IEEE Trans. Multim. | 3 |
| 2025 | D-LORD for Motion StylizationabstractThis article introduces a novel framework named double-latent optimization for representation disentanglement (D-LORD), which is designed for motion stylization (motion style transfer and motion retargeting). The primary objective of this framework is to separate the class and content information from a given motion sequence using a data-driven latent optimization approach. Here, class refers to person-specific style, such as a particular emotion or an individual’s identity, while content relates to the style-agnostic aspect of an action, such as walking or jumping, as universally understood concepts. The key advantage of D-LORD is its ability to perform style transfer without needing paired motion data. Instead, it utilizes class and content labels during the latent optimization process. By disentangling the representation, the framework enables the transformation of one motion sequence’s style to another’s style using adaptive instance normalization. The proposed D-LORD framework is designed with a focus on generalization, allowing it to handle different class and content labels for various applications. In addition, it can generate diverse motion sequences when specific class and content labels are provided. The framework’s efficacy is demonstrated through experimentation on three datasets: 1) the CMU XIA dataset for motion style transfer; 2) the multimodal human action database dataset; and 3) the RRIS Ability dataset for motion retargeting. Notably, this article presents the first generalized framework for motion style transfer and motion retargeting, showcasing its potential contributions in this area. Meenakshi Gupta, Mingyuan Lei, Tat-Jen Cham, Hwee Kuan Lee |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Dynamic Eyebox Steering for Improved Pinlight AR Near-Eye DisplaysabstractAn optical-see-through near-eye display (NED) for augmented reality (AR) allows the user to perceive virtual and real imagery simultaneously. Existing technologies for optical-see-through AR NEDs involve trade-offs between key metrics such as field of view (FOV), eyebox size, form factor, etc. We have enhanced an existing compact wide-FOV pinlight AR NED design with real-time 3D pupil localization in order to dynamically steer and thus effectively enlarge the usable eyebox. This is achieved with a dual-camera rig that captures stereoscopic views of the pupils. The 3D pupil location is used to dynamically calculate a display pattern that spatio-temporally modulates the light entering the wearer's eyes. We have built a demonstrable compact prototype and have conducted a user study that indicates the effectiveness of our eyebox steering method (e.g., without eyebox steering, in 10.5% of our tests, users were unable to perceive the test pattern correctly before experiment timeout; with eyebox steering, that fraction decreased dramatically to 1.25%). This is a small yet crucial step in making simple wide-FOV pinlight NEDs usable for human users and not just as demonstration prototypes filmed with a precisely positioned camera standing in for the user's eye. Further contributions of this paper include a detailed description of display design, calibration technique, and user study design, all of which may benefit other NED research. Xinxing Xia, Zheye Yu, Dongyu Qiu, Andrei State, Tat-Jen Cham, Frank Guan, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | One More Step: A Versatile Plug-and-Play Module for Rectifying Diffusion Schedule Flaws and Enhancing Low-Frequency ControlsabstractIt is well known that many open-released foundational diffusion models have difficulty in generating images that substantially depart from average brightness, despite such images being present in the training data. This is due to an inconsistency: while denoising starts from pure Gaus-sian noise during inference, the training noise schedule retains residual data even in the final timestep distribution, due to difficulties in numerical conditioning in main-stream formulation, leading to unintended bias during in-ference. To mitigate this issue, certain ∊-prediction mod-els are combined with an ad-hoc offset-noise methodology. In parallel, some contemporary models have adopted zero-terminal SNR noise schedules together with v -prediction, which necessitate major alterations to pre-trained models. However, such changes risk destabilizing a large multitude of community-driven applications anchored on these pre-trained models. In light of this, our investigation revisits the fundamental causes, leading to our proposal of an inno-vative and principled remedy, called One More Step (OMS). By integrating a compact network and incorporating an ad-ditional simple yet effective step during inference, OMS ele-vates image fidelity and harmonizes the dichotomy between training and inference, while preserving original model pa-rameters. Once trained, various pre-trained diffusion mod-els with the same latent domain can share the same OMS module. Codes and models are released at here. Minghui Hu 0001, Chuanxia Zheng, Dacheng Tao, Tat-Jen Cham |
CVPR | 6 |
| 2024 | MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-view Images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger 0001, Tat-Jen Cham, Jianfei Cai 0001 |
ECCV (21) | 7 |
| 2024 | 3iGS: Factorised Tensorial Illumination for 3D Gaussian Splatting
Zhe Jun Tang, Tat-Jen Cham |
ECCV (14) | 2 |
| 2024 | ClusteringSDF: Self-Organized Neural Implicit Surfaces for 3D Decomposition
Tianhao Wu 0015, Chuanxia Zheng, Qianyi Wu, Tat-Jen Cham |
ECCV (57) | 4 |
| 2024 | PanoDiffusion: 360-degree Panorama Outpainting via DiffusionabstractGenerating complete 360\textdegree{} panoramas from narrow field of view images is ongoing research as omnidirectional RGB data is not readily available. Existing GAN-based approaches face some barriers to achieving higher quality output, and have poor generalization performance over different mask types. In this paper, we present our 360\textdegree{} indoor RGB panorama outpainting model using latent diffusion models (LDM), called PanoDiffusion. We introduce a new bi-modal latent diffusion structure that utilizes both RGB and depth panoramic data during training, which works surprisingly well to outpaint depth-free RGB images during inference. We further propose a novel technique of introducing progressive camera rotations during each diffusion denoising step, which leads to substantial improvement in achieving panorama wraparound consistency. Results show that our PanoDiffusion not only significantly outperforms state-of-the-art methods on RGB panorama outpainting by producing diverse well-structured results for different types of masks, but can also synthesize high-quality depth panoramas to provide realistic 3D indoor models. Tianhao Wu 0015, Chuanxia Zheng, Tat-Jen Cham |
ICLR | 3 |
| 2024 | MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse ViewsabstractWe introduce MVSplat360, a feed-forward approach for 360° novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and insufficient visual information provided, making it challenging for conventional methods to achieve high-quality results. Our MVSplat360 addresses this by effectively combining geometry-aware 3D reconstruction with temporally consistent video generation. Specifically, it refactors a feed-forward 3D Gaussian Splatting (3DGS) model to render features directly into the latent space of a pre-trained Stable Video Diffusion (SVD) model, where these features then act as pose and visual cues to guide the denoising process and produce photorealistic 3D-consistent views. Our model is end-to-end trainable and supports rendering arbitrary views with as few as 5 sparse input views. To evaluate MVSplat360's performance, we introduce a new benchmark using the challenging DL3DV-10K dataset, where MVSplat360 achieves superior visual quality compared to state-of-the-art methods on wide-sweeping or even 360° NVS tasks. Experiments on the existing benchmark RealEstate10K also confirm the effectiveness of our model. Readers are highly recommended to view the video results at [donydchen.github.io/mvsplat360](https://donydchen.github.io/mvsplat360). Yuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang, Andrea Vedaldi, Tat-Jen Cham, Jianfei Cai 0001 |
NeurIPS | 6 |
| 2024 | Bridging Global Context Interactions for High-Fidelity Pluralistic Image CompletionabstractWe introduce PICFormer, a novel framework for Pluralistic Image Completion using a transFormer based architecture, that achieves both high quality and diversity at a much faster inference speed. Our key contribution is to introduce a code-shared codebook learning using a restrictive CNN on small and non-overlapping receptive fields (RFs) for the local visible token representation. This results in a compact yet expressive discrete representation, facilitating efficient modeling of global visible context relations by the transformer. Unlike the prevailing autoregressive approaches, we proposed to sample all tokens simultaneously, leading to more than 100× faster inference speed. To enhance appearance consistency between visible and generated regions, we further propose a novel attention-aware layer (AAL), designed to better exploit distantly related high-frequency features. Through extensive experiments, we demonstrate that the PICFormer efficiently learns semantically-rich discrete codes, resulting in significantly improved image quality. Moreover, our diverse image completion framework surpasses State-of-the-Art methods on multiple image completion datasets. Chuanxia Zheng, Guoxian Song, Tat-Jen Cham, Jianfei Cai 0001, Linjie Luo, Dinh Q. Phung |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | ABLE-NeRF: Attention-Based Rendering with Learnable Embeddings for Neural Radiance FieldabstractNeural Radiance Field (NeRF) is a popular method in representing 3D scenes by optimising a continuous volumetric scene function. Its large success which lies in applying volumetric rendering (VR) is also its Achilles' heel in producing view-dependent effects. As a consequence, glossy and transparent surfaces often appear murky. A remedy to reduce these artefacts is to constrain this VR equation by excluding volumes with back-facing normal. While this approach has some success in rendering glossy surfaces, translucent objects are still poorly represented. In this paper, we present an alternative to the physics-based VR approach by introducing a self-attention-based framework on volumes along a ray. In addition, inspired by modern game engines which utilise Light Probes to store local lighting passing through the scene, we incorporate Learnable Embeddings to capture view dependent effects within the scene. Our method, which we call ABLE-NeRF, significantly reduces ‘blurry’ glossy surfaces in rendering and produces realistic translucent surfaces which lack in prior art. In the Blender dataset, ABLE-NeRF achieves SOTA results and surpasses Ref-NeRF in all 3 image quality metrics PSNR, SSIM, LPIPS. Zhe Jun Tang, Tat-Jen Cham, Haiyu Zhao |
CVPR | 2 |
| 2023 | Unified Discrete Diffusion for Simultaneous Vision-Language Generation
Minghui Hu 0001, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham, Heliang Zheng, Dacheng Tao, Ponnuthurai N. Suganthan |
ICLR | 4 |
| 2023 | Cocktail: Mixing Multi-Modality Control for Text-Conditional Image GenerationabstractText-conditional diffusion models are able to generate high-fidelity images with diverse contents.
However, linguistic representations frequently exhibit ambiguous descriptions of the envisioned objective imagery, requiring the incorporation of additional control signals to bolster the efficacy of text-guided diffusion models.
In this work, we propose Cocktail, a pipeline to mix various modalities into one embedding, amalgamated with a generalized ControlNet (gControlNet), a controllable normalisation (ControlNorm), and a spatial guidance sampling method, to actualize multi-modal and spatially-refined control for text-conditional diffusion models.
Specifically, we introduce a hyper-network gControlNet, dedicated to the alignment and infusion of the control signals from disparate modalities into the pre-trained diffusion model.
gControlNet is capable of accepting flexible modality signals, encompassing the simultaneous reception of any combination of modality signals, or the supplementary fusion of multiple modality signals.
The control signals are then fused and injected into the backbone model according to our proposed ControlNorm.
Furthermore, our advanced spatial guidance sampling methodology proficiently incorporates the control signal into the designated region, thereby circumventing the manifestation of undesired objects within the generated image.
We demonstrate the results of our method in controlling various modalities, proving high-quality synthesis and fidelity to multiple external signals. Minghui Hu 0001, Daqing Liu, Chuanxia Zheng, Dacheng Tao, Tat-Jen Cham |
NeurIPS | 7 |
| 2022 | Global Context with Discrete Diffusion in Vector Quantised Modelling for Image GenerationabstractThe integration of Vector Quantised Variational AutoEncoder (VQ-VAE) with autoregressive models as generation part has yielded high-quality results on image generation. However, the autoregressive models will strictly follow the progressive scanning order during the sampling phase. This leads the existing VQ series models to hardly escape the trap of lacking global information. Denoising Diffusion Probabilistic Models (DDPM) in the continuous domain have shown a capability to capture the global context, while generating high-quality images. In the discrete state space, some works have demonstrated the potential to perform text generation and low resolution image generation. We show that with the help of a content-rich discrete visual codebook from VQ-VAE, the discrete diffusion model can also generate high fidelity images with global context, which compensates for the deficiency of the classical autoregressive model along pixel space. Meanwhile, the integration of the discrete VAE with the diffusion model resolves the drawback of conventional autoregressive models being oversized, and the diffusion model which demands excessive time in the sampling process when generating images. It is found that the quality of the generated images is heavily dependent on the discrete visual codebook. Extensive experiments demonstrate that the proposed Vector Quantised Discrete Diffusion Model (VQ-DDM) is able to achieve comparable performance to top-tier methods with low complexity. It also demonstrates outstanding advantages over other vectors quantised with autoregressive models in terms of image inpainting tasks without additional training. Minghui Hu 0001, Tat-Jen Cham, Jianfei Yang 0001, Ponnuthurai N. Suganthan |
CVPR | 3 |
| 2022 | Bridging Global Context Interactions for High-Fidelity Image CompletionabstractBridging global context interactions correctly is important for high-fidelity image completion with large masks. Previous methods attempting this via deep or large receptive field (RF) convolutions cannot escape from the dominance of nearby interactions, which may be inferior. In this paper, we propose to treat image completion as a directionless sequence-to-sequence prediction task, and deploy a transformer to directly capture long-range depen-dence. Crucially, we employ a restrictive CNN with small and non-overlapping RF for weighted token representation, which allows the transformer to explicitly model the long-range visible context relations with equal importance in all layers, without implicitly confounding neighboring tokens when larger RFs are used. To improve appearance consistency between visible and generated regions, a novel attention-aware layer (AAL) is introduced to better exploit distantly related high-frequency features. Overall, extensive experiments demonstrate superior performance compared to state-of-the-art methods on several datasets. Code is available at https://github.com/lyndonzheng/TFill. Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001, Dinh Q. Phung |
CVPR | 2 |
| 2022 | Sem2NeRF: Converting Single-View Semantic Masks to Neural Radiance Fields
Yuedong Chen, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
ECCV (14) | 4 |
| 2022 | Entry-Flipped Transformer for Inference and Prediction of Participant Behavior
Tat-Jen Cham |
ECCV (4) | 2 |
| 2022 | MPT-Net: Mask Point Transformer Network for Large Scale Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation is important for road scene perception, a task for driverless vehicles to achieve full fledged autonomy. In this work, we introduce Mask Point Transformer Network (MPT-Net), a novel architecture for point cloud segmentation that is simple to implement. MPT-Net consists of a local and global feature encoder and a transformer based decoder; a 3D Point-Voxel Convolution encoder backbone with voxel self attention to encode features and a Mask Point Transformer module to decode point features and segment the point cloud. Firstly, we introduce the novel MPT designed to specifically handle point cloud segmentation. MPT offers two benefits. It attends to every point in the point cloud using mask tokens to extract class specific features globally with cross attention, and provide inter-class feature information exchange using self attention on the learned mask tokens. Secondly, we design a backbone to use sparse point voxel convolutional blocks and a self attention block using transformers to learn local and global contextual features. We evaluate MPT-Net on large scale outdoor driving scene point cloud datasets, SemanticKITTI and nuScenes. Our experiments show that by replacing the standard segmentation head with MPT, MPT-Net achieves a state-of-the-art performance over our baseline approach by 3.8% in SemanticKITTI and is highly effective in detecting 'stuffs' in point cloud. Zhe Jun Tang, Tat-Jen Cham |
IROS | 2 |
| 2022 | Real-time Shadow-aware Portrait Relighting in Virtual Backgrounds for Realistic TelepresenceabstractWhile using virtual backgrounds has recently become a very popular feature in videoconferencing, there often exists a jarring mismatch between the lighting of the user and the illumination condition of the virtual background. Existing portrait relighting methods can alleviate the problem, but do not have the capacity to deal with difficult shadow effects. In this paper, we present a new shadow-aware portrait relighting system that can relight an input portrait to be consistent with a given desired background image with shadow effects. Our system consists of four major components: portrait neutralization, illumination estimation, shadow generation and hierarchical neural rendering, which are all based on deep neural networks, and the whole system is end-to-end trainable. In addition, we created a large-scale photorealistic synthetic dataset with shadow, illumination and depth annotations for training, which allows our model to generalize well to real images. The extensive experiments demonstrate that our shadow-aware relight system outperforms the state-of-the-art portrait relighting solutions in terms of producing more lighting-consistent relighted images with shadow effects. Guoxian Song, Tat-Jen Cham, Jianfei Cai 0001, Jianmin Zheng |
ISMAR | 2 |
| 2022 | Towards Unbiased Visual Emotion Recognition via Causal InterventionabstractAlthough much progress has been made in visual emotion recognition, researchers have realized that modern deep networks tend to exploit dataset characteristics to learn spurious statistical associations between the input and the target. Such dataset characteristics are usually treated as dataset bias, which damages the robustness and generalization performance of these recognition systems. In this work, we scrutinize this problem from the perspective of causal inference, where such dataset characteristic is termed as a confounder which misleads the system to learn the spurious correlation. To alleviate the negative effects brought by the dataset bias, we propose a novel Interventional Emotion Recognition Network (IERN) to achieve the backdoor adjustment, which is one fundamental deconfounding technique in causal inference. Specifically, IERN starts by disentangling the dataset-related context feature from the actual emotion feature, where the former forms the confounder. The emotion feature will then be forced to see each confounder stratum equally before being fed into the classifier. A series of designed tests validate the efficacy of IERN, and experiments on three emotion benchmarks demonstrate that IERN outperforms state-of-the-art approaches for unbiased visual emotion recognition. Yuedong Chen, Xu Yang 0021, Tat-Jen Cham, Jianfei Cai 0001 |
ACM Multimedia | 3 |
| 2022 | GeoConv: Geodesic guided convolution for facial action unit recognition
Yuedong Chen, Guoxian Song, Zhiwen Shao, Jianfei Cai 0001, Tat-Jen Cham, Jianmin Zheng |
Pattern Recognit. | 5 |
| 2022 | Unconstrained Facial Action Unit Detection via Latent Feature DomainabstractFacial action unit (AU) detection in the wild is a challenging problem, due to the unconstrained variability in facial appearances and the lack of accurate annotations. Most existing methods depend on either impractical labor-intensive labeling or inaccurate pseudo labels. In this paper, we propose an end-to-end unconstrained facial AU detection framework based on domain adaptation, which transfers accurate AU labels from a constrained source domain to an unconstrained target domain by exploiting labels of AU-related facial landmarks. Specifically, we map a source image with label and a target image without label into a latent feature domain by combining source landmark-related feature with target landmark-free feature. Due to the combination of source AU-related information and target AU-free information, the latent feature domain with transferred source label can be learned by maximizing the target-domain AU detection performance. Moreover, we introduce a novel landmark adversarial loss to disentangle the landmark-free feature from the landmark-related feature by treating the adversarial learning as a multi-player minimax game. Our framework can also be naturally extended for use with target-domain pseudo AU labels. Extensive experiments show that our method soundly outperforms lower-bounds and upper-bounds of the basic model, as well as state-of-the-art approaches on the challenging in-the-wild benchmarks. The code is available athttps://github.com/ZhiwenShao/ADLD. Zhiwen Shao, Jianfei Cai 0001, Tat-Jen Cham, Xuequan Lu, Lizhuang Ma |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | The Spatially-Correlative Loss for Various Image Translation TasksabstractWe propose a novel spatially-correlative loss that is simple, efficient and yet effective for preserving scene structure consistency while supporting large appearance changes during unpaired image-to-image (I2I) translation. Previous methods attempt this by using pixel-level cycle-consistency or feature-level matching losses, but the domain-specific nature of these losses hinder translation across large domain gaps. To address this, we exploit the spatial patterns of self-similarity as a means of defining scene structure. Our spatially-correlative loss is geared towards only capturing spatial relationships within an image rather than domain appearance. We also introduce a new self-supervised learning method to explicitly learn spatially-correlative maps for each specific translation task. We show distinct improvement over baseline models in all three modes of unpaired I2I translation: single-modal, multi-modal, and even single-image translation. This new loss can easily be integrated into existing network architectures and thus allows wide applicability. The code is available at https://github.com/lyndonzheng/F-LSeSim. Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
CVPR | 2 |
| 2021 | A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder∗abstractWe present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requirements, most existing approaches only cater to a specific task or use different architectures to address various tasks. Here we propose a unified framework based on Conditional Variational Auto-Encoder (CVAE), where we treat any arbitrary input as a masked motion series. Notably, by considering this problem as a conditional generation process, we estimate a parametric distribution of the missing regions based on the input conditions, from which to sample and synthesize the full motion series. To further allow the flexibility of manipulating the motion style of the generated series, we design an Action-Adaptive Modulation (AAM) to propagate the given semantic guidance through the whole sequence. We also introduce a cross-attention mechanism to exploit distant relations among decoder and encoder features for better realism and global consistency. We conducted extensive experiments on Human 3.6M and CMU-Mocap. The results show that our method produces coherent and realistic results for various motion synthesis tasks, with the synthesized motions distinctly adapted by the given action labels. Yujun Cai, Yiwei Wang 0001, Yiheng Zhu 0003, Tat-Jen Cham, Jianfei Cai 0001, Junsong Yuan 0001, Jun Liu 0036, Chuanxia Zheng, Sijie Yan, Henghui Ding, Xiaohui Shen, Ding Liu 0001, Nadia Magnenat-Thalmann |
ICCV | 4 |
| 2021 | Half-body Portrait Relighting with Overcomplete Lighting RepresentationabstractAbstract We present a neural‐based model for relighting a half‐body portrait image by simply referring to another portrait image with the desired lighting condition. Rather than following classical inverse rendering methodology that involves estimating normals, albedo and environment maps, we implicitly encode the subject and lighting in a latent space, and use these latent codes to generate relighted images by neural rendering. A key technical innovation is the use of a novel overcomplete lighting representation, which facilitates lighting interpolation in the latent space, as well as helping regularize the self‐organization of the lighting latent space during training. In addition, we propose a novel multiplicative neural render that more effectively combines the subject and lighting latent codes for rendering. We also created a large‐scale photorealistic rendered relighting dataset for training, which allows our model to generalize well to real images. Extensive experiments demonstrate that our system not only outperforms existing methods for referral‐based portrait relighting, but also has the capability generate sequences of relighted images via lighting rotations. Guoxian Song, Tat-Jen Cham, Jianfei Cai 0001, Jianmin Zheng |
Comput. Graph. Forum | 2 |
| 2021 | Pluralistic Free-Form Image Completion
Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Visiting the Invisible: Layer-by-Layer Completed Scene Decomposition
Chuanxia Zheng, Duy-Son Dao, Guoxian Song, Tat-Jen Cham, Jianfei Cai 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | AgileGAN: stylizing portraits by inversion-consistent transfer learningabstractPortraiture as an art form has evolved from realistic depiction into a plethora of creative styles. While substantial progress has been made in automated stylization, generating high quality stylistic portraits is still a challenge, and even the recent popular Toonify suffers from several artifacts when used on real input images. Such StyleGAN-based methods have focused on finding the best latent inversion mapping for reconstructing input images; however, our key insight is that this does not lead to good generalization to different portrait styles. Hence we propose AgileGAN, a framework that can generate high quality stylistic portraits via inversion-consistent transfer learning. We introduce a novel hierarchical variational autoencoder to ensure the inverse mapped distribution conforms to the original latent Gaussian distribution, while augmenting the original space to a multi-resolution latent space so as to better encode different levels of detail. To better capture attribute-dependent stylization of facial features, we also present an attribute-aware generator and adopt an early stopping strategy to avoid overfitting small training datasets. Our approach provides greater agility in creating high quality and high resolution (1024×1024) portrait stylization models, requiring only a limited number of style exemplars (~100) and short training time (~1 hour). We collected several style datasets for evaluation including 3D cartoons, comics, oil paintings and celebrities. We show that we can achieve superior portrait stylization quality to previous state-of-the-art methods, with comparisons done qualitatively, quantitatively and through a perceptual user study. We also demonstrate two applications of our method, image editing and motion retargeting. Guoxian Song, Linjie Luo, Wan-Chun Ma, Chun-Pong Lai, Chuanxia Zheng, Tat-Jen Cham |
ACM Trans. Graph. | 7 |
| 2020 | Learning Progressive Joint Propagation for Human Motion Prediction
Yujun Cai, Lin Huang 0004, Yiwei Wang 0001, Tat-Jen Cham, Jianfei Cai 0001, Junsong Yuan 0001, Jun Liu 0036, Xu Yang 0021, Yiheng Zhu 0003, Xiaohui Shen, Ding Liu 0001, Jing Liu 0050, Nadia Magnenat-Thalmann |
ECCV (7) | 4 |
| 2020 | Towards Eyeglass-style Holographic Near-eye Displays with StaticallyabstractHolography is perhaps the only method demonstrated so far that can achieve a wide field of view (FOV) and a compact eyeglass-style form factor for augmented reality (AR) near-eye displays (NEDs). Unfortunately, the eyebox of such NEDs is impractically small (~ <; 1mm). In this paper, we introduce and demonstrate a design for holographic NEDs with a practical, wide eyebox of ~ 10mm and without any moving parts, based on holographic lenslets. In our design, a holographic optical element (HOE) based on a lenslet array was fabricated as the image combiner with expanded eyebox. A phase spatial light modulator (SLM) alters the phase of the incident laser light projected onto the HOE combiner such that the virtual image can be perceived at different focus distances, which can reduce the vergence-accommodation conflict (VAC). We have successfully implemented a bench-top prototype following the proposed design. The experimental results show effective eyebox expansion to a size of ~ 10mm. With further work, we hope that these design concepts can be incorporated into eyeglass-size NEDs. Xinxing Xia, Frank Guan, Andrei State, Praneeth Chakravarthula, Tat-Jen Cham, Henry Fuchs |
ISMAR | 5 |
| 2020 | Recovering facial reflectance and geometry from multi-view images
Guoxian Song, Jianmin Zheng, Jianfei Cai 0001, Tat-Jen Cham |
Image Vis. Comput. | 4 |
| 2019 | Pluralistic Image CompletionabstractMost image completion methods produce only one result for each masked input, although there may be many reasonable possibilities. In this paper, we present an approach for pluralistic image completion - the task of generating multiple and diverse plausible solutions for image completion. A major challenge faced by learning-based approaches is that usually only one ground truth training instance per label. As such, sampling from conditional VAEs still leads to minimal diversity. To overcome this, we propose a novel and probabilistically principled framework with two parallel paths. One is a reconstructive path that utilizes the only one given ground truth to get prior distribution of missing parts and rebuild the original image from this distribution. The other is a generative path for which the conditional prior is coupled to the distribution obtained in the reconstructive path. Both are supported by GANs. We also introduce a new short+long term attention layer that exploits distant relations among decoder and encoder features, improving appearance consistency. When tested on datasets with buildings (Paris), faces (CelebA-HQ), and natural images (ImageNet), our method not only generated higher-quality completion results, but also with multiple and diverse plausible outputs. Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
CVPR | 2 |
| 2019 | Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksabstractDespite great progress in 3D pose estimation from single-view images or videos, it remains a challenging task due to the substantial depth ambiguity and severe self-occlusions. Motivated by the effectiveness of incorporating spatial dependencies and temporal consistencies to alleviate these issues, we propose a novel graph-based method to tackle the problem of 3D human body and 3D hand pose estimation from a short sequence of 2D joint detections. Particularly, domain knowledge about the human hand (body) configurations is explicitly incorporated into the graph convolutional operations to meet the specific demand of the 3D pose estimation. Furthermore, we introduce a local-to-global network architecture, which is capable of learning multi-scale features for the graph-based representations. We evaluate the proposed method on challenging benchmark datasets for both 3D hand pose estimation and 3D body pose estimation. Experimental results show that our method achieves state-of-the-art performance on both tasks. Yujun Cai, Liuhao Ge, Jun Liu 0036, Jianfei Cai 0001, Tat-Jen Cham, Junsong Yuan 0001, Nadia Magnenat-Thalmann |
ICCV | 5 |
| 2019 | Shading-Based Surface Recovery Using Subdivision-Based RepresentationabstractAbstract This paper presents subdivision‐based representations for both lighting and geometry in shape‐from‐shading. A very recent shading‐based method introduced a per‐vertex overall illumination model for surface reconstruction, which has advantage of conveniently handling complicated lighting condition and avoiding explicit estimation of visibility and varied albedo. However, due to its discrete nature, the per‐vertex overall illumination requires a large amount of memory and lacks intrinsic coherence. To overcome these problems, in this paper we propose to use classic subdivision to define the basic smooth lighting function and surface, and introduce additional independent variables into the subdivision to adaptively model sharp changes of illumination and geometry. Compared to previous works, the new model not only preserves the merits of the per‐vertex illumination model, but also greatly reduces the number of variables required in surface recovery and intrinsically regularizes the illumination vectors and the surface. These features make the new model very suitable for multi‐view stereo surface reconstruction under general, unknown illumination condition. Particularly, a variational surface reconstruction method built upon the subdivision representations for lighting and geometry is developed. The experiments on both synthetic and real‐world data sets have demonstrated that the proposed method can achieve memory efficiency and improve surface detail recovery. Teng Deng, Jianmin Zheng, Jianfei Cai 0001, Tat-Jen Cham |
Comput. Graph. Forum | 4 |
| 2019 | Conditional adversarial synthesis of 3D facial action units
Zhilei Liu, Guoxian Song, Jianfei Cai 0001, Tat-Jen Cham, Juyong Zhang |
Neurocomputing | 4 |
| 2019 | Visibility Constrained Generative Model for Depth-Based 3D Facial Pose TrackingabstractIn this paper, we propose a generative framework that unifies depth-based 3D facial pose tracking and face model adaptation on-the-fly, in the unconstrained scenarios with heavy occlusions and arbitrary facial expression variations. Specifically, we introduce a statistical 3D morphable model that flexibly describes the distribution of points on the surface of the face model, with an efficient switchable online adaptation that gradually captures the identity of the tracked subject and rapidly constructs a suitable face model when the subject changes. Moreover, unlike prior art that employed ICP-based facial pose estimation, to improve robustness to occlusions, we propose a ray visibility constraint that regularizes the pose based on the face model's visibility with respect to the input point cloud. Ablation studies and experimental results on Biwi and ICT-3DHP datasets demonstrate that the proposed framework is effective and outperforms completing state-of-the-art depth-based methods. Lu Sheng, Jianfei Cai 0001, Tat-Jen Cham, Vladimir Pavlovic 0001, King Ngi Ngan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Towards a Switchable AR/VR Near-eye Display with Accommodation-Vergence and Eyeglass Prescription SupportabstractIn this paper, we present our novel design for switchable AR/VR near-eye displays which can help solve the vergence-accommodation-conflict issue. The principal idea is to time-multiplex virtual imagery and real-world imagery and use a tunable lens to adjust focus for the virtual display and the see-through scene separately. With this novel design, prescription eyeglasses for near- and far-sighted users become unnecessary. This is achieved by integrating the wearer's corrective optical prescription into the tunable lens for both virtual display and see-through environment. We built a prototype based on the design, comprised of micro-display, optical systems, a tunable lens, and active shutters. The experimental results confirm that the proposed near-eye display design can switch between AR and VR and can provide correct accommodation for both. Xinxing Xia, Frank Guan, Andrei State, Praneeth Chakravarthula, Kishore Rathinavel, Tat-Jen Cham, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | Towards Efficient 3D Calibration for Different Types of Multi-view Autostereoscopic 3D DisplaysabstractA novel and efficient 3D calibration method for different types of autostereoscopic multi-view 3D displays is presented in this paper. In our method, a camera is placed at different locations within the viewing volume of a 3D display to capture a series of images that relate to the subset of light rays emitted by the 3D display and arriving at each of the camera positions. Gray code patterns modulate the images shown on the 3D display, helping to significantly reduce the number of images captured by the camera and thereby accelerate the process of calculating the correspondence relationship between the pixels on the 3D display and the locations of the capturing camera. The proposed calibration method has been successfully tested on two different types of multi-view 3D displays and can be easily generalized for calibrating other types of such displays. The experimental results show that this novel 3D calibration method can also be used to improve the image quality by reducing the frequently observed crosstalk that typically exists when multiple users are simultaneously viewing multi-view 3D displays from a range of viewing positions. Xinxing Xia, Frank Guan, Andrei State, Tat-Jen Cham, Henry Fuchs |
CGI | 4 |
| 2018 | T ^2 2 Net: Synthetic-to-Realistic Translation for Solving Single-Image Depth Estimation Tasks
Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
ECCV (7) | 2 |
| 2018 | Background Subtraction Based on Deep Pixel Distribution LearningabstractPrevious approaches to background subtraction typically address the problem by formulating a representation of the background, and comparing the background to new frames. In this work, we focus on the essence of background subtraction, which is the classification of a pixel's current observation in comparison to historical observations, and propose a Deep Pixel Distribution Learning (DPDL) model for background subtraction. In the DPDL model, a novel pixel-based feature, called the Random Permutation of Temporal Pixels (RPoTP), is used to represent the distribution of past observations for a particular pixel, in which the temporal correlation between observations is deliberately obfuscated. Subsequently a convolutional neural network (CNN) is used to learn the distribution for determining whether the current observation is foreground or background, with the random permutation enabling the framework to focus primarily on the distribution of observations, rather than be misled by learning spurious temporal correlations. In addition, the pixel-wise representation allows for a large number of RPoTP features to be captured even with a limited number of groundtruth frames, with the DPDL model being effective even with only a single groundtruth frame. The proposed framework is able to achieve promising results in diverse natural scenes, and a comprehensive evaluation on standard benchmarks demonstrates the superiority of our work to state-of-the-art methods. The source code ispublicly available at https://github.com/zhaochenqiu/DPDL Chenqiu Zhao, Tat-Jen Cham, Xinyu Ren, Jianfei Cai 0001, Haichen Zhu |
ICME | 2 |
| 2018 | Real-time 3D Face-Eye Performance Capture of a Person Wearing VR HeadsetabstractTeleconference or telepresence based on virtual reality (VR) head-mount display (HMD) device is a very interesting and promising application since HMD can provide immersive feelings for users. However, in order to facilitate face-to-face communications for HMD users, real-time 3D facial performance capture of a person wearing HMD is needed, which is a very challenging task due to the large occlusion caused by HMD. The existing limited solutions are very complex either in setting or in approach as well as lacking the performance capture of 3D eye gaze movement. In this paper, we propose a convolutional neural network (CNN) based solution for real-time 3D face-eye performance capture of HMD users without complex modification to devices. To address the issue of lacking training data, we generate massive pairs of HMD face-label dataset by data synthesis as well as collecting VR-IR eye dataset from multiple subjects. Then, we train a dense-fitting network for facial region and an eye gaze network to regress 3D eye model parameters. Extensive experimental results demonstrate that our system can efficiently and effectively produce in real time a vivid personalized 3D avatar with the correct identity, pose, expression and eye motion corresponding to the HMD user. Guoxian Song, Jianfei Cai 0001, Tat-Jen Cham, Jianmin Zheng, Juyong Zhang, Henry Fuchs |
ACM Multimedia | 3 |
| 2018 | SubdSH: Subdivision-based Spherical Harmonics Field for Real-time Shading-based Refinement under Challenging Unknown IlluminationabstractThis paper presents a spatial-varying illumination model for shading-based depth refinement that based on a smooth Spherical Harmonics (SH) lighting field. The proposed lighting model is able to recover shading under challenging unknown lighting conditions, thus improving the quality of recovered surface detail. To avoid over-parameterization, local lighting coefficients are treated as a vector-valued function which is represented by subdivided surfaces using Catmull-Clark subdivision. We solve our lighting model utilizing a highly parallelized scheme that recovers lighting in a few milliseconds. A real-time shading-based depth recovery system is implemented with the integration of our proposed lighting model. We conduct quantitative and qualitative evaluations on both synthetic and real world datasets under challenging illumination. The experimental results show our method outperforms the state-of-the-art real-time shading-based depth refinement system. Teng Deng, Jianmin Zheng, Jianfei Cai 0001, Tat-Jen Cham |
VCIP | 4 |
| 2018 | Shading-Based Surface Detail Recovery Under General Unknown IlluminationabstractReconstructing the shape of a 3D object from multi-view images under unknown, general illumination is a fundamental problem in computer vision. High quality reconstruction is usually challenging especially when fine detail is needed and the albedo of the object is non-uniform. This paper introduces vertex overall illumination vectors to model the illumination effect and presents a total variation (TV) based approach for recovering surface details using shading and multi-view stereo (MVS). Behind the approach are the two important observations: (1) the illumination over the surface of an object often appears to be piecewise smooth and (2) the recovery of surface orientation is not sufficient for reconstructing the surface, which was often overlooked previously. Thus we propose to use TV to regularize the overall illumination vectors and use visual hull to constrain partial vertices. The reconstruction is formulated as a constrained TV-minimization problem that simultaneously treats the shape and illumination vectors as unknowns. An augmented Lagrangian method is proposed to quickly solve the TV-minimization problem. As a result, our approach is robust, stable and is able to efficiently recover high-quality surface details even when starting with a coarse model obtained using MVS. These advantages are demonstrated by extensive experiments on the state-of-the-art MVS database, which includes challenging objects with varying albedo. Di Xu 0012, Qi Duan, Jianmin Zheng, Juyong Zhang, Jianfei Cai 0001, Tat-Jen Cham |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2018 | Structure-Aware Multimodal Feature Fusion for RGB-D Scene Classification and BeyondabstractWhile convolutional neural networks (CNNs) have been excellent for object recognition, the greater spatial variability in scene images typically means that the standard full-image CNN features are suboptimal for scene classification. In this article, we investigate a framework allowing greater spatial flexibility, in which the Fisher vector (FV)-encoded distribution of local CNN features, obtained from a multitude of region proposals per image, is considered instead. The CNN features are computed from an augmented pixel-wise representation consisting of multiple modalities of RGB, HHA, and surface normals, as extracted from RGB-D data. More significantly, we make two postulates: (1) component sparsity—that only a small variety of region proposals and their corresponding FV GMM components contribute to scene discriminability, and (2) modal nonsparsity—that features from all modalities are encouraged to coexist. In our proposed feature fusion framework, these are implemented through regularization terms that apply group lasso to GMM components and exclusive group lasso across modalities. By learning and combining regressors for both proposal-based FV features and global CNN features, we are able to achieve state-of-the-art scene classification performance on the SUNRGBD Dataset and NYU Depth Dataset V2. Moreover, we further apply our feature fusion framework on an action recognition task to demonstrate that our framework can be generalized for other multimodal well-structured features. In particular, for action recognition, we enforce interpart sparsity to choose more discriminative body parts, and intermodal nonsparsity to make informative features from both appearance and motion modalities coexist. Experimental results on the JHMDB and MPII Cooking Datasets show that our feature fusion is also very effective for action recognition, achieving very competitive performance compared with the state of the art. Anran Wang 0001, Jianfei Cai 0001, Jiwen Lu, Tat-Jen Cham |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | A Generative Model for Depth-Based Robust 3D Facial Pose TrackingabstractWe consider the problem of depth-based robust 3D facial pose tracking under unconstrained scenarios with heavy occlusions and arbitrary facial expression variations. Unlike the previous depth-based discriminative or data-driven methods that require sophisticated training or manual intervention, we propose a generative framework that unifies pose tracking and face model adaptation on-the-fly. Particularly, we propose a statistical 3D face model that owns the flexibility to generate and predict the distribution and uncertainty underlying the face model. Moreover, unlike prior arts employing the ICP-based facial pose estimation, we propose a ray visibility constraint that regularizes the pose based on the face models visibility against the input point cloud, which augments the robustness against the occlusions. The experimental results on Biwi and ICT-3DHP datasets reveal that the proposed framework is effective and outperforms the state-of-the-art depth-based methods. Lu Sheng, Jianfei Cai 0001, Tat-Jen Cham, Vladimir Pavlovic 0001, King Ngi Ngan |
CVPR | 3 |
| 2017 | FaceCollage: A Rapidly Deployable System for Real-time Head Reconstruction for On-The-Go 3D TelepresenceabstractThis paper presents FaceCollage, a robust and real-time system for head reconstruction that can be used to create easy-to-deploy telepresence systems, using a pair of consumer-grade RGBD cameras that provide a wide range of views of the reconstructed user. A key feature is that the system is very simple to rapidly deploy, with autonomous calibration and requiring minimal intervention from the user, other than casually placing the cameras. This system is realized through three technical contributions: (1) a fully automatic calibration method, which analyzes and correlates the left and right RGBD faces just by the face features; (2) an implementation that exploits the parallel computation capability of GPU throughout most of the system pipeline, in order to attain real-time performance; and (3) a complete integrated system on which we conducted various experiments to demonstrate its capability, robustness, and performance, including testing the system on twelve participants with visually-pleasing results. Fuwen Tan, Chi-Wing Fu, Teng Deng, Jianfei Cai 0001, Tat-Jen Cham |
ACM Multimedia | 5 |
| 2017 | Multiple consumer-grade depth camera registration using everyday objects
Teng Deng, Jianfei Cai 0001, Tat-Jen Cham, Jianmin Zheng |
Image Vis. Comput. | 3 |
| 2016 | Modality and Component Aware Feature Fusion for RGB-D Scene ClassificationabstractWhile convolutional neural networks (CNN) have been excellent for object recognition, the greater spatial variability in scene images typically meant that the standard full-image CNN features are suboptimal for scene classification. In this paper, we investigate a framework allowing greater spatial flexibility, in which the Fisher vector (FV) encoded distribution of local CNN features, obtained from a multitude of region proposals per image, is considered instead. The CNN features are computed from an augmented pixel-wise representation comprising multiple modalities of RGB, HHA and surface normals, as extracted from RGB-D data. More significantly, we make two postulates: (1) component sparsity - that only a small variety of region proposals and their corresponding FV GMM components contribute to scene discriminability, and (2) modal non-sparsity - within these discriminative components, all modalities have important contribution. In our framework, these are implemented through regularization terms applying group lasso to GMM components and exclusive group lasso across modalities. By learning and combining regressors for both proposal-based FV features and global CNN features, we were able to achieve state-of-the-art scene classification performance on the SUNRGBD Dataset and NYU Depth Dataset V2. Anran Wang 0001, Jianfei Cai 0001, Jiwen Lu, Tat-Jen Cham |
CVPR | 4 |
| 2016 | Robust real-time performance-driven 3D face trackingabstractWe introduce a novel robust hybrid 3D face tracking framework from RGBD video streams, which is capable of tracking head pose and facial actions without pre-calibration or intervention from a user. In particular, we emphasize on improving the tracking performance in instances where the tracked subject is at a large distance from the cameras, and the quality of point cloud deteriorates severely. This is accomplished by the combination of a flexible 3D shape regressor and the joint 2D+3D optimization on shape parameters. Our approach fits facial blendshapes to the point cloud of the human head, while being driven by an efficient and rapid 3D shape regressor trained on generic RGB datasets. As an on-line tracking system, the identity of the unknown user is adapted on-the-fly resulting in improved 3D model reconstruction and consequently better tracking performance. The result is a robust RGBD face tracker capable of handling a wide range of target scene depths, whose performances are demonstrated in our extensive experiments better than those of the state-of-the-arts. Hai Xuan Pham, Vladimir Pavlovic 0001, Jianfei Cai 0001, Tat-Jen Cham |
ICPR | 4 |
| 2015 | Objects co-segmentation: Propagated from simpler imagesabstractRecent works on image co-segmentation aim to segment common objects among image sets. These methods can co-segment simple images well, but their performance may degrade significantly on more cluttered images. In order to co-segment both simple and complex images well, this paper proposes a novel paradigm to rank images and to propagate the segmentation results from the simple images to more and more complex ones. In the experiments, the proposed paradigm demonstrates its effectiveness in segmenting large image sets with a wide variety in object appearance, sizes, orientations, poses, and multiple objects in one image. It outperformed the current state-of-the-art algorithms significantly, especially in difficult images. Marcus Chen, Santiago Velasco-Forero, Ivor W. Tsang, Tat-Jen Cham |
ICASSP | 4 |
| 2015 | MMSS: Multi-modal Sharable and Specific Feature Learning for RGB-D Object RecognitionabstractMost of the feature-learning methods for RGB-D object recognition either learn features from color and depth modalities separately, or simply treat RGB-D as undifferentiated four-channel data, which cannot adequately exploit the relationship between different modalities. Motivated by the intuition that different modalities should contain not only some modal-specific patterns but also some shared common patterns, we propose a multi-modal feature learning framework for RGB-D object recognition. We first construct deep CNN layers for color and depth separately, and then connect them with our carefully designed multi-modal layers, which fuse color and depth information by enforcing a common part to be shared by features of different modalities. In this way, we obtain features reflecting shared properties as well as modal-specific properties in different modalities. The information of the multi-modal learning frameworks is back-propagated to the early CNN layers. Experimental results show that our proposed multi-modal feature learning method outperforms state-of-the-art approaches on two widely used RGB-D object benchmark datasets. Anran Wang 0001, Jianfei Cai 0001, Jiwen Lu, Tat-Jen Cham |
ICCV | 4 |
| 2015 | Unsupervised Joint Feature Learning and Encoding for RGB-D Scene LabelingabstractMost existing approaches for RGB-D indoor scene labeling employ hand-crafted features for each modality independently and combine them in a heuristic manner. There has been some attempt on directly learning features from raw RGB-D data, but the performance is not satisfactory. In this paper, we propose an unsupervised joint feature learning and encoding (JFLE) framework for RGB-D scene labeling. The main novelty of our learning framework lies in the joint optimization of feature learning and feature encoding in a coherent way, which significantly boosts the performance. By stacking basic learning structure, higher level features are derived and combined with lower level features for better representing RGB-D data. Moreover, to explore the nonlinear intrinsic characteristic of data, we further propose a more general joint deep feature learning and encoding (JDFLE) framework that introduces the nonlinear mapping into JFLE. The experimental results on the benchmark NYU depth dataset show that our approaches achieve competitive performance, compared with the state-of-the-art methods, while our methods do not need complex feature handcrafting and feature combination and can be easily applied to other data sets. Anran Wang 0001, Jiwen Lu, Jianfei Cai 0001, Gang Wang 0012, Tat-Jen Cham |
IEEE Trans. Image Process. | 5 |
| 2015 | Kinect Depth Recovery Using a Color-Guided, Region-Adaptive, and Depth-Selective FrameworkabstractConsidering that the existing depth recovery approaches have different limitations when applied to Kinect depth data, in this article, we propose to integrate their effective features including adaptive support region selection, reliable depth selection, and color guidance together under an optimization framework for Kinect depth recovery. In particular, we formulate our depth recovery as an energy minimization problem, which solves the depth hole filling and denoising simultaneously. The energy function consists of a fidelity term and a regularization term, which are designed according to the Kinect characteristics. Our framework inherits and improves the idea of guided filtering by incorporating structure information and prior knowledge of the Kinect noise model. Through analyzing the solution to the optimization framework, we also derive a local filtering version that provides an efficient and effective way of improving the existing filtering techniques. Quantitative evaluations on our developed synthesized dataset and experiments on real Kinect data show that the proposed method achieves superior performance in terms of recovery accuracy and visual quality. Chongyu Chen, Jianfei Cai 0001, Jianmin Zheng, Tat-Jen Cham, Guangming Shi |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2015 | A Unified Feature Selection Framework for Graph Embedding on High Dimensional DataabstractAlthough graph embedding has been a powerful tool for modeling data intrinsic structures, simply employing all features for data structure discovery may result in noise amplification. This is particularly severe for high dimensional data with small samples. To meet this challenge, this paper proposes a novel efficient framework to perform feature selection for graph embedding, in which a category of graph embedding methods is cast as a least squares regression problem. In this framework, a binary feature selector is introduced to naturally handle the feature cardinality in the least squares formulation. The resultant integral programming problem is then relaxed into a convex Quadratically Constrained Quadratic Program (QCQP) learning problem, which can be efficiently solved via a sequence of accelerated proximal gradient (APG) methods. Since each APG optimization is w.r.t. only a subset of features, the proposed method is fast and memory efficient. The proposed framework is applied to several graph embedding learning problems, including supervised, unsupervised, and semi-supervised graph embedding. Experimental results on several high dimensional data demonstrated that the proposed method outperformed the considered state-of-the-art methods. Marcus Chen, Ivor W. Tsang, Mingkui Tan, Tat-Jen Cham |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Large-Margin Multi-Modal Deep Learning for RGB-D Object RecognitionabstractMost existing feature learning-based methods for RGB-D object recognition either combine RGB and depth data in an undifferentiated manner from the outset, or learn features from color and depth separately, which do not adequately exploit different characteristics of the two modalities or utilize the shared relationship between the modalities. In this paper, we propose a general CNN-based multi-modal learning framework for RGB-D object recognition. We first construct deep CNN layers for color and depth separately, which are then connected with a carefully designed multi-modal layer. This layer is designed to not only discover the most discriminative features for each modality, but is also able to harness the complementary relationship between the two modalities. The results of the multi-modal layer are back-propagated to update parameters of the CNN layers, and the multi-modal feature learning and the back-propagation are iteratively performed until convergence. Experimental results on two widely used RGB-D object datasets show that our method for general multi-modal learning achieves comparable performance to state-of-the-art methods specifically designed for RGB-D data. Anran Wang 0001, Jiwen Lu, Jianfei Cai 0001, Tat-Jen Cham, Gang Wang 0012 |
IEEE Trans. Multim. | 4 |
| 2014 | Recovering Surface Details under General Unknown Illumination Using Shading and Coarse Multi-view StereoabstractSummary form only given. Reconstructing the shape of a 3D object from multi-view images under unknown, general illumination is a fundamental problem in computer vision and high quality reconstruction is usually challenging especially when high detail is needed. This paper presents a total variation (TV) based approach for recovering surface details using shading and multi-view stereo (MVS). Behind the approach are our two important observations: (1) the illumination over the surface of an object tends to be piecewise smooth and (2) the recovery of surface orientation is not sufficient for reconstructing geometry, which were previously overlooked. Thus we introduce TV to regularize the lighting and use visual hull to constrain partial vertices. The reconstruction is formulated as a constrained TV minimization problem that treats the shape and lighting as unknowns simultaneously. An augmented Lagrangian method is proposed to quickly solve the TV-minimization problem. As a result, our approach is robust, stable and is able to efficiently recover high quality of surface details even starting with a coarse MVS. These advantages are demonstrated by the experiments with synthetic and real world examples. Di Xu 0012, Qi Duan, Jianming Zheng, Juyong Zhang, Jianfei Cai 0001, Tat-Jen Cham |
CVPR | 6 |
| 2014 | Multi-modal Unsupervised Feature Learning for RGB-D Scene Labeling
Anran Wang 0001, Jiwen Lu, Gang Wang 0012, Jianfei Cai 0001, Tat-Jen Cham |
ECCV (5) | 5 |
| 2014 | Estimating spatial layout of rooms from RGB-D videosabstractSpatial layout estimation of indoor rooms plays an important role in many visual analysis applications such as robotics and human-computer interaction. While many methods have been proposed for recovering spatial layout of rooms in recent years, their performance is still far from satisfactory due to high occlusion caused by the presence of objects that clutter the scene. In this paper, we propose a new approach to estimate the spatial layout of rooms from RGB-D videos. Unlike most existing methods which estimate the layout from still images, RGB-D videos provide more spatial-temporal and depth information, which are helpful to improve the estimation performance because more contextual information can be exploited in RGB-D videos. Given a RGB-D video, we first estimate the spatial layout of the scene in each single frame and compute the camera trajectory using the simultaneous localization and mapping (SLAM) algorithm. Then, the estimated spatial layouts of different frames are integrated to infer temporally consistent layouts of the room throughout the whole video. Our method is evaluated on the NYU RGB-D dataset, and the experimental results show the efficacy of the proposed approach. Anran Wang 0001, Jiwen Lu, Jianfei Cai 0001, Gang Wang 0012, Tat-Jen Cham |
MMSP | 5 |
| 2013 | High-quality Kinect depth filtering for real-time 3D telepresenceabstract3D telepresence is a next-generation multimedia application, offering remote users an immersive and natural video-conferencing environment with real-time 3D graphics. Kinect sensor, a consumer-grade range camera, facilitates the implementation of some recent 3D telepresence systems. However, conventional data filtering methods are insufficient to handle Kinect depth error because such error is quantized rather than just randomly-distributed. Hence, one could often observe large irregularly-shaped patches of pixels that receive the same depth values from Kinect. To enhance visual quality in 3D telepresence, we propose a novel depth data filtering method for Kinect by means of multi-scale and direction-aware support windows. In addition, we develop a GPU-based CUDA implementation that can perform real-time depth filtering. Results from the experiments show that our method can reconstruct hole-free surfaces that are smoother and less bumpy compared to existing methods like bilateral filtering. Mengyao Zhao, Fuwen Tan, Chi-Wing Fu, Chi-Keung Tang, Jianfei Cai 0001, Tat-Jen Cham |
ICME | 6 |
| 2013 | A color-guided, region-adaptive and depth-selective unified framework for Kinect depth recoveryabstractConsidering the existing depth recovery approaches that have different limitations when applying to Kinect depth data, in this paper, we propose to integrate their effective features including adaptive support region selection, reliable depth selection and color guidance together under a unified framework for Kinect depth recovery. In particular, we formulate our depth recovery as an energy minimization problem, which solves the depth hole-filling and denoising simultaneously. The energy function consists of a fidelity term and a regularization term. The fidelity term takes into account the characteristics of Kinect data. The regularization term is designed to incorporate the joint bilateral filtering (JBF) kernel and the joint trilateral filtering (JTF) kernel so as to facilitate both depth hole-filling and denoising. Moreover, the JBF kernel is modified to incorporate the structure information. Both simulations on the benchmark Middlebury dataset and experiments on real Kinect data show that our proposed method achieves state-of-the-art performance in terms of recovery accuracy and visual quality. Chongyu Chen, Jianfei Cai 0001, Jianmin Zheng, Tat-Jen Cham, Guangming Shi |
MMSP | 4 |
| 2011 | Visual tracking with generative template model based on Riemannian manifold of covariances
Marcus Chen, Sze Kim Pang, Tat-Jen Cham, Alvina Goh |
FUSION | 3 |
| 2010 | Estimating camera pose from a single urban ground-view omnidirectional image and a 2D building outline mapabstractA framework is presented for estimating the pose of a camera based on images extracted from a single omnidirectional image of an urban scene, given a 2D map with building outlines with no 3D geometric information nor appearance data. The framework attempts to identify vertical corner edges of buildings in the query image, which we term VCLH, as well as the neighboring plane normals, through vanishing point analysis. A bottom-up process further groups VCLH into elemental planes and subsequently into 3D structural fragments modulo a similarity transformation. A geometric hashing lookup allows us to rapidly establish multiple candidate correspondences between the structural fragments and the 2D map building contours. A voting-based camera pose estimation method is then employed to recover the correspondences admitting a camera pose solution with high consensus. In a dataset that is even challenging for humans, the system returned a top-30 ranking for correct matches out of 3600 camera pose hypotheses (0.83% selectivity) for 50.9% of queries. Tat-Jen Cham, Arridhana Ciptadi, Wei-Chian Tan, Minh-Tri Pham, Liang-Tien Chia |
CVPR | 1 |
| 2010 | Fast polygonal integration and its application in extending haar-like features to improve object detectionabstractThe integral image is typically used for fast integrating a function over a rectangular region in an image. We propose a method that extends the integral image to do fast integration over the interior of any polygon that is not necessarily rectilinear. The integration time of the method is fast, independent of the image resolution, and only linear to the polygon's number of vertices. We apply the method to Viola and Jones' object detection framework, in which we propose to improve classical Haar-like features with polygonal Haar-like features. We show that the extended feature set improves object detection's performance. The experiments are conducted in three domains: frontal face detection, fixed-pose hand detection, and rock detection for Mars' surface terrain assessment. Minh-Tri Pham, Yang Gao 0002, Viet-Dung Hoang, Tat-Jen Cham |
CVPR | 4 |
| 2010 | Face and Human Gait Recognition Using Image-to-Class DistanceabstractWe propose a new distance measure for face recognition and human gait recognition. Each probe image (a face image or an average human silhouette image) is represented as a set of local features uniformly sampled over a grid with fixed spacing, and each gallery image is represented as a set of local features sampled at each pixel. We formulate an integer programming problem to compute the distance (referred to as the image-to-class distance) from one probe image to all the gallery images belonging to a certain class, in which any feature of the probe image can be matched to only one feature from one of the gallery images. Considering computational efficiency as well as the fact that face images or average human silhouette images are roughly aligned in the preprocessing step, we also enforce a spatial neighborhood constraint by only allowing neighboring features that are within a given spatial distance to be considered for feature matching. The integer programming problem is further treated as a classical minimum-weight bipartite graph matching problem, which can be efficiently solved with the Kuhn-Munkres algorithm. We perform comprehensive experiments on three benchmark face databases: 1) the CMU PIE database; 2) the FERET database; and 3) the FRGC database, as well as the USF Human ID gait database. The experiments clearly demonstrate the effectiveness of our image-to-class distance. Yi Huang 0014, Dong Xu 0001, Tat-Jen Cham |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Near Duplicate Identification With Spatially Aligned Pyramid MatchingabstractA new framework, termed spatially aligned pyramid matching, is proposed for near duplicate image identification. The proposed method robustly handles spatial shifts as well as scale changes, and is extensible for video data. Images are divided into both overlapped and non-overlapped blocks over multiple levels. In the first matching stage, pairwise distances between blocks from the examined image pair are computed using earth mover's distance (EMD) or the visual word with$\chi^{2}$distance based method with scale-invariant feature transform (SIFT) features. In the second stage, multiple alignment hypotheses that consider piecewise spatial shifts and scale variation are postulated and resolved using integer-flow EMD. Moreover, to compute the distances between two videos, we conduct the third step matching (i.e., temporal matching) after spatial matching. Two application scenarios are addressed—near duplicate retrieval (NDR) and near duplicate detection (NDD). For retrieval ranking, a pyramid-based scheme is constructed to fuse matching results from different partition levels. For NDD, we also propose a dual-sample approach by using the multilevel distances as features and support vector machine for binary classification. The proposed methods are shown to clearly outperform existing methods through extensive testing on the Columbia Near Duplicate Image Database and two new datasets. In addition, we also discuss in depth our framework in terms of the extension for video NDR and NDD, the sensitivity to parameters, the utilization of multiscale dense SIFT descriptors, and the test of scalability in image NDD. Dong Xu 0001, Tat-Jen Cham, Shuicheng Yan, Lixin Duan, Shih-Fu Chang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Detection with multi-exit asymmetric boostingabstractWe introduce a generalized representation for a boosted classifier with multiple exit nodes, and propose a method to training which combines the idea of propagating scores across boosted classifiers [14, 17] and the use of asymmetric goals [13]. A means for determining the ideal constant asymmetric goal is provided, which is theoretically justified under a conservative bound on the ROC operating point target and empirically near-optimal under the exact bound. Moreover, our method automatically minimizes the number of weak classifiers, avoiding the need to retrain a boosted classifier multiple times for empirical best performance as in conventional methods. Experimental results shows significant reduction in training time and number of weak classifiers, as well as better accuracy, compared to conventional cascades and multi-exit boosted classifiers. Minh-Tri Pham, V-D. D. Hoang, Tat-Jen Cham |
CVPR | 3 |
| 2008 | Near duplicate image identification with patially Aligned Pyramid MatchingabstractA new framework, termed spatially aligned pyramid matching, is proposed for near duplicate image identification. The proposed method robustly handles spatial shifts as well as scale changes. Images are divided into both overlapped and non-overlapped blocks over multiple levels. In the first matching stage, pairwise distances between blocks from the examined image pair are computed using SIFT features and Earth Moverpsilas distance (EMD). In the second stage, multiple alignment hypotheses that consider piecewise spatial shifts and scale variation are postulated and resolved using integer-flow EMD. Two application scenarios are addressed - retrieval ranking and binary classification. For retrieval ranking, a pyramid-based scheme is constructed to fuse matching results from different partition levels. For binary classification, a novel generalized neighborhood component analysis method is formulated that can be effectively used in tandem with SVMs to select the most critical matching components. The proposed methods are shown to clearly outperform existing methods through extensive testing on the Columbia near duplicate image database and another new dataset. Dong Xu 0001, Tat-Jen Cham, Shuicheng Yan, Shih-Fu Chang |
CVPR | 2 |
| 2007 | High Distortion and Non-Structural Image Matching via Feature Co-occurrenceabstractWe propose a novel approach for determining if a pair of images match each other under the effect of a high-distortion transformation or non-structural relation. The co-occurrence statistics between features across a pair of images are learned from a training set comprising matched and mismatched image pairs - these are expressed in the form of a cross-feature ratio table. The proposed method does not require feature-to-feature correspondences, but instead identifies and exploits feature co-occurrences that are able to provide discriminative result from the transformation. The method not only allows for the matching of test image pairs that have substantially different visual content as compared to those present in the training set, but also caters for transformations and relations that do not preserve image structure. Xi Chen 0069, Tat-Jen Cham |
CVPR | 2 |
| 2007 | Online Learning Asymmetric Boosted Classifiers for Object DetectionabstractWe present an integrated framework for learning asymmetric boosted classifiers and online learning to address the problem of online learning asymmetric boosted classifiers, which is applicable to object detection problems. In particular, our method seeks to balance the skewness of the labels presented to the weak classifiers, allowing them to be trained more equally. In online learning, we introduce an extra constraint when propagating the weights of the data points from one weak classifier to another, allowing the algorithm to converge faster. In compared with the Online Boosting algorithm recently applied to object detection problems, we observed about 0-10% increase in accuracy, and about 5-30% gain in learning speed. Minh-Tri Pham, Tat-Jen Cham |
CVPR | 2 |
| 2007 | Fast training and selection of Haar features using statistics in boosting-based face detectionabstractTraining a cascade-based face detector using boosting and Haar features is computationally expensive, often requiring weeks on single CPU machines. The bottleneck is at training and selecting Haar features for a single weak classifier, currently in minutes. Traditional techniques for training a weak classifier usually run in 0(NT log N), with N examples (approximately 10,000), and T features (approximately 40,000). We present a method to train a weak classifier in time 0(Nd2+ T), where d is the number of pixels of the probed image sub-window (usually from 350 to 500), by using only the statistics of the weighted input data. Experimental results revealed a significantly reduced training time of a weak classifier to the order of seconds. In particular, this method suffers very minimal immerse in training time with very large increases in members of Haar features, enjoying a significant gain in accuracy, even with reduced training time. Minh-Tri Pham, Tat-Jen Cham |
ICCV | 2 |
| 2007 | Click4BuildingID@NTU: Click for Building Identification with GPS-enabled Camera Cell PhoneabstractA working prototype of a building identification service which can be used on any camera cell phones equipped with GPS capability has been developed. Users can simply snap photos of architectures and send them, together with the corresponding GPS coordinates, via MMS to a remote server. The server will match the photos with the stored, GPS-tagged images using a combination of scale saliency algorithm for feature matching and earth movers distance measure for scene matching. The estimated location and other information are then sent back to the users via MMS. This prototype will have better accuracy than systems which rely solely on photo recognition given the exploitation of GPS information. Moreover, it is computationally lighter since the recognition engine only needs to compare stored images which lie within the GPS coordinates error range. It is relatively inexpensive as no special phones or subscriptions to telecommunication providers for provision of GPS equivalent location data (i.e.cell location) are needed. Chai Kiat Yeo, Liang-Tien Chia, Tat-Jen Cham, D. Rajon |
ICME | 3 |
| 2007 | Shadow Elimination and Blinding Light Suppression for Interactive Projected DisplaysabstractA major problem with interactive displays based on front projection is that users cast undesirable shadows on the display surface. This paper demonstrates that shadows can be muted by redundantly illuminating the display surface using multiple projectors, all mounted at different locations. However, this technique alone does not eliminate shadows: Multiple projectors create multiple dark regions on the surface (penumbral occlusions) and cast undesirable light onto the users. These problems can be solved by eliminating shadows and suppressing the light that falls on occluding users by actively modifying the projected output. This paper categorizes various methods that can be used to achieve redundant illumination, shadow elimination, and blinding light suppression and evaluates their performance. Jay Summet, Matthew Flagg, Tat-Jen Cham, James M. Rehg, Rahul Sukthankar |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2006 | Alignment of 3D Models to Images Using Region-Based Mutual Information and Neighborhood Extended Gaussian Images
Hon-Keat Pong, Tat-Jen Cham |
ACCV (1) | 2 |
| 2006 | Object Detection Using a Cascade of 3D Models
Hon-Keat Pong, Tat-Jen Cham |
ACCV (2) | 2 |
| 2006 | Image Pre-Conditioning for Out-of-Focus Projector BlurabstractWe present a technique to reduce image blur caused by out-of-focus regions in projected imagery. Unlike traditional restoration algorithms that operate on a blurred image to recover the original, the nature of our problem requires that the correction be applied to the original image before blurring. To accomplish this, a camera is used to estimate a series of spatially varying point-spread-functions (PSF) across the projector’s image. These discrete PSFs are then used to guide a pre-processing algorithm based on Wiener filtering to condition the image before projection. Results show that using this technique can help ameliorate the visual effects from out-of-focus projector blur. Michael S. Brown, Peng Song 0010, Tat-Jen Cham |
CVPR (2) | 3 |
| 2005 | Learning Feature Distance Measures for Image CorrespondencesabstractStandard but ad hoc measures such as sum-of-squared pixel differences (SSD) are often used when comparing and registering two images that have not been previously observed before. In this paper, we propose a framework to address the problem of learning a parametric feature distance measure to measure the dissimilarity between pairs of images. The method is based on optimizing the parameters of the distance measure in order to minimize correspondence classification errors on training data. Because the learning process involves relative (rather than absolute) visual content between image pairs, the learned distance measure may also be applied to other images with very different visual content. Results on matching classification with a wide variety of image content show that the learned feature distance measure clearly outperforms the standard measures of SSD, chamfer and Bhattacharyya histogram distances. Xi Chen 0069, Tat-Jen Cham |
CVPR (2) | 2 |
| 2003 | Shadow Elimination and Occluder Light Suppression for Multi-Projector DisplaysabstractTwo related problems of front projection displays, which occur when users obscure a projector, are: (i) undesirable shadows cast on the display by the users, and (ii) projected light falling on and distracting the users. This paper provides a computational framework for solving these two problems based on multiple overlapping projectors and cameras. The overlapping projectors are automatically aligned to display the same dekeystoned image. The system detects when and where shadows are cast by occluders and is able to determine the pixels, which are occluded in different projectors. Through a feedback control loop, the contributions of unoccluded pixels from other projectors are boosted in the shadowed regions, thereby eliminating the shadows. In addition, pixels, which are being occluded, are blanked, thereby preventing the projected light from falling on a user when they occlude the display. This can be accomplished even when the occluders are not visible to the camera. The paper presents results from a number of experiments demonstrating that the system converges rapidly with low steady-state errors. Tat-Jen Cham, James M. Rehg, Rahul Sukthankar, Gita Reese Sukthankar |
CVPR (2) | 1 |
| 2002 | Analogous view transfer for gaze correction in video sequencesabstractThis paper provides a framework for doing facial gaze correction in video sequences. The proposed framework involves stages of face registration, face parameter mapping, and face synthesis. We introduce the concept of analogous views, and derive a novel formulation which extends view transfers based on epipolar geometry to cope with non-rigid motion. Additionally, a disparity mapping function is derived which is learned from training data and handles both spatial disparities as well as pixel-value changes. The disparity mapping function generalizes to facial expressions, illumination conditions and individuals not in the training set, as shown by the results obtained. Tat-Jen Cham, Shyam Krishnamoorthy |
ICARCV | 1 |
| 2002 | Projected light displays using visual feedbackabstractA system of coordinated projectors and cameras enables the creation of projected light displays that are robust to environmental disturbances. This paper describes approaches for tackling both geometric and photometric aspects of the problem: (1) the projected image remains stable even when the system components (projector, camera or screen) are moved; (2) the display automatically removes shadows caused by users moving between a projector and the screen, while simultaneously suppressing projected light on the user. The former can be accomplished without knowing the positions of the system components. The latter can be achieved without direct observation of the occluder. We demonstrate that the system responds quickly to environmental disturbances and achieves low steady-state errors. James M. Rehg, Matthew Flagg, Tat-Jen Cham, Rahul Sukthankar, Gita Reese Sukthankar |
ICARCV | 3 |
| 2001 | Reconstruction of 3-D Figure Motion from 2-D CorrespondencesabstractWe present a method for computing the 3D motion of articulated models from 2D correspondences. An iterative batch algorithm is proposed which estimates the maximum a posteriori trajectory based on 2D measurements subject to a number of constraints. These include (i) kinematic constraints based on a 3D kinematic model, (ii) joint angle limits, (iii) dynamic smoothing, and (iv) 3D key frames which can be specified by the user. The framework handles any variation in the number of constraints as well as partial or missing data. This method is shown to obtain favorable reconstruction results on a number of complex human motion sequences. David E. DiFranco, Tat-Jen Cham, James M. Rehg |
CVPR (1) | 2 |
| 2001 | Dynamic Shadow Elimination for Multi-Projector DisplaysabstractA major problem with interactive displays based on front-projection is that users cast undesirable shadows on the display surface. This situation is only partially addressed by mounting a single projector at an extreme angle and pre-warping the projected image to undo keystoning distortions. This paper demonstrates that shadows can be muted by redundantly illuminating the display surface using multiple projectors, all mounted at different locations. However, this technique alone does not eliminate shadows: multiple projectors create multiple dark regions on the surface (penumbral occlusions). We solve the problem by using cameras to automatically identify occlusions as they occur and dynamically adjust each projector's output so that additional light is projected onto each partially-occluded patch. The system is self-calibrating: relevant homographies relating projectors, cameras and the display surface are recovered by observing the distortions induced in projected calibration patterns. The resulting redundantly-projected display retains the high image quality of a single-projector system while dynamically correcting for all penumbral occlusions. Our initial two-projector implementation operates at 3 Hz. Rahul Sukthankar, Tat-Jen Cham, Gita Reese Sukthankar |
CVPR (2) | 2 |
| 2001 | Self-Calibrating Camera Projector Systems for Interactive Displays and PresentationsabstractThe authors demonstrate a self-calibrating system that employs uncalibrated cameras and microportable projectors to create novel interactive displays and presentations. Three benefits of ther system are detailed. Rahul Sukthankar, Tat-Jen Cham, Gita Reese Sukthankar, James M. Rehg, David Hsu, Thomas K. Leung |
ICCV | 2 |
| 2000 | Video Editing Using Figure Tracking and Image-Based RenderingabstractWe describe a new approach to video editing based on the semi-automatic segmentation of video into multiple layers and the composition of layers using image-based rendering. Using figure tracking and background motion estimation, we can segment a moving figure and reconstruct the background. Using geometrically-correct pixel reprojection, layers can be composited on the basis of the geometry of the underlying scene and the position of a virtual camera. We have implemented a prototype editing system called SpliceWorld. James M. Rehg, Sing Bing Kang, Tat-Jen Cham |
ICIP | 3 |
| 1999 | A Multiple Hypothesis Approach to Figure TrackingabstractThis paper describes a probabilistic multiple-hypothesis framework for tracking highly articulated objects. In this framework, the probability density of the tracker state is represented as a set of modes with piecewise Gaussians characterizing the neighborhood around these modes. The temporal evolution of the probability density is achieved through sampling from the prior distribution, followed by local optimization of the sample positions to obtain updated modes. This method of generating hypotheses from state-space search does not require the use of discrete features unlike classical multiple-hypothesis tracking. The parametric form of the model is suited for high dimensional state-spaces which cannot be efficiently modeled using non-parametric approaches. Results are shown for tracking Fred Astaire in a movie dance sequence. Tat-Jen Cham, James M. Rehg |
CVPR | 1 |
| 1999 | Dynamic Feature Ordering for Efficient RegistrationabstractExisting sequential feature based registration algorithms involving search typically either select features randomly (e.g. the RANSAC approach (M. Fischler and R. Bolles, 1981)) or assume a predefined, intuitive ordering for the features (e.g. based on size or resolution). The paper presents a formal framework for computing an ordering for features which maximizes search efficiency. Features are ranked according to matching ambiguity measure, and an algorithm is proposed which couples the feature selection with the parameter estimation, resulting in a dynamic feature ordering. The analysis is extended to template features where the matching is non discrete and a sample refinement process is proposed. The framework is demonstrated effectively on the localization of a person in an image, using a kinematic model with template features. Different priors are used on the model parameters and the results demonstrate nontrivial variations in the optimal feature hierarchy. Tat-Jen Cham, James M. Rehg |
ICCV | 1 |
| 1999 | A Dynamic Bayesian Network Approach to Figure Tracking using Learned Dynamic ModelsabstractThe human figure exhibits complex and rich dynamic behavior that is both nonlinear and time-varying. However most work on tracking and synthesizing figure motion has employed either simple, generic dynamic models or highly specific hand-tailored ones. Recently, a broad class of learning and inference algorithms for time-series models have been successfully cast in the framework of dynamic Bayesian networks (DBNs). This paper describes a novel DBN-based switching linear dynamic system (SLDS) model and presents its application to figure motion analysis. A key feature of our approach is an approximate Viterbi inference technique for overcoming the intractability of exact inference in mixed-state DBNs. We present experimental results for learning figure dynamics from video data and show promising initial results for tracking, interpolation, synthesis, and classification using learned models. Vladimir Pavlovic 0001, James M. Rehg, Tat-Jen Cham, Kevin Murphy 0002 |
ICCV | 3 |
| 1999 | Automated B-Spline Curve Representation Incorporating MDL and Error-Minimizing Control Point Insertion StrategiesabstractThe main issues of developing an automatic and reliable scheme for spline-fitting are discussed and addressed in this paper, which are not fully covered in previous papers or algorithms. The proposed method incorporates B-spline active contours, the minimum description length (MDL) principle, and a novel control point insertion strategy based on maximizing the potential for energy-reduction maximization (PERM). A comparison of test results shows that it outperforms one of the better existing methods. Tat-Jen Cham, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | A Statistical Framework for Long-Range Feature Matching in Uncalibrated Image MosaicingabstractThe problem considered is that of estimating the projective transformation between two images in situations where the image motion is large and feature matching is not aided by a proximity heuristic. The overall algorithm designed is based on a multiresolution, multihypothesis scheme, and similarities between tracking and matching through multiple resolution levels are exploited. Two major tools are developed in this paper: (i) a Bayesian framework for incorporating similarity measures of feature correspondences in regression to specify the different levels of confidence in the correspondences; and (ii) a Bayesian version of RANSAC, which is able to utilise prior estimates and matching probabilities. The algorithm is tested on a number of real images with large image motion and promising results were obtained. Tat-Jen Cham, Roberto Cipolla |
CVPR | 1 |
| 1997 | Stereo Coupled Active ContoursabstractWe consider how tracking in stereo may be enhanced by coupling pairs of active contours in different views via affine epipolar geometry and various subsets of planar affine transformations, as well as by implementing temporal constraints imposed by curve rigidity. 3D curve tracking is achieved using a submanifold model, where it is shown how the coupling mechanisms can be decomposed to cater for fired and variable epipolar geometries. In the case of tracking planar curves, the canonical frame model is developed such that the various geometrical constraints needed in different situations may be efficiently selected. The results show that coupled active contours add consistency and robustness to tracking in stereo. Tat-Jen Cham, Roberto Cipolla |
CVPR | 1 |
| 1996 | Automated B-Spline Curve Representation with MDL-based Active ContoursabstractPresent spline-fitting methods used in computer vision do not fully address the main issues of developing an automatic and reliable algorithm, which are discussed in this paper. A paradigm for spline fitting is proposed, with features of the algorithm selected such that the main issues are resolved. This is achieved through the use of Bspline active contours, the minimum description length principle, and in conjunction with a control point insertion strategy based on the Potential for Energy-Reduction Maximisation (PERM). This strategy selects control points such that the formation of compatible collapse mechanisms for the splines is encouraged. An implementation of the algorithm is carried and tested on various images. A comparison with one of the better existing methods for spline fitting demonstrates that there is considerable potential for the algorithm to outperform current algorithms. 1 Introduction Representing curves by analytic functions instead of sets of data points has man... Tat-Jen Cham, Roberto Cipolla |
BMVC | 1 |
| 1996 | Geometric Saliency of Curve Correspondances and Grouping of Symmetric Comntours
Tat-Jen Cham, Roberto Cipolla |
ECCV (1) | 1 |
| 1995 | Symmetry detection through local skewed symmetries
Tat-Jen Cham, Roberto Cipolla |
Image Vis. Comput. | 1 |
| 1994 | Skewed Symmetry Detection Through Local Skewed SymmetriesabstractWe explore how global symmetry can be detected prior to segmentation and under noise and occlusion. The definition of local symmetries is extended to affine geometries by considering the tangents and curvatures of local structures, and a quantitative measure of local symmetry known as symmetricity is introduced, which is based on Mahalanobis distances from the tangent-curvature states of local structures to the local skewed symmetry state-subspace. These symmetricity values, together with the associated local axes of symmetry, are spatially related in the local skewed symmetry field (LSSF). In the implementation, a fast, local symmetry detection algorithm allows initial hypotheses for the symmetry axis to be generated through the use of a modified Hough transform. This is then improved upon by maximising a global symmetry measure based on accumulated local support in the LSSF — a straight active contour model is used for this purpose. This produces useful estimates for the axis of symmetry and the angle of skew in the presence of contour fragmentation, artifacts and occlusion. Tat-Jen Cham, Roberto Cipolla |
BMVC | 1 |
| 1994 | A local approach to recovering global skewed symmetryabstractA local approach is adopted to allow recovery of global skewed symmetry in the presence of occlusion. Local skewed symmetries are established by extending the definition of local symmetries to affine geometries through the use of local derivatives. Symmetricity, a quantitative gauge of local symmetry based on Mahalanobis distances from the tangent-curvature states of local structures to the local skewed symmetry state-subspace, is also introduced to cope with noise. The symmetricity values and local symmetry axes for each pair of points are then spatially related in the local skewed symmetry field. The global symmetry detection algorithm implemented involves obtaining fast, initial estimates of the symmetry axis from a separate Hough transform technique, followed by maximising a global symmetry measure via a straight active contour model which is driven by effective symmetricity values. This produces useful estimates for the axis of symmetry and the angle of skew in the presence of contour fragmentation, artifacts and occlusion. Tat-Jen Cham, Roberto Cipolla |
ICPR (1) | 1 |