Wei-Chao Chen

dblp:37/1413 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-6165-7706ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 4 since 2021Systems, architecture and hardware · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-authorSecurity and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 How Bias Binds: Measuring Hidden Associations for Bias Control in Text-to-Image Compositions
abstract
Text-to-image generative models often exhibit bias related to sensitive attributes. However, current research tends to focus narrowly on single-object prompts with limited contextual diversity. In reality, each object or attribute within a prompt can contribute to bias. For example, the prompt ``an assistant wearing a pink hat'' may reflect female-inclined biases associated with a pink hat. Neglecting joint semantic bindings in prompts leads to significant failures of current debiasing methods. We present a preliminary investigation into how bias manifests under semantic binding, where contextual associations between objects and attributes affect generative outcomes. We demonstrate that the underlying bias distribution can be amplified based on these associations. Therefore, we introduce a bias adherence score that quantifies how specific object-attribute bindings activate bias. To delve deeper, we develop a training-free context-bias control framework to explore how token decoupling can facilitate the debiasing of semantic bindings. This framework achieves over 10% debiasing improvement in compositional generation tasks. Our analysis of bias scores across various attribute-object bindings and token decorrelation highlights a fundamental challenge: reducing bias without disrupting essential semantic relationships. These findings expose critical limitations in current debiasing approaches when applied to semantically bound contexts, underscoring the need to reassess prevailing bias mitigation strategies.
Jeng-Lin Li, Ming-Ching Chang, Wei-Chao Chen
AAAI3
2026 PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
Jeng-Lin Li, Po-Hsuan Huang, Ming-Ching Chang, Wei-Chao Chen
WACV5
2025 Latent Orthogonal Perturbation: A Data Poisoning Approach to Prevent Unauthorized LoRA Fine-Tuning
abstract
Generative AI (GenAI) enables creators to produce high-quality images without advanced artistic skills, but also raises concerns over the unauthorized replication of distinctive styles. With the rise of fine-tuning techniques like LoRA, mimicking specific artists’ styles has become increasingly accessible. To address this, we propose Latent Orthogonal Perturbation (LO-Perturbation), a data poisoning method that subtly injects orthogonal noise into the latent space of artworks. These perturbations introduce only subtle visual changes but significantly degrade the model’s ability to learn fine-grained stylistic features during LoRA fine-tuning. We evaluate our method using five standard image quality metrics. Results show that, while Glaze and Nightshade prioritize preserving visual fidelity, LO-Perturbation leads to significantly greater degradation in image quality and alignment under LoRA fine-tuning—achieving a 56.92% lower Q-Align score and a 46.18% higher FID, along with the lowest CLIP-IQA and LIQE scores. Our method directly disrupts the fine-tuning process, a threat that is less emphasized in existing protection techniques, thus offering an additional layer of protection for artists.
Hong Bin Tan, Ching-Yun Fu, Ming-Ching Chang, Wei-Chao Chen
AVSS4
2025 Enhancing LLM Question Answering with RAG through Dense Vector Search and Re-Ranking
abstract
Retrieval-Augmented Generation (RAG) has emerged as a powerful framework for enhancing Large Language Models (LLMs) by incorporating external knowledge through information retrieval (IR) techniques. However, in question-answering tasks, RAG often retrieves documents that are only semantically similar to the query, which may not provide the most relevant information for generating accurate responses. To address this limitation, we propose an improved retrieval pipeline that combines dense vector search with a re-ranking mechanism to more effectively identify and extract highly relevant knowledge from the retrieved content. We evaluated our approach on two Chinese datasets, TTQA and TMMLU+, using 17 different LLMs. Experimental results show that our method improves performance by up to 21.24% over baseline approaches, particularly on two finance-related subsets, after incorporating domain-specific financial regulations to enhance the knowledge base used in the TMMLU+ dataset.
Te-Lun Yang, Jyi-Shane Liu, Yuen-Hsien Tseng, Jyh-Shing Roger Jang, Ming-Ching Chang, Wei-Chao Chen
AVSS6
2025 A Comprehensive Evaluation of Encrypted DNN Inference Methods
abstract
Fully Homomorphic Encryption (FHE) in the realm of deep neural network (DNN) encrypted inference represents a pivotal advancement in privacy-preserving machine learning. This technology allows users to securely access DNN inference services hosted on remote servers without compromising their personal privacy. Given its wide range of potential applications, FHE has garnered significant research attention. Despite its rapid progress, FHE still faces considerable challenges, particularly the high computational resource demands. Moreover, varying configurations of FHE schemes can lead to notable differences in performance, whether in terms of efficiency or inference accuracy, making it difficult to strike an optimal balance tailored to specific application requirements.To tackle this challenge, we introduce a new approach that simulates FHE-induced errors to assess the impact of different FHE architectures and parameter configurations on encrypted inference during the model testing phase. Our simulation framework enables efficient approximation of homomorphic computation outcomes on a DNN model, specifically for third-generation FHEs, without the need for executing the complete homomorphic process. This method significantly streamlines the research process for optimizing parameters based on specific application needs. We validate the effectiveness of our approach through extensive performance benchmarking across a variety of experimental settings.
Yu-Te Ku, Ming-Chien Ho, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, Wei-Chao Chen
ISCAS8
2025 Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
abstract
Recent advancements in large vision-language models (LVLM) have significantly enhanced their ability to comprehend visual inputs alongside natural language. How-ever, a major challenge in their real-world application is hallucination, where LVLMs generate non-existent visual elements, eroding user trust. The underlying mechanism driving this multimodal hallucination is poorly understood. Minimal research has illuminated whether contexts such as sky, tree, or grass field involve the LVLM in hallucinating a frisbee. We hypothesize that hiddenfactors, such as objects, contexts, and semantic foreground-background structures, induce hallucination. This study proposes a novel causal approach: a hallucination probing system to identify these hidden factors. By analyzing the causality between images, text prompts, and network saliency, we systematically ex-plore interventions to block these factors. Our experimen-tal findings show that a straightforward technique based on our analysis can significantly reduce hallucinations. Additionally, our analyses indicate the potential to edit network internals to minimize hallucinated outputs.
Po-Hsuan Huang, Jeng-Lin Li, Chin-Po Chen, Ming-Ching Chang, Wei-Chao Chen
WACV5
2025 Optimizing Encrypted Neural Networks: Model Design, Quantization and Fine-Tuning Using FHEW/TFHE
abstract
Third-generation Fully Homomorphic Encryption (FHE), particularly the FHEW/TFHE schemes, is recognized for its balanced security requirements, small parameters, and low memory usage, though the current methods in the scenarios of Deep Neural Network (DNN) inference still have high computational costs, limiting the practical applicability. This work demonstrates how to improve practicality of the third-generation technologies for DNN tasks while preserving its key advantages. Our work focuses on two main contributions. First, we developed a computational architecture called FHE-Neuron, which reconfigures the parameters and bootstrapping structure of traditional FHEW/TFHE Boolean operations. This architecture significantly reducing the cost of encrypted DNN inference by dynamically switching the precision of encrypted data during computation—using high precision for cost-effective linear operations and low precision for computationally expensive nonlinear operations. Second, we introduced an FHE-aware Quantization and Fine-tuning framework that optimizes model parameters to align with FHE-Neuron’s constraints, ensuring high accuracy in encrypted inference. We validate our approach on various neural network models across several computing platforms. In our experiments, our method achieves one-image inference time on average 4.5 milliseconds for MNIST and 17 milliseconds for Fashion MNIST, achieving accuracy rates of 96.52% and 88.57% respectively. For the CIFAR-10 dataset, our system completes one image inference in 30 seconds with a 90.5% accuracy rate.
Yu-Te Ku, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, I-Ping Tu, Wei-Chao Chen
Proc. Priv. Enhancing Technol.7
2024 Invited Paper: Efficient Design of FHEW/TFHE Bootstrapping Implementation with Scalable Parameters
abstract
Fully Homomorphic Encryption (FHE) is vital for computing over encrypted data, thereby enabling numerous privacy-preserving applications. This work focuses on the third generation FHE schemes (e.g., FHEW and TFHE), known for their fast bootstrapping, small FHE parameters, and robust security built on milder assumptions.
Ming-Chien Ho, Yu-Te Ku, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, Wei-Chao Chen
ICCAD8
2024 Learning With Instance-Dependent Noisy Labels By Anchor Hallucination And Hard Sample Label Correction
abstract
Learning from noisy-labeled data is crucial for real-world applications. Traditional Noisy-Label Learning (NLL) methods categorize training data into clean and noisy sets based on the loss distribution of training samples. However, they often neglect that clean samples, especially those with intricate visual patterns, may also yield substantial losses. This oversight is particularly significant in datasets with Instance-Dependent Noise (IDN), where mislabeling probabilities correlate with visual appearance. Our approach explicitly distinguishes between clean $v s$. noisy and easy $v s$. hard samples. We identify training samples with small losses, assuming they have simple patterns and correct labels. Utilizing these easy samples, we hallucinate multiple anchors to select hard samples for label correction. Corrected hard samples, along with the easy samples, are used as labeled data in subsequent semi-supervised training. Experiments on synthetic and real-world IDN datasets demonstrate the superior performance of our method over other state-of-the-art NLL methods.
Po-Hsuan Huang, Chia-Ching Lin, Chih-Fan Hsu, Ming-Ching Chang, Wei-Chao Chen
ICIP5
2024 Expert Composer Policy: Scalable Skill Repertoire for Quadruped Robots
abstract
We propose the expert composer policy, a framework to reliably expand the skill repertoire of quadruped agents. The composer policy links pair of experts via transitions to a sampled target state, allowing experts to be composed sequentially. Each expert specializes in a single skill, such as a locomotion gait or a jumping motion. Instead of a hierarchical or mixture-of-experts architecture, we train a single composer policy in an independent process that is not conditioned on the other expert policies. By reusing the same composer policy, our approach enables adding new experts without affecting existing ones, enabling incremental repertoire expansion and preserving original motion quality. We measured the transition success rate of 72 transition pairs and achieved an average success rate of 99.99%, which is over 10% higher than the baseline random approach, and outperforms other state-of-the-art methods. Using domain randomization during training we ensure a successful transfer to the real world, where we achieve an average transition success rate of 97.22% (N=360) in our experiments.
Guilherme Christmann, Ying-Sheng Luo, Wei-Chao Chen
ICRA3
2024 Benchmarking Smoothness and Reducing High-Frequency Oscillations in Continuous Control Policies
abstract
Reinforcement learning (RL) policies are prone to high-frequency oscillations, especially undesirable when deploying to hardware in the real-world. In this paper, we identify, categorize, and compare methods from the literature that aim to mitigate high-frequency oscillations in deep RL. We define two broad classes: loss regularization and architectural methods. At their core, these methods incentivize learning a smooth mapping, such that nearby states in the input space produce nearby actions in the output space. We present benchmarks in terms of policy performance and control smoothness on traditional RL environments from the Gymnasium and a complex manipulation task, as well as three robotics locomotion tasks that include deployment and evaluation with real-world hardware. Finally, we also propose hybrid methods that combine elements from both loss regularization and architectural methods. We find that the best-performing hybrid outperforms other methods, and improves control smoothness by 26.8% over the baseline, with a worst-case performance degradation of just 2.8%.
Guilherme Christmann, Ying-Sheng Luo, Hanjaya Mandala, Wei-Chao Chen
IROS4
2023 Expanding Versatility of Agile Locomotion through Policy Transitions Using Latent State Representation
abstract
This paper proposes the transition-net, a robust transition strategy that expands the versatility of robot locomotion in the real-world setting. To this end, we start by distributing the complexity of different gaits into dedicated locomotion policies applicable to real-world robots. Next, we expand the versatility of the robot by unifying the policies with robust transitions into a single coherent meta-controller by examining the latent state representations. Our approach enables the robot to iteratively expand its skill repertoire and robustly transition between any policy pair in a library. In our framework, adding new skills does not introduce any process that alters the previously learned skills. Moreover, training of a locomotion policy takes less than an hour with a single consumer GPU. Our approach is effective in the real-world and achieves a 19% higher average success rate for the most challenging transition pairs in our experiments compared to existing approaches.
Guilherme Christmann, Ying-Sheng Luo, Jonathan Hans Soeseno, Wei-Chao Chen
ICRA4
2021 A Dense Tensor Accelerator with Data Exchange Mesh for DNN and Vision Workloads
abstract
We propose a dense tensor accelerator called VectorMesh, a scalable, memory-efficient architecture that can support a wide variety of DNN and computer vision workloads. Its building block is a tile execution unit (TEU), which includes dozens of processing elements (PEs) and SRAM buffers connected through a butterfly network. A mesh of FIFOs between the TEUs facilitates data exchange between tiles and promote local data to global visibility. Our design performs better according to the roofline model for CNN, GEMM, and spatial matching algorithms compared to state-of-the-art architectures. It can reduce global buffer and DRAM fetches by 2-22 times and up to 5 times, respectively.
Wei-Chao Chen, Chia-Lin Yang, Shao-Yi Chien
ISCAS2
2021 TrustMAE: A Noise-Resilient Defect Classification Framework using Memory-Augmented Auto-Encoders with Trust Regions
abstract
In this paper, we propose a framework called Trust-MAE to address the problem of product defect classification. Instead of relying on defective images that are difficult to collect and laborious to label, our framework can accept datasets with unlabeled images. Moreover, unlike most anomaly detection methods, our approach is robust against noises, or defective images, in the training dataset. Our framework uses a memory-augmented auto-encoder with a sparse memory addressing scheme to avoid over-generalizing the auto-encoder, and a novel trust-region memory updating scheme to keep the noises away from the memory slots. The result is a framework that can reconstruct defect-free images and identify the defective regions using a perceptual distance network. When compared against various state-of-the-art baselines, our approach performs competitively under noise-free MVTec datasets. More importantly, it remains effective at a noise level up to 40% while significantly outperforming other baselines.
Daniel Stanley Tan, Trista Pei-Chun Chen, Wei-Chao Chen
WACV4
2020 CARL: controllable agent with reinforcement learning for quadruped locomotion
abstract
Motion synthesis in a dynamic environment has been a long-standing problem for character animation. Methods using motion capture data tend to scale poorly in complex environments because of their larger capturing and labeling requirement. Physics-based controllers are effective in this regard, albeit less controllable. In this paper, we present CARL, a quadruped agent that can be controlled with high-level directives and react naturally to dynamic environments. Starting with an agent that can imitate individual animation clips, we use Generative Adversarial Networks to adapt high-level controls, such as speed and heading, to action distributions that correspond to the original animations. Further fine-tuning through the deep reinforcement learning enables the agent to recover from unseen external perturbations while producing smooth transitions. It then becomes straightforward to create autonomous agents in dynamic environments by adding navigation modules over the entire process. We evaluate our approach by measuring the agent's ability to follow user control and provide a visual analysis of the generated motion to show its effectiveness.
Ying-Sheng Luo, Jonathan Hans Soeseno, Trista Pei-Chun Chen, Wei-Chao Chen
ACM Trans. Graph.4
2020 MERIT: Tensor Transform for Memory-Efficient Vision Processing on Parallel Architectures
abstract
Computationally intensive deep neural networks (DNNs) are well- suited to run on GPUs, but newly developed algorithms usually require the heavily optimized DNN routines to work efficiently, and this problem could be even more difficult for specialized DNN architectures. In this article, we propose a mathematical formulation that can be useful for transferring the algorithm optimization knowledge across computing platforms. We discover that data movement and storage inside parallel processor architectures can be viewed as tensor transforms across memory hierarchies, making it possible to describe many memory optimization techniques mathematically. Such transform, which we call memory-efficient ranged inner-product tensor (MERIT) transform, can be applied to not only DNN tasks but also many traditional machine learning and computer vision computations. Moreover, the tensor transforms can be readily mapped to existing vector processor architectures. In this article, we demonstrate that many popular applications can be converted to a succinct MERIT notation on GPUs, speeding up GPU kernels up to 20 times while using only half as many code tokens. We also use the principle of the proposed transform to design a specialized hardware unit called MERIT-z processor. This processor can be applied to a variety of DNN tasks as well as other computer vision tasks while providing comparable area and power efficiency to dedicated DNN application-specific integrated circuits (ASICs).
Wei-Chao Chen, Shao-Yi Chien
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Unrolled Memory Inner-Products: An Abstract GPU Operator for Efficient Vision-Related Computations
abstract
Recently, convolutional neural networks (CNNs) have achieved great success in fields such as computer vision, natural language processing, and artificial intelligence. Many of these applications utilize parallel processing in GPUs to achieve higher performance. However, it remains a daunting task to optimize for GPUs, and most researchers have to rely on vendor-provided libraries for such purposes. In this paper, we discuss an operator that can be used to succinctly express computational kernels in CNNs and various scientific and vision applications. This operator, called Unrolled-Memory-Inner-Product (UMI), is a computationally-efficient operator with smaller code token requirement. Since a naive UMI implementation would increase memory requirement through input data unrolling, we propose a method to achieve optimal memory fetch performance in modern GPUs. We demonstrate this operator by converting several popular applications into the UMI representation, and achieve 1.3x-26.4x speedup against frameworks such as OpenCV and Caffe.
Wei-Chao Chen, Shao-Yi Chien
ICCV2
2013 Dynamic Media Assemblage
abstract
This paper presents dynamic media assemblage, which is a new presentation and summarization method for images and videos on a 2-D canvas. Instead of using the keyframes of the videos to generate a still image summarization, our method allows the videos to play simultaneously on the canvas while utilizing the limited space efficiently. This technique uses an efficient iterative packing algorithm, and as a result is well suited for interactive manipulations of media files within the assemblages in real time, such as insertion, deletion, and rearrangement. Our method starts by detecting shot boundaries and dividing longer input videos into individual shots. Within each shot, its temporal-spatial salient regions are extracted and used to recover camera motions. These saliency information further defines important regions within individual videos or images, which allows us to preserve visually important regions while packing the media files more efficiently into media assemblages. Our algorithm is iterative and can therefore quickly adjust to new canvas sizes or other user intentions. We demonstrate the effectiveness of our techniques by applying our methods to several applications, including media collection presentation, single video dynamic summary, personal media file browser, and interactive video wall.
Sheng-Jie Luo, Chun-Yu Tsai, Wei-Chao Chen, Bing-Yu Chen 0004
IEEE Trans. Circuits Syst. Video Technol.3
2012 Warping-Based Novel View Synthesis from a Binocular Image for Autostereoscopic Displays
abstract
This paper presents a warping-based method for synthesizing multiple views from a binocular stereoscopic image. Auto stereoscopic displays require multiple views while most stereoscopic cameras can only capture two. Popular novel view synthesis methods, such as depth image based rendering (DIBR), often heavily rely on accurate depth maps, which are still difficult to obtain. The proposed method requires neither depth maps nor user intervention. It extracts dense and reliable features. Feature correspondences guide image warping to synthesize novel views while simultaneously maintaining stereoscopic properties and preserving image structures. Compared to DIBR, the proposed method produces higher-quality multi-view images more efficiently without tedious parameter tuning. The method can be used to convert stereoscopic images taken by binocular cameras into multi-view images ready to be displayed on auto stereoscopic displays.
Yu-Hsiang Huang, Tzu-Kuei Huang, Yan-Hsiang Huang, Wei-Chao Chen, Yung-Yu Chuang
ICME4
2010 TouchTone: Interactive Local Image Adjustment Using Point-and-Swipe
abstract
Abstract Recent proliferation of camera phones, photo sharing and social network services has significantly changed how we process our photos. Instead of going through the traditional download‐edit‐share cycle using desktop editors, an increasing number of photos are taken with camera phones and published through cellular networks. The immediacy of the sharing process means that on‐device image editing, if needed, should be quick and intuitive. However, due to the limited computational resources and vastly different user interaction model on small screens, most traditional local selection methods can not be directly adapted to mobile devices. To address this issue, we present TouchTone, a new method for edge‐aware image adjustment using simple finger gestures. Our method enables users to select regions within the image and adjust their corresponding photographic attributes simultaneously through a simple point‐and‐swipe interaction. To enable fast interaction, we develop a memory‐ and computation‐efficient algorithm which samples a collection of 1D paths from the image, computes the adjustment solution along these paths, and interpolates the solutions to entire image through bilateral filtering. Our system is intuitive to use, and can support several local editing tasks, such as brightness, contrast, and color balance adjustments, within a minute on a mobile device.
Chia-Kai Liang, Wei-Chao Chen, Natasha Gelfand
Comput. Graph. Forum2
2009 SURFTrac: Efficient tracking and continuous object recognition using local feature descriptors
abstract
We present an efficient algorithm for continuous image recognition and feature descriptor tracking in video which operates by reducing the search space of possible interest points inside of the scale space image pyramid. Instead of performing tracking in 2D images, we search and match candidate features in local neighborhoods inside the 3D image pyramid without computing their feature descriptors. The candidates are further validated by fitting to a motion model. The resulting tracked interest points are more repeatable and resilient to noise, and descriptor computation becomes much more efficient because only those areas of the image pyramid that contain features are searched. We demonstrate our method on real-time object recognition and label augmentation running on a mobile device.
Duy-Nguyen Ta, Wei-Chao Chen, Natasha Gelfand, Kari Pulli
CVPR2
2009 Visual summaries of popular landmarks from community photo collections
abstract
We present a novel data-driven algorithm that leverages online image repositories such as Flickr for automatically generating tourist maps. Our hypothesis is that, given a large enough dataset of images with geo-based metadata, clusters of matching images from that dataset tend to provide reliable cues as to what the popular tourist spots may be. Our algorithm takes the geographical area of interest as input and retrieves geotagged photos from online photo collections. By clustering the photos based on their locations and identifying the popular tags for each cluster, our algorithm generates a set of points of interest (POIs) for the area. After retrieving additional photos based on these discovered POI tags, we use image matching to find the most representative landmark view for each POI. Finally, we remove clutter from the representative image and apply tooning to generate a map icon for each landmark.
Wei-Chao Chen, Agathe Battestini, Natasha Gelfand, Vidya Setlur
ACM Multimedia1
2007 Efficient Extraction of Robust Image Features on Mobile Devices
abstract
Recent convergence of imaging sensors and general purpose processors on mobile phones creates an opportunity for a new class of augmented reality applications. Robust image feature extraction is a crucial enabler of this type of systems. In this article, we discuss an efficient mobile phone implementation of a state-of-the-art algorithm for computing robust image features called SURF. We implement several improvements to the basic algorithm that significantly improve its performance and reduce its memory footprint making the use of this algorithm on the mobile phone practical. Our prototype implementation has been applied to several practical applications such as image search, object recognition and augmented reality applications.
Wei-Chao Chen, Yingen Xiong, Jiang Gao, Natasha Gelfand, Ramakrishna Grzeszczuk
ISMAR1
2006 SDG Cut: 3D Reconstruction of Non-lambertian Objects Using Graph Cuts on Surface Distance Grid
abstract
We show that the approaches to 3D reconstruction that use volumetric graph cuts to minimize a cost function over the object surface have two types of biases, the minimal surface bias and the discretization bias. These biases make it difficult to recover surface extrusions and other details, especially when a non-lambertian photo-consistency measure is used. To reduce these biases, we propose a new iterative graph cuts based algorithm that operates on the Surface Distance Grid (SDG), which is a special discretization of the 3Dspace, constructed using a signed distance transform of the current surface estimate. It can be shown that SDG significantly reduces the minimal surface bias, and transforms the discretization bias into a controllable degree of surface smoothness. Experiments on 3D reconstruction of non-lambertian objects confirm the effectiveness of our algorithm over previous methods.
Tian-Li Yu 0002, Narendra Ahuja, Wei-Chao Chen
CVPR (2)3
2006 Sparse Lumigraph Relighting by Illumination and Reflectance Estimation from Multi-View Images
Tian-Li Yu 0002, Narendra Ahuja, Wei-Chao Chen
Rendering Techniques4
2003 Acquisition of large-scale surface light fields
abstract
Abstract: This sketch reports the development of an acquisition methodology to acquire high-resolution surface light fields of an office-size environment. 1
Wei-Chao Chen, Lars S. Nyland, Anselmo Lastra, Henry Fuchs
SIGGRAPH1
2002 Light field mapping: efficient representation and hardware rendering of surface light fields
abstract
A light field parameterized on the surface offers a natural and intuitive description of the view-dependent appearance of scenes with complex reflectance properties. To enable the use of surface light fields in real-time rendering we develop a compact representation suitable for an accelerated graphics pipeline. We propose to approximate the light field data by partitioning it over elementary surface primitives and factorizing each part into a small set of lower-dimensional functions. We show that our representation can be further compressed using standard image compression techniques leading to extremely compact data sets that are up to four orders of magnitude smaller than the input data. Finally, we develop an image-based rendering method, light field mapping, that can visualize surface light fields directly from this compact representation at interactive frame rates on a personal computer. We also implement a new method of approximating the light field data that produces positive only factors allowing for faster rendering using simpler graphics hardware than earlier methods. We demonstrate the results for a variety of non-trivial synthetic scenes and physical objects scanned through 3D photography.
Wei-Chao Chen, Jean-Yves Bouguet, Michael H. Chu, Radek Grzeszczuk
ACM Trans. Graph.1
2000 Toward a compelling sensation of telepresence: demonstrating a portal to a distant (static) office
abstract
In 1998 we introduced the idea for a project we call the Office of the Future. Our long-term vision is to provide a better every-day working environment, with high-fidelity scene reconstruction for life-sized 3D tele-collaboration. In particular, we want a true sense of presence with our remote collaborator and their real surroundings. The challenges related to this vision are enormous and involve many technical tradeoffs. This is true in particular for scene reconstruction. Researchers have been striving to achieve real-time approaches, and while they have made respectable progress, the limitations of conventional technologies relegate them to relatively low resolution in a restricted volume. We present a significant step toward our ultimate goal, via a slightly different path. In lieu of low-fidelity dynamic scene modeling we present an exceedingly high fidelity reconstruction of a real but static office. By assembling the best of available hardware and software technologies in static scene acquisition, modeling algorithms, rendering, tracking and stereo projective display, we are able to demonstrate a portal to a real office, occupied today by a mannequin, and in the future by a real remote collaborator. We now have both a compelling sense of just how good it could be, and a framework into which we will later incorporate dynamic scene modeling, as we continue to head toward our ultimate goal of 3D collaborative telepresence.
Wei-Chao Chen, Herman Towles, Lars S. Nyland, Greg Welch, Henry Fuchs
IEEE Visualization1
1999 Multi-Projector Displays Using Camera-Based Registration
abstract
Conventional projector-based display systems are typically designed around precise and regular configurations of projectors and display surfaces. While this results in rendering simplicity and speed, it also means painstaking construction and ongoing maintenance. In previously published work, we introduced a vision of projector-based displays constructed from a collection of casually-arranged projectors and display surfaces. In this paper, we present flexible yet practical methods for realizing this vision, enabling low-cost mega-pixel display systems with large physical dimensions, higher resolution, or both. The techniques afford new opportunities to build personal 3D visualization systems in offices, conference rooms, theaters, or even your living room. As a demonstration of the simplicity and effectiveness of the methods that we continue to perfect, we show in the included video that a 10-year old child can construct and calibrate a two-camera, two-projector, head-tracked display system, all in about 15 minutes.
Ramesh Raskar, Michael S. Brown, Ruigang Yang, Wei-Chao Chen, Greg Welch, Herman Towles, W. Brent Seales, Henry Fuchs
IEEE Visualization4
1996 Image Shading Taking into Account Relativistic Effects
abstract
This article is concerned with creating more realistic images of 3D scenes which are moving relative to the viewer at such high speeds that the propagation delay of light signals and other relativistic effects can not be neglected. Creating images of 3D scenes in relativistic motion might have important applications to science-fiction films, computer games, and virtual environments. We shall discuss the following problems: (1) how to determine the visual appearance of a rapidly moving object, (2) how to determine the apparent radiance of a scene point on a moving object, (3) how to determine the incident irradiance at a scene point coming from a moving light source, (4) how to determine the color of a rapidly moving object, and (5) how to generate shadows when there are relative motions between the viewer, the scenes, and the light sources. Detailed examples are also given to show the result of shading with the relativistic effects taken into account.
Meng-chou Chang, Feipei Lai, Wei-Chao Chen
ACM Trans. Graph.3