Yee-Hong Yang

dblp:01/184 · DBLP profile ↗
← Back
130ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0002-7194-3327ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 80 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 14 since 2021Human-computer interaction and ubiquitous computing · 8Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2026 SciceVPR: Stable cross-image correlation enhanced model for visual place recognition
Shanshan Wan, Yingmei Wei, Lai Kang, Tianrui Shen, Haixuan Wang, Yee-Hong Yang
Neurocomputing6
2026 V-Sparse: From temporal-spatial visual semantic compression to coarse-to-fine interaction for text-video retrieval
Shibai Yin, Jun Wang 0089, Xingyang Wang, Yubing Shen, Yee-Hong Yang
Neural Networks8
2026 DGA-GCN: Dynamic Global Adaptive Graph Convolutional Networks for skeleton-based action recognition
Zhao Pei, Yanni Xue, Zhichao Ren, Chengcai Leng, Yee-Hong Yang
Pattern Recognit.6
2025 DUQ: Dual Uncertainty Quantification for Text-Video Retrieval
abstract
Text-video retrieval establishes accurate similarity relationships between text and video through feature enhancement and granularity alignment. However, relying solely on similarity to associate intra-pair features and distinguish inter-pair features is insufficient, \textit{e.g.}, when querying a multi-scene video with sparse text or selecting the most relevant video from many similar candidates. In this paper, we propose a novel Dual Uncertainty Quantification (DUQ) model that separately handles uncertainties in intra-pair interaction and inter-pair exclusion. Specifically, to enhance intra-pair interaction, we propose an intra-pair similarity uncertainty module to provide similarity-based trustworthy predictions and explicitly model this uncertainty. To increase inter-pair exclusion, we propose an inter-pair distance uncertainty module to construct a distance-based diversity probability embeding, thereby widening the gap between similar features. The two components work synergistically, jointly improving the calculation of similarity between features. We evaluate our model on six benchmark datasets: MSRVTT (51.2%), DiDeMo, MSVD, LSMDC, Charades, and VATEX, achieving state-of-the-art retrieval performance.
Shibai Yin, Xingyang Wang, Yee-Hong Yang
IJCAI6
2025 Dual-Branch Wavelet Diffusion models with Dual-Prior Refinement for Underwater Image Enhancement
Yiwei Shi, Shibai Yin, Yanfang Fu, Yee-Hong Yang
J. Vis. Commun. Image Represent.6
2025 FishDetectLLM: Multimodal instruction tuning with large language models for fish detection
abstract
Aquatic species play crucial roles in global ecosystems but are increasingly threatened by factors such as overfishing, coastal development and climate change . Existing deep learning methods address these challenges by employing powerful networks and large-scale, diverse datasets, separately tackling species recognition and trait identification during ongoing monitoring. However, they often exhibit limited generalization ability. Inspired by the human ability to quickly identify fish species and their locations with just a glance at an underwater image or scene, we introduce FishDetectLLM—a framework built on the lightweight TinyLLaVA architecture. FishDetectLLM utilizes the powerful reasoning capabilities and vast world knowledge of large language models (LLMs) to address the fish detection problem, providing both fish classification results and predicted bounding boxes for fish. Specifically, we create instruction dialogues for fish detection that connect fish taxonomy with classification descriptions and map location descriptions to the corresponding coordinates of bounding box in the input images from the recently released large-scale FishNet dataset. Then, we pretrain and fine-tune FishDetectLLM to achieve fish detection using the created dataset, leveraging the principle of augmenting human knowledge. Our results show that FishDetectLLM significantly outperforms existing multimodal LLMs and task-specific methods. Unlike conventional detection architectures that struggle to generalize beyond the training data, FishDetectLLM exhibits strong generalization capabilities, achieving robust performance on unseen data. This innovation paves the way for future applications of MLLMs in full research and offers valuable tools for the conservation of fish biodiversity.
Shibai Yin, Xingyang Wang, Yee-Hong Yang
Knowl. Based Syst.5
2025 When Aware Haze Density Meets Diffusion Model for Synthetic-to-Real Dehazing
abstract
Image dehazing is an important preliminary step for downstream vision tasks. Existing deep learning-based methods have limited generalization capabilities for real hazy images because they are trained on synthetic data and exhibit high domain-specific properties. This work proposes a new Diffusion Model for Synthetic-to-Real dehazing (DMSR) based on the haze-aware density. DMSR mainly comprises of a physics-based dehazing model and a Conditional Denoising Diffusion Model (CDDM)-based model. The coarse transmission map and coarse dehazing result estimated by the physics-based dehazing model serve as conditions for the subsequent CDDM-based model. In this process, the CDDM-based dehazing model progressively refines the coarse transmission map while generating the dehazing result, enabling the model to remove haze with accurate haze density information. Next, we propose a haze density-aware resampling strategy that incorporates the coarse dehazed result into the resampling process using the transmission map, thereby fully leveraging the diffusion model for heavy haze removal. Moreover, a new synthetic-to-real training strategy with the prior-based loss function and the memory loss function is applied to DMSR for improving generalization capabilities and narrowing the gap between the synthetic and real domains with low computational cost. Extensive experiments on various real datasets demonstrate the effectiveness and superiority of the proposed DMSR over state-of-the-art methods.
Shibai Yin, Yiwei Shi, Yibin Wang 0001, Yee-Hong Yang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Spherical Pseudo-Cylindrical Representation for Omnidirectional Image Super-resolution
abstract
Omnidirectional images have attracted significant attention in recent years due to the rapid development of virtual reality technologies. Equirectangular projection (ERP), a naive form to store and transfer omnidirectional images, however, is challenging for existing two-dimensional (2D) image super-resolution (SR) methods due to its inhomogeneous distributed sampling density and distortion across latitude. In this paper, we make one of the first attempts to design a spherical pseudo-cylindrical representation, which not only allows pixels at different latitudes to adaptively adopt the best distinct sampling density but also is model-agnostic to most off-the-shelf SR methods, enhancing their performances. Specifically, we start by upsampling each latitude of the input ERP image and design a computationally tractable optimization algorithm to adaptively obtain a (sub)-optimal sampling density for each latitude of the ERP image. Addressing the distortion of ERP, we introduce a new viewport-based training loss based on the original 3D sphere format of the omnidirectional image, which inherently lacks distortion. Finally, we present a simple yet effective recursive progressive omnidirectional SR network to showcase the feasibility of our idea. The experimental results on public datasets demonstrate the effectiveness of the proposed method as well as the consistently superior performance of our method over most state-of-the-art methods both quantitatively and qualitatively.
Dongwei Ren, Haiyong Zheng, Junyu Dong, Yee-Hong Yang
AAAI7
2024 TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling
Dong Huo, Zixin Guo, Xinxin Zuo, Zhihao Shi, Juwei Lu, Peng Dai 0002, Songcen Xu, Li Cheng 0001, Yee-Hong Yang
ECCV (38)9
2024 Visual Attention and ODE-inspired Fusion Network for image dehazing
Shibai Yin, Ruyuan Lu, Zhen Deng, Yee-Hong Yang
Eng. Appl. Artif. Intell.5
2024 A multi-level wavelet-based underwater image enhancement network with color compensation prior
Yibin Wang 0001, Shuhao Hu, Shibai Yin, Zhen Deng, Yee-Hong Yang
Expert Syst. Appl.5
2024 Convolution-transformer blend pyramid network for underwater image enhancement
Lunpeng Ma, Dongyang Hong, Shibai Yin, Wanqiu Deng, Yang Yang 0229, Yee-Hong Yang
J. Vis. Commun. Image Represent.6
2024 Autofocusing for Synthetic Aperture Imaging Based on Pedestrian Trajectory Prediction
abstract
Occlusions and complex backgrounds are common factors that hinder many computer vision applications. In a street scene, the challenge of accurately predicting pedestrian trajectories comes from the complexity of human behavior and the diversity of the external environment. It is difficult, if not impossible, to extract relevant information to accurately predict pedestrian trajectories in dynamic scenes. Synthetic aperture imaging (SAI) uses an array of cameras to mimic a camera with a large virtual convex lens by projecting images of a scene from different views onto a virtual focal plane. It is commonly used to reconstruct occluded objects, and in a street scene, can provide observation of pedestrians occluded by other objects and pedestrians. In this paper, we propose a joint prediction method based on autofocusing of SAI to predict pedestrian trajectories in dynamic scenes. The main contributions of this paper include: 1) The task of pedestrian trajectory prediction in dynamic scenarios is redefined as pedestrian trajectory prediction and SAI autofocusing from a practical but more challenging perspective. 2) The proposed method is based on an existing SAI-based method to extract information in heavily occluded views, which can obtain more accurate results but with less computational cost and without using other sensors such as LiDAR or depth cameras. 3) A new pedestrian trajectory prediction model, an attention-based trajectory prediction variational autoencoder (ATP-VAE), is proposed to extract complex human behavior and social interactions in dynamic scenes through a new Intention Attention Unit. The experimental results on multiple public datasets show that the proposed method achieves state-of-the-art results in the first-person perspective and in aerial view.
Zhao Pei, Jianing Wang 0003, Yee-Hong Yang
IEEE Trans. Circuits Syst. Video Technol.6
2024 MCCG: A ConvNeXt-Based Multiple-Classifier Method for Cross-View Geo-Localization
abstract
The key to crossview geolocalization is to match images of the same target from different viewpoints, e.g., images from drones and satellites. It is a challenging problem due to the changing appearance of objects from variable viewpoints. Most existing methods focus mainly on extracting global features or on segmenting feature maps, causing the loss of information contained in the images. To address the above issues, we propose a new ConvNeXt-based method called MCCG, which stands for Multiple Classifier for Cross-view Geolocalization. The proposed method captures rich discriminative information by cross-dimension interaction and acquires multiple feature representations, realizing a comprehensive feature representation. Additionally, the robustness of the model is improved crediting the multiple feature representations exploiting more contextual information despite position shifting or scale variations. Extensive experiments on the widely used public benchmarks University-1652 and SUES-200 demonstrate that the proposed method achieves state-of-the-art performance in both drone-view target localization and drone navigation applications by over 3% compared to existing methods. Our code and model are available athttps://github.com/mode-str/crossview.
Tianrui Shen, Yingmei Wei, Lai Kang, Shanshan Wan, Yee-Hong Yang
IEEE Trans. Circuits Syst. Video Technol.5
2024 Learning to Recover Spectral Reflectance From RGB Images
abstract
This paper tackles spectral reflectance recovery (SRR) from RGB images. Since capturing ground-truth spectral reflectance and camera spectral sensitivity are challenging and costly, most existing approaches are trained on synthetic images and utilize the same parameters for all unseen testing images, which are suboptimal especially when the trained models are tested on real images because they never exploit the internal information of the testing images. To address this issue, we adopt a self-supervised meta-auxiliary learning (MAXL) strategy that fine-tunes the well-trained network parameters with each testing image to combine external with internal information. To the best of our knowledge, this is the first work that successfully adapts the MAXL strategy to this problem. Instead of relying on naive end-to-end training, we also propose a novel architecture that integrates the physical relationship between the spectral reflectance and the corresponding RGB images into the network based on our mathematical analysis. Besides, since the spectral reflectance of a scene is independent to its illumination while the corresponding RGB images are not, we recover the spectral reflectance of a scene from its RGB images captured under multiple illuminations to further reduce the unknown. Qualitative and quantitative evaluations demonstrate the effectiveness of our proposed network and of the MAXL. Our code and data are available at https://github.com/Dong-Huo/SRR-MAXL.
Dong Huo, Jian Wang 0100, Yiming Qian, Yee-Hong Yang
IEEE Trans. Image Process.4
2024 PU-Ray: Domain-Independent Point Cloud Upsampling via Ray Marching on Neural Implicit Surface
abstract
While recent advancements in deep-learning point cloud upsampling methods have improved the input to intelligent transportation systems, they still suffer from issues of domain dependency between synthetic and real-scanned point clouds. This paper addresses the above issues by proposing a new ray-based upsampling approach with an arbitrary rate, where a depth prediction is made for each query ray and its corresponding patch. Our novel method simulates the sphere-tracing ray marching algorithm on the neural implicit surface defined with an unsigned distance function (UDF) to achieve more precise and stable ray-depth predictions by training a point-transformer-based network. The rule-based mid-point query sampling method generates more evenly distributed points without requiring an end-to-end model trained using a nearest-neighbor-based reconstruction loss function, which may bias towards the training dataset. Self-supervised learning becomes possible with accurate ground truths within the input point cloud. The results demonstrate the method’s versatility across domains and training scenarios with limited computational resources and training data. Comprehensive analyses of synthetic and real-scanned applications provide empirical evidence for the significance of the upsampling task across the computer vision and graphics domains to real-world applications of ITS.
Sangwon Lim, Karim El-Basyouny, Yee-Hong Yang
IEEE Trans. Intell. Transp. Syst.3
2023 Adams-based hierarchical features fusion network for image dehazing
Shibai Yin, Shuhao Hu, Yibin Wang 0001, Weixing Wang 0001, Yee-Hong Yang
Neural Networks5
2023 Blind Image Deconvolution Using Variational Deep Image Prior
abstract
Conventional deconvolution methods utilize hand-crafted image priors to constrain the optimization. While deep-learning-based methods have simplified the optimization by end-to-end training, they fail to generalize well to blurs unseen in the training dataset. Thus, training image-specific models is important for higher generalization. Deep image prior (DIP) provides an approach to optimize the weights of a randomly initialized network with a single degraded image by maximum a posteriori (MAP), which shows that the architecture of a network can serve as the hand-crafted image prior. Unlike conventional hand-crafted image priors, which are obtained through statistical methods, finding a suitable network architecture is challenging due to the unclear relationship between images and their corresponding architectures. As a result, the network architecture cannot provide enough constraint for the latent sharp image. This paper proposes a new variational deep image prior (VDIP) for blind image deconvolution, which exploits additive hand-crafted image priors on latent sharp images and approximates a distribution for each pixel to avoid suboptimal solutions. Our mathematical analysis shows that the proposed method can better constrain the optimization. The experimental results further demonstrate that the generated images have better quality than that of the original DIP on benchmark datasets.
Dong Huo, Abbas Masoumzadeh, Rafsanjany Kushol, Yee-Hong Yang
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 HIPA: Hierarchical Patch Transformer for Single Image Super Resolution
abstract
Transformer-based architectures start to emerge in single image super resolution (SISR) and have achieved promising performance. However, most existing vision Transformer-based SISR methods still have two shortcomings: (1) they divide images into the same number of patches with a fixed size, which may not be optimal for restoring patches with different levels of texture richness; and (2) their position encodings treat all input tokens equally and hence, neglect the dependencies among them. This paper presents a HIPA, which stands for a novel Transformer architecture that progressively recovers the high resolution image using a hierarchical patch partition. Specifically, we build a cascaded model that processes an input image in multiple stages, where we start with tokens with small patch sizes and gradually merge them to form the full resolution. Such a hierarchical patch mechanism not only explicitly enables feature aggregation at multiple resolutions but also adaptively learns patch-aware features for different image regions, e.g., using a smaller patch for areas with fine details and a larger patch for textureless regions. Meanwhile, a new attention-based position encoding scheme for Transformer is proposed to let the network focus on which tokens should be paid more attention by assigning different weights to different tokens, which is the first time to our best knowledge. Furthermore, we also propose a multi-receptive field attention module to enlarge the convolution receptive field from different branches. The experimental results on several public datasets demonstrate the superior performance of the proposed HIPA over previous methods quantitatively and qualitatively. We will share our code and models when the paper is accepted.
Yiming Qian, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.5
2023 Glass Segmentation With RGB-Thermal Image Pairs
abstract
This paper proposes a new glass segmentation method utilizing paired RGB and thermal images. Due to the large difference between the transmission property of visible light and that of the thermal energy through the glass where most glass is transparent to the visible light but opaque to thermal energy, glass regions of a scene are made more distinguishable with a pair of RGB and thermal images than solely with an RGB image. To exploit such a unique property, we propose a neural network architecture that effectively combines an RGB-thermal image pair with a new multi-modal fusion module based on attention, and integrate CNN and transformer to extract local features and non-local dependencies, respectively. As well, we have collected a new dataset containing 5551 RGB-thermal image pairs with ground-truth segmentation annotations. The qualitative and quantitative evaluations demonstrate the effectiveness of the proposed approach on fusing RGB and thermal data for glass segmentation. Our code and data are available at https://github.com/Dong-Huo/RGB-T-Glass-Segmentation.
Dong Huo, Jian Wang 0100, Yiming Qian, Yee-Hong Yang
IEEE Trans. Image Process.4
2022 Degradation-aware and color-corrected network for underwater image enhancement
Shibai Yin, Shuhao Hu, Yibin Wang 0001, Weixing Wang 0001, Yee-Hong Yang
Knowl. Based Syst.6
2022 Online learnable keyframe extraction in videos and its application with semantic word vector in action recognition
G. M. Mashrur E Elahi, Yee-Hong Yang
Pattern Recognit.2
2022 Online temporal classification of human action using action inference graph
G. M. Mashrur E Elahi, Yee-Hong Yang
Pattern Recognit.2
2022 Multi-scale attention-based pseudo-3D convolution neural network for Alzheimer's disease diagnosis using structural MRI
Zhao Pei, Zhiyang Wan, Yanning Zhang 0001, Miao Wang 0008, Chengcai Leng, Yee-Hong Yang
Pattern Recognit.6
2022 TDPN: Texture and Detail-Preserving Network for Single Image Super-Resolution
abstract
Single image super-resolution (SISR) using deep convolutional neural networks (CNNs) achieves the state-of-the-art performance. Most existing SISR models mainly focus on pursuing high peak signal-to-noise ratio (PSNR) and neglect textures and details. As a result, the recovered images are often perceptually unpleasant. To address this issue, in this paper, we propose a texture and detail-preserving network (TDPN), which focuses not only on local region feature recovery but also on preserving textures and details. Specifically, the high-resolution image is recovered from its corresponding low-resolution input in two branches. First, a multi-reception field based branch is designed to let the network fully learn local region features by adaptively selecting local region features in different reception fields. Then, a texture and detail-learning branch supervised by the textures and details decomposed from the ground-truth high resolution image is proposed to provide additional textures and details for the super-resolution process to improve the perceptual quality. Finally, we introduce a gradient loss into the SISR field and define a novel hybrid loss to strengthen boundary information recovery and to avoid overly smooth boundary in the final recovered high-resolution image caused by using only the MAE loss. More importantly, the proposed method is model-agnostic, which can be applied to most off-the-shelf SISR networks. The experimental results on public datasets demonstrate the superiority of our TDPN on most state-of-the-art SISR methods in PSNR, SSIM and perceptual quality. We will share our code on https://github.com/tocaiqing/TDPN.
Jinxing Li 0003, Huafeng Li 0001, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.4
2022 A Novel Hybrid Level Set Model for Non-Rigid Object Contour Tracking
abstract
Most existing trackers use bounding boxes for object tracking. However, the background contained in the bounding box inevitably decreases the accuracy of the target model, which affects the performance of the tracker and is particularly pronounced for non-rigid objects. To address the above issue, this paper proposes a novel hybrid level set model, which can robustly address the issue of topology changing, occlusions and abrupt motion in non-rigid object tracking by accurately tracking the object contour. In particular, an appearance model is first obtained by repeatedly training and relabeling the initial labeled frame using competing one-class SVMs. Then, by integrating the trained appearance model, an edge detector and image spatial information into the level set model, a new hybrid level set model is presented, which accurately locates the object contour and feeds back to the competing one-class SVMs to update the appearance model of the next frame. In addition, a motion model is defined to predict the accurate location of the object when occlusion and abrupt motion occur in the next frame. Finally, the experimental results on state-of-the-art benchmarks demonstrate the feasibility and effectiveness of the proposed model and the superiority of the proposed method over existing trackers in terms of accuracy and robustness.
Yiming Qian, Sanping Zhou, Jinjun Wang, Yee-Hong Yang
IEEE Trans. Image Process.6
2022 AVLSM: Adaptive Variational Level Set Model for Image Segmentation in the Presence of Severe Intensity Inhomogeneity and High Noise
abstract
Intensity inhomogeneity and noise are two common issues in images but inevitably lead to significant challenges for image segmentation and is particularly pronounced when the two issues simultaneously appear in one image. As a result, most existing level set models yield poor performance when applied to this images. To this end, this paper proposes a novel hybrid level set model, named adaptive variational level set model (AVLSM) by integrating an adaptive scale bias field correction term and a denoising term into one level set framework, which can simultaneously correct the severe inhomogeneous intensity and denoise in segmentation. Specifically, an adaptive scale bias field correction term is first defined to correct the severe inhomogeneous intensity by adaptively adjusting the scale according to the degree of intensity inhomogeneity while segmentation. More importantly, the proposed adaptive scale truncation function in the term is model-agnostic, which can be applied to most off-the-shelf models and improves their performance for image segmentation with severe intensity inhomogeneity. Then, a denoising energy term is constructed based on the variational model, which can remove not only common additive noise but also multiplicative noise often occurred in medical image during segmentation. Finally, by integrating the two proposed energy terms into a variational level set framework, the AVLSM is proposed. The experimental results on synthetic and real images demonstrate the superiority of AVLSM over most state-of-the-art level set models in terms of accuracy, robustness and running time.
Yiming Qian, Sanping Zhou, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.5
2021 Attentive U-recurrent encoder-decoder network for image dehazing
Shibai Yin, Yibin Wang 0001, Yee-Hong Yang
Neurocomputing3
2021 All-in-focus synthetic aperture imaging using generative adversarial network-based semantic inpainting
Zhao Pei, Yanning Zhang 0001, Miao Ma, Yee-Hong Yang
Pattern Recognit.5
2021 Visual Attention Dehazing Network with Multi-level Features Refinement and Fusion
Shibai Yin, Yibin Wang 0001, Yee-Hong Yang
Pattern Recognit.4
2020 A novel image-dehazing network with a parallel attention block
Shibai Yin, Yibin Wang 0001, Yee-Hong Yang
Pattern Recognit.3
2020 Automated Colorization of a Grayscale Image With Seed Points Propagation
abstract
In this paper, we propose a fully automatic image colorization method for grayscale images using neural network and optimization. For a determined training set including the gray images and its corresponding color images, our method segments grayscale images into superpixels and then extracts features of particular points of interest in each superpixel. The obtained features and their RGB values are given as input for, the training colorization neural network of each pixel. To achieve a better image colorization effect in shorter running time, our method further propagates the resulting color points to neighboring pixels for improved colorization results. In the propagation of color, we present a cost function to formalize the premise that neighboring pixels should have the maximum positive similarity of intensities and colors; we then propose our solution to solving the optimization problem. At last, a guided image filter is employed to refine the colorized image. Experiments on a wide variety of images show that the proposed algorithms can achieve superior performance over the state-of-the-art algorithms.
Shaohua Wan 0001, Yu Xia 0033, Lianyong Qi, Yee-Hong Yang, Mohammed Atiquzzaman
IEEE Trans. Multim.4
2019 Saliency-guided level set model for automatic object segmentation
Yiming Qian, Sanping Zhou, Xiaojun Duan, Yee-Hong Yang
Pattern Recognit.6
2019 Human trajectory prediction in crowded scene using social-affinity Long Short-Term Memory
Zhao Pei, Xiaoning Qi, Yanning Zhang 0001, Miao Ma, Yee-Hong Yang
Pattern Recognit.5
2018 Simultaneous 3D Reconstruction for Water Surface and Underwater Scene
Yiming Qian, Yinqiang Zheng, Minglun Gong, Yee-Hong Yang
ECCV (3)4
2018 All-In-Focus Synthetic Aperture Imaging Using Image Matting
abstract
“Seeing through” occluders is one of the most important effects that can be achieved with synthetic aperture imaging. As well, the occlusion problem, a challenging task for many computer vision applications, can be easily handled. Synthetic aperture imaging takes advantage of the property that only objects on the focal plane are sharp. The resulting image that is obtained by averaging images from different views consists of blurry objects away from the focal plane and sharp objects on the focal plane. Removing the blurriness caused by defocusing in synthetic aperture images to achieve an all-in-focus “seeing through” image is a challenging research problem. In this paper, we propose a novel method to improve the image quality of synthetic aperture imaging using image matting via energy minimization by estimating the foreground and the background. In particular, we first estimate the out-of-focus region by focusing on the background objects in each camera view using energy minimization. Next, we utilize a labeling method to create a sharp “see through” synthetic aperture image of the hidden objects. Then, image matting is used to extract the alpha matte of the hidden objects. Finally, by compositing the hidden objects with the estimated background regions, a sharp “see through” synthetic aperture image is created. The experimental results show that the proposed method outperforms the traditional synthetic aperture imaging method [1] as well as its improved versions [2]-[4], which simply dim and blur the area in the image that is out of focus, and a recent all-in-focus method [5]. We show that both the occluded objects and the background can be combined using our method to create a sharp synthetic aperture image.
Zhao Pei, Xida Chen, Yee-Hong Yang
IEEE Trans. Circuits Syst. Video Technol.3
2017 Stereo-Based 3D Reconstruction of Dynamic Fluid Surfaces by Global Optimization
abstract
3D Reconstruction of dynamic fluid surfaces is an open and challenging problem in computer vision. Unlike previous approaches that reconstruct each surface point independently and often return noisy depth maps, we propose a novel global optimization-based approach that recovers both depths and normals of all 3D points simultaneously. Using the traditional refraction stereo setup, we capture the wavy appearance of a pre-generated random pattern, and then estimate the correspondences between the captured images and the known background by tracking the pattern. Assuming that the light is refracted only once through the fluid interface, we minimize an objective function that incorporates both the cross-view normal consistency constraint and the single-view normal consistency constraints. The key idea is that the normals required for light refraction based on Snells law from one view should agree with not only the ones from the second view, but also the ones estimated from local 3D geometry. Moreover, an effective reconstruction error metric is designed for estimating the refractive index of the fluid. We report experimental results on both synthetic and real data demonstrating that the proposed approach is accurate and shows superiority over the conventional stereo-based method.
Yiming Qian, Minglun Gong, Yee-Hong Yang
CVPR3
2017 Two-view underwater 3D reconstruction for cameras with unknown poses under flat refractive interfaces
Lai Kang, Lingda Wu, Yingmei Wei, Songyang Lao, Yee-Hong Yang
Pattern Recognit.5
2017 Synthetic aperture photography using a moving camera-IMU system
Xiaoqiang Zhang 0002, Yanning Zhang 0001, Tao Yang 0006, Yee-Hong Yang
Pattern Recognit.4
2017 A Closed-Form Solution to Single Underwater Camera Calibration Using Triple Wavelength Dispersion and Its Application to Single Camera 3D Reconstruction
abstract
In this paper, we present a new method to estimate the housing parameters of an underwater camera by making full use of triple wavelength dispersion. Our method is based on an important finding that there is a closed-form solution to the distance from the camera center to the refractive interface once the refractive normal is known. The correctness of this finding is mathematically proved in this paper. To the best of our knowledge, such a finding has not been studied or reported, and hence is never proved theoretically. As well, the refractive normal can be estimated by solving a set of linear equations using wavelength dispersion. Our method does not require any calibration target, such as a checkerboard pattern, which may be difficult to manipulate when the camera is deployed deep undersea. Extensive experiments have been carried out which include simulations to verify the correctness and robustness to noise of our method and real experiments. The results of real experiments show that our method works as expected. The accuracy of our results is evaluated against the ground truth in both simulated and real experiments. Finally, we also show how we can apply dispersion to compute the 3D shape of an object using one single camera.
Xida Chen, Yee-Hong Yang
IEEE Trans. Image Process.2
2016 3D Reconstruction of Transparent Objects with Position-Normal Consistency
abstract
Estimating the shape of transparent and refractive objects is one of the few open problems in 3D reconstruction. Under the assumption that the rays refract only twice when traveling through the object, we present the first approach to simultaneously reconstructing the 3D positions and normals of the object's surface at both refraction locations. Our acquisition setup requires only two cameras and one monitor, which serves as the light source. After acquiring the ray-ray correspondences between each camera and the monitor, we solve an optimization function which enforces a new position-normal consistency constraint. That is, the 3D positions of surface points shall agree with the normals required to refract the rays under Snell's law. Experimental results using both synthetic and real data demonstrate the robustness and accuracy of the proposed approach.
Yiming Qian, Minglun Gong, Yee-Hong Yang
CVPR3
2015 Frequency-Based Environment Matting by Compressive Sensing
abstract
Extracting environment mattes using existing approaches often requires either thousands of captured images or a long processing time, or both. In this paper, we propose a novel approach to capturing and extracting the matte of a real scene effectively and efficiently. Grown out of the traditional frequency-based signal analysis, our approach can accurately locate contributing sources. By exploiting the recently developed compressive sensing theory, we simplify the data acquisition process of frequency-based environment matting. Incorporating phase information in a frequency signal into data acquisition further accelerates the matte extraction procedure. Compared with the state-of-the-art method, our approach achieves superior performance on both synthetic and real data, while consuming only a fraction of the processing time.
Yiming Qian, Minglun Gong, Yee-Hong Yang
ICCV3
2015 An easy-to-implement Benchmarking Tool for Mobile Tablet-PC Visual Pose Estimation
abstract
Recently, researchers are interested in mobile device based computer vision applications. An accurate visual pose estimation is a common and important subtask. To quantitatively evaluate the accuracy, ground truth visual pose datasets are needed. However, the lack of inexpensive and easy benchmarking tool for mobile device based visual poses estimation makes it difficult, if not impossible, to quantitatively evaluate the estimated visual poses. In this paper, a novel and easy-to-implement experimental setup is proposed to generate ground truth visual pose data for handheld tablet-PC. The tablet-PC screen is leveraged to display a calibration pattern every time the on-board camera captures an image. The tablet-PC screen image is captured by another camera and is used to estimate the visual pose of the tablet-PC. An experimental environment is setup for parameter calibration and pose accuracy verification. Extensive experimental results with quantitative analysis demonstrate the accuracy and the generality of our tool.
Xiaoqiang Zhang 0002, Yanning Zhang 0001, Tao Yang 0006, Ting Chen 0004, Yee-Hong Yang
MoMM5
2015 Scene adaptive structured light using error detection and correction
Xida Chen, Yee-Hong Yang
Pattern Recognit.2
2014 Robust Edge Aware Descriptor for Image Matching
Rouzbeh Maani, Sanjay Kalra, Yee-Hong Yang
ACCV (1)3
2014 Two-View Camera Housing Parameters Calibration for Multi-layer Flat Refractive Interface
abstract
In this paper, we present a novel refractive calibration method for an underwater stereo camera system where both cameras are looking through multiple parallel flat refractive interfaces. At the heart of our method is an important finding that the thickness of the interface can be estimated from a set of pixel correspondences in the stereo images when the refractive axis is given. To our best knowledge, such a finding has not been studied or reported. Moreover, by exploring the search space for the refractive axis and using reprojection error as a measure, both the refractive axis and the thickness of the interface can be recovered simultaneously. Our method does not require any calibration target such as a checkerboard pattern which may be difficult to manipulate when the cameras are deployed deep undersea. The implementation of our method is simple. In particular, it only requires solving a set of linear equations of the form Ax = b and applies sparse bundle adjustment to refine the initial estimated results. Extensive experiments have been carried out which include simulations with and without outliers to verify the correctness of our method as well as to test its robustness to noise and outliers. The results of real experiments are also provided. The accuracy of our results is comparable to that of a state-of-the-art method that requires known 3D geometry of a scene.
Xida Chen, Yee-Hong Yang
CVPR2
2014 Frequency-Based 3D Reconstruction of Transparent and Specular Objects
abstract
3D reconstruction of transparent and specular objects is a very challenging topic in computer vision. For transparent and specular objects, which have complex interior and exterior structures that can reflect and refract light in a complex fashion, it is difficult, if not impossible, to use either passive stereo or the traditional structured light methods to do the reconstruction. We propose a frequency-based 3D reconstruction method, which incorporates the frequency-based matting method. Similar to the structured light methods, a set of frequency-based patterns are projected onto the object, and a camera captures the scene. Each pixel of the captured image is analyzed along the time axis and the corresponding signal is transformed to the frequency-domain using the Discrete Fourier Transform. Since the frequency is only determined by the source that creates it, the frequency of the signal can uniquely identify the location of the pixel in the patterns. In this way, the correspondences between the pixels in the captured images and the points in the patterns can be acquired. Using a new labelling procedure, the surface of transparent and specular objects can be reconstructed with very encouraging results.
Xida Chen, Yee-Hong Yang
CVPR3
2014 Robust multi-view L2 triangulation via optimal inlier selection and 3D structure refinement
Lai Kang, Lingda Wu, Yee-Hong Yang
Pattern Recognit.3
2014 Robust Volumetric Texture Classification of Magnetic Resonance Images of the Brain Using Local Frequency Descriptor
abstract
This paper presents a method for robust volumetric texture classification. It also proposes 2D and 3D gradient calculation methods designed to be robust to imaging effects and artifacts. Using the proposed 2D method, the gradient information is extracted on the XYZ orthogonal planes at each voxel and used to form a local coordinate system. The local coordinate system and the local 3D gradient computed by the proposed 3D gradient calculator are then used to define volumetric texture features. It is shown that the presented gradient calculation methods can be efficiently implemented by convolving with 2D and 3D kernels. The experimental results demonstrate that the proposed gradient operators and the texture features are robust to imaging effects and artifacts, such as blurriness and noise in 2D and 3D images. The proposed method is compared with three state-of- the-art volumetric texture classification methods the 3D gray level cooccurance matrix, 3D local binary patterns, and second orientation pyramid on magnetic resonance imaging data of the brain. The experimental results show the superiority of the proposed method in accuracy, robustness, and speed.
Rouzbeh Maani, Sanjay Kalra, Yee-Hong Yang
IEEE Trans. Image Process.3
2013 Underwater Camera Calibration Using Wavelength Triangulation
abstract
In underwater imagery, the image formation process includes refractions that occur when light passes from water into the camera housing, typically through a flat glass port. We extend the existing work on physical refraction models by considering the dispersion of light, and derive new constraints on the model parameters for use in calibration. This leads to a novel calibration method that achieves improved accuracy compared to existing work. We describe how to construct a novel calibration device for our method and evaluate the accuracy of the method through synthetic and real experiments.
Timothy Yau, Minglun Gong, Yee-Hong Yang
CVPR3
2013 A novel unsupervised approach for multilevel image clustering from unordered image collection
Lai Kang, Lingda Wu, Yee-Hong Yang
Frontiers Comput. Sci.3
2013 Practical structure and motion recovery from two uncalibrated images using ε Constrained Adaptive Differential Evolution
Lai Kang, Lingda Wu, Xida Chen, Yee-Hong Yang
Pattern Recognit.4
2013 Noise robust rotation invariant features for texture classification
Rouzbeh Maani, Sanjay Kalra, Yee-Hong Yang
Pattern Recognit.3
2013 Synthetic aperture imaging using pixel labeling via energy minimization
Zhao Pei, Yanning Zhang 0001, Xida Chen, Yee-Hong Yang
Pattern Recognit.4
2013 Rotation Invariant Local Frequency Descriptors for Texture Classification
abstract
This paper presents a novel rotation invariant method for texture classification based on local frequency components. The local frequency components are computed by applying 1-D Fourier transform on a neighboring function defined on a circle of radius R at each pixel. We observed that the low frequency components are the major constituents of the circular functions and can effectively represent textures. Three sets of features are extracted from the low frequency components, two based on the phase and one based on the magnitude. The proposed features are invariant to rotation and linear changes of illumination. Moreover, by using low frequency components, the proposed features are very robust to noise. While the proposed method uses a relatively small number of features, it outperforms state-of-the-art methods in three well-known datasets: Brodatz, Outex, and CUReT. In addition, the proposed method is very robust to noise and can remarkably improve the classification accuracy especially in the presence of high levels of noise.
Rouzbeh Maani, Sanjay Kalra, Yee-Hong Yang
IEEE Trans. Image Process.3
2012 Two-View Underwater Structure and Motion for Cameras under Flat Refractive Interfaces
Lai Kang, Lingda Wu, Yee-Hong Yang
ECCV (4)3
2012 Automatic Real-Time Video Matting Using Time-of-Flight Camera and Multichannel Poisson Equations
Liang Wang 0002, Minglun Gong, Ruigang Yang, Cha Zhang, Yee-Hong Yang
Int. J. Comput. Vis.6
2012 A novel multi-object detection method in complex scene using synthetic aperture imaging
Zhao Pei, Yanning Zhang 0001, Tao Yang 0006, Xiuwei Zhang 0001, Yee-Hong Yang
Pattern Recognit.5
2011 Physically based baking animations with smoothed particle hydrodynamics
Omar Rodriguez-Arenas, Yee-Hong Yang
Graphics Interface2
2011 A Component-Wise Analysis of Constructible Match Cost Functions for Global Stereopsis
abstract
Match cost functions are common elements of every stereopsis algorithm that are used to provide a dissimilarity measure between pixels in different images. Global stereopsis algorithms incorporate assumptions about the smoothness of the resulting distance map that can interact with match cost functions in unpredictable ways. In this paper, we present a large-scale study on the relative performance of a structured set of match cost functions within several global stereopsis frameworks. We compare 272 match cost functions that are built from component parts in the context of four global stereopsis frameworks with a data set consisting of 57 stereo image pairs at three different variances of synthetic sensor noise. From our analysis, we infer a set of general rules that can be used to guide derivation of match cost functions for use in global stereopsis algorithms.
Daniel Neilson, Yee-Hong Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Near-real-time stereo matching with slanted surface modeling and sub-pixel accuracy
Minglun Gong, Yee-Hong Yang
Pattern Recognit.3
2010 Background estimation using graph cuts and inpainting
Xida Chen, Yufeng Shen, Yee-Hong Yang
Graphics Interface3
2010 Real-time video matting using multichannel poisson equations
Minglun Gong, Liang Wang 0002, Ruigang Yang, Yee-Hong Yang
Graphics Interface4
2009 Feature based classification of computer graphics and real images
abstract
Photorealistic images can now be created using advanced techniques in computer graphics (CG). Synthesized elements could easily be mistaken for photographic (real) images. Therefore we need to differentiate between CG and real images. In our work, we propose and develop a new framework based on an aggregate of existing features. Our framework has a classification accuracy of 90% when tested on the de facto standard Columbia dataset, which is 4% better than the best results obtained by other prominent methods in this area. We further show that using feature selection it is possible to reduce the feature dimension of our framework from 557 to 80 without a significant loss in performance (Lt 1%). We also investigate different approaches that attackers can use to fool the classification system, including creation of hybrid images and histogram manipulations. We then propose and develop filters to effectively detect such attacks, thereby limiting the effect of such attacks to our classification system.
Gopinath Sankar, H. Vicky Zhao, Yee-Hong Yang
ICASSP3
2009 A new multiview spacetime-consistent depth recovery framework for free viewpoint video rendering
abstract
In this paper, we present a new approach for recovering spacetime-consistent depth maps from multiple video sequences captured by stationary, synchronized and calibrated cameras for depth based free viewpoint video rendering. Our two-pass approach is generalized from the recently proposed region-tree based binocular stereo matching method. In each pass, to enforce temporal consistency between successive depth maps, the traditional region-tree is extended into a temporal one by including connections to “temporal neighbor regions” in previous video frames, which are identified using estimated optical flow information. For enforcing spatial consistency, multi-view geometric constraints are used to identify inconsistencies between depth maps among different views which are captured in an inconsistency map for each view. Iterative optimizations are performed to progressively correct inconsistencies through inconsistency maps based depth hypotheses pruning and visibility reasoning. Furthermore, the background depth and color information is generated from the results of the first pass and is used in the second pass to enforce sequence-wise temporal consistency and to aid in identifying and correcting spatial inconsistencies. The extensive experimental evaluations have shown that our proposed approach is very effective in producing spatially and temporally consistent depth maps.
Xida Chen, Yee-Hong Yang
ICCV3
2009 Optical flow estimation on coarse-to-fine region-trees using discrete optimization
abstract
In this paper, we propose a new region-based method for accurate motion estimation using discrete optimization. In particular, the input image is represented as a tree of over-segmented regions and the optical flow is estimated by optimizing an energy function defined on such a region-tree using dynamic programming. To accommodate the sampling-inefficiency problem intrinsic to discrete optimization compared to the continuous optimization based methods, both spatial and solution domain coarse-to-fine (C2F) strategies are used. That is, multiple region-trees are built using different over-segmentation granularities. Starting from a global displacement label discretization, optical flow estimation on the coarser level region-tree is used for defining region-wise finer displacement samplings for finer level region-trees. Furthermore, cross-checking based occlusion detection and correction and continuous optimization are also used to improve accuracy. Extensive experiments using the Middlebury benchmark datasets have shown that our proposed method can produce top-ranking results.
Yee-Hong Yang
ICCV2
2009 Music-driven character animation
abstract
Music-driven character animation extracts musical features from a song and uses them to create an animation. This article presents a system that builds a new animation directly from musical attributes, rather than simply synchronizing it to the music like similar systems. Using a simple script that identifies the movements involved in the performance and their timing, the user can easily control the animation of characters. Another unique feature of the system is its ability to incorporate multiple characters into the same animation, both with synchronized and unsynchronized movements. A system that integrates Celtic dance movements is developed in this article. An evaluation of the results shows that the majority of animations are found to be appealing to viewers and that altering the music can change the attractiveness of the final result.
Danielle Sauer, Yee-Hong Yang
ACM Trans. Multim. Comput. Commun. Appl.2
2008 Evaluation of constructable match cost measures for stereo correspondence using cluster ranking
abstract
Stereo correspondence research often involves the comparison of techniques to determine which are better under different circumstances. The methods of comparison employed often take the form of applying the techniques to a few stereo image pairs with the technique with the lowest error rate declared superior. However, the majority of these comparisons do not contain any discussion of statistical significance; making the declared superiority of a technique statistically unreliable. In this paper we present a new evaluation method called cluster ranking that yields a statistically significant comparison of the stereo techniques being compared. Cluster ranking leverages statistical inference techniques to first rank the performance of stereo techniques on a single stereo image pair and then combine the rankings from multiple stereo pairs into an over-all ranking; in both of these rankings, only stereo techniques that are statistically different are given different ranks. We demonstrate our framework with a comparison of constructable match cost measures (those that can be assembled from a base set of components) on a data set consisting of 30 synthetic stereo pairs, with varying amounts of noise, and 18 scenes from the 2005 and 2006 Middlebury data sets. Our analysis reveals match cost measures, and measure components, that are statistically superior to all other measures depending on amount of noise, illumination, or exposure time.
Daniel Neilson, Yee-Hong Yang
CVPR2
2008 Local stereo matching with 3D adaptive cost aggregation for slanted surface modeling and sub-pixel accuracy
abstract
This paper presents a new local binocular stereo algorithm which takes into consideration plane fitting at the per-pixel level. Two disparity calculation passes are used. The first pass assumes that surfaces in the scene are fronto-parallel and generates an initial disparity map, from which the disparity plane orientations of all pixels are extracted and refined. In the second pass, the cost aggregation for each pixel is conducted along the estimated disparity plane orientations, rather than the fronto-parallel ones. Large window size with adaptive support weights is used to ensure the effectiveness of the slanted surface modeling. The disparity search space is also quantized at sub-pixel level to improve the accuracy of the disparity results. The experimental results demonstrate the validity of our presented approach.
Minglun Gong, Yee-Hong Yang
ICPR3
2007 Real-time backward disparity-based rendering for dynamic scenes using programmable graphics hardware
abstract
This paper presents a backward disparity-based rendering algorithm, which runs at real-time speed on programmable graphics hardware. The algorithm requires only a handful of image samples of the scene and estimated noisy disparity maps, whereas most existing techniques need either dense samples or accurate depth information. To color a given pixel in the novel view, a backward searching process is conducted to find the corresponding pixels from the closest four reference images. The use of backward searching process makes the algorithm more robust to errors in estimated disparity maps than existing forward warping-based approaches. In addition, since the computations for different pixels are independent, they can be performed in parallel on the Graphics Processing Units of modern graphics hardware. Experiment results demonstrate that our algorithm can synthesize accurate novel views for dynamic real scenes at a high frame rate.
Minglun Gong, Jason M. Selzer, Yee-Hong Yang
Graphics Interface4
2007 Real-Time Stereo Matching Using Orthogonal Reliability-Based Dynamic Programming
abstract
A novel algorithm is presented in this paper for estimating reliable stereo matches in real time. Based on the dynamic programming-based technique we previously proposed, the new algorithm can generate semi-dense disparity maps using as few as two dynamic programming passes. The iterative best path tracing process used in traditional dynamic programming is replaced by a local minimum searching process, making the algorithm suitable for parallel execution. Most computations are implemented on programmable graphics hardware, which improves the processing speed and makes real-time estimation possible. The experiments on the four new Middlebury stereo datasets show that, on an ATI Radeon X800 card, the presented algorithm can produce reliable matches for 60% approximately 80% of pixels at the rate of 10 approximately 20 frames per second. If needed, the algorithm can be configured for generating full density disparity maps.
Minglun Gong, Yee-Hong Yang
IEEE Trans. Image Process.2
2007 Aura 3D Textures
abstract
This paper presents a new technique, called aura 3D textures, for generating solid textures based on input examples. Our method is fully automatic and requires no user interactions in the process. Given an input texture sample, our method first creates its aura matrix representations and then generates a solid texture by sampling the aura matrices of the input sample constrained in multiple view directions. Once the solid texture is generated, any given object can be textured by the solid texture. We evaluate the results of our method based on extensive user studies. Based on the evaluation results using human subjects, we conclude that our algorithm can generate faithful results of both stochastic and structural textures with an average successful rate of 76.4 percent. Our experimental results also show that the new method outperforms Wei and Levoy's method and is comparable to that proposed by Jagnow et al. [21].
Xuejie Qin, Yee-Hong Yang
IEEE Trans. Vis. Comput. Graph.2
2006 Region-Tree Based Stereo Using Dynamic Programming Optimization
abstract
In this paper, we present a novel stereo algorithm that combines the strengths of region-based stereo and dynamic programming on a tree approaches. Instead of formulating an image as individual scan-lines or as a pixel tree, a new region tree structure, which is built as a minimum spanning tree on the adjacency-graph of an over-segmented image, is used for the global dynamic programming optimization. The resulting disparity maps do not contain any streaking problem as is common in scanline-based algorithms because of the tree structure. The performance evaluation using the Middlebury benchmark datasets shows that the performance of our algorithm is comparable in accuracy and efficiency with top ranking algorithms.
Jason M. Selzer, Yee-Hong Yang
CVPR (2)3
2006 Particle-based immiscible fluid-fluid collision
Hai Mao, Yee-Hong Yang
Graphics Interface2
2006 Estimate Large Motions Using the Reliability-Based Motion Estimation Algorithm
Minglun Gong, Yee-Hong Yang
Int. J. Comput. Vis.2
2006 Tri-focal tensor-based multiple video synchronization with subframe optimization
abstract
In this paper, we present a novel method for synchronizing multiple (more than two) uncalibrated video sequences recording the same event by free-moving full-perspective cameras. Unlike previous synchronization methods, our method takes advantage of tri-view geometry constraints instead of the commonly used two-view one for their better performance in measuring geometric alignment when video frames are synchronized. In particular, the tri-ocular geometric constraint of point/line features, which is evaluated by tri-focal transfer, is enforced when building the timeline maps for sequences to be synchronized. A hierarchical approach is used to reduce the computational complexity. To achieve subframe synchronization accuracy, the Levenberg-Marquardt method-based optimization is performed. The experimental results on several synthetic and real video datasets demonstrate the effectiveness and robustness of our method over previous methods in synchronizing full-perspective videos.
Yee-Hong Yang
IEEE Trans. Image Process.2
2005 Near Real-Time Reliable Stereo Matching Using Programmable Graphics Hardware
abstract
A near-real-time stereo matching technique is presented in this paper, which is based on the reliability-based dynamic programming algorithm we proposed earlier. The new algorithm can generate semi-dense disparity maps using only two dynamic programming passes, while our previous approach requires 20-30 passes. We also implement the algorithm on programmable graphics hardware, which further improves the processing speed. The experiments on the four Middlebury stereo datasets show that the new algorithm can produce dense (>85% of the pixels) and reliable (error rate <0.3%) matches in near real-time (0.05-0.1 sec). If needed, it can also be used to generate dense disparity maps. Based on the evaluation conducted by the Middlebury Stereo Vision Research Website, the new algorithm is ranked between the variable window and the graph cuts approaches and currently is the most accurate dynamic programming based technique. When more than one reference images are available, the accuracy can be further improved with little extra computation time.
Minglun Gong, Yee-Hong Yang
CVPR (1)2
2005 Basic Gray Level Aura Matrices: Theory and its Application to Texture Synthesis
abstract
In this paper, we present a new mathematical framework for modeling texture images using independent basic gray level aura matrices (BGLAMs). We prove that independent BGLAMs are the basis of gray level aura matrices (GLAMs), and that an image can be uniquely represented by its independent BGLAMs. We propose a new BGLAM distance measure for automatically evaluating synthesis results w.r.t. input textures to determine if the output is a successful synthesis of the input. For the application to texture synthesis, we present a new algorithm to synthesize textures by sampling only the independent BGLAMs of an input texture. With respect to synthesis of textures and evaluation of the results, the performance of our approach is extensively evaluated and compared with symmetric GLAMs that are used in existing techniques and with gray level cooccurrence matrices (GLCMs). Experimental results have shown that (1) our approach significantly outperforms both symmetric GLAMs and GLCMs; (2) the new BGLAM distance measure has the ability to evaluate synthesis results, which can be used to automate the conventional visual inspection process for determining whether or not the output texture is a successful synthesis of the input; and (3) a broad range of textures can be faithfully synthesized using independent BGLAMs and the synthesis results are comparable to existing techniques.
Xuejie Qin, Yee-Hong Yang
ICCV2
2005 Camera field rendering for static and dynamic scenes
Minglun Gong, Yee-Hong Yang
Graph. Model.2
2005 Fast Unambiguous Stereo Matching Using Reliability-Based Dynamic Programming
abstract
An efficient unambiguous stereo matching technique is presented in this paper. Our main contribution is to introduce a new reliability measure to dynamic programming approaches in general. For stereo vision application, the reliability of a proposed match on a scanline is defined as the cost difference between the globally best disparity assignment that includes the match and the globally best assignment that does not include the match. A reliability-based dynamic programming algorithm is derived accordingly, which can selectively assign disparities to pixels when the corresponding reliabilities exceed a given threshold. The experimental results show that the new approach can produce dense (> 70 percent of the unoccluded pixels) and reliable (error rate < 0.5 percent) matches efficiently (< 0.2 sec on a 2GHz P4) for the four Middlebury stereo data sets.
Minglun Gong, Yee-Hong Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Similarity Measure and Learning with Gray Level Aura Matrices (GLAM) for Texture Image Retrieval
Xuejie Qin, Yee-Hong Yang
CVPR (1)2
2004 Object Representation using 1D Displacement Mapping
Yee-Hong Yang
Graphics Interface2
2004 Estimate large motions using reliability-based dynamic programming
abstract
Detecting and estimating motions of fast moving objects has many important applications. However, most existing motion estimation techniques have difficulties in handling large motions in the scene. In this paper, the reliability-based dynamic programming technique proposed by Gong and Yang is extended and applied to large motion estimation problem. Compared with the Gong and Yang approach, the extended algorithm removes the constant penalty assumption and also explicitly enforces the inter-scanline consistency constraint. The experimental results indicate that the new algorithm can effectively estimate velocities for fast moving objects. The algorithm can also be configured to produce sparse but reliable flow fields.
Minglun Gong, Yee-Hong Yang
ICIP2
2004 Frequency-Based Environment Matting
abstract
Environment matting is a technique to extract the environment matte which is used to describe how an object reflects and refracts the environment light. In this paper, we propose a novel environment matting method to obtain the environment matte of a real scene. Previous methods use different backdrops as the calibration patterns and search for the environment matte in the spatial domain. In our method, however, a series of background images displayed on a screen sequentially in time are interpreted as signals. The frequency similarity of these signals is used as the searching criterion. The frequencies of these signals are not changed when they interact with the foreground objects and thus can be used to extract the environment matte. While using correspondence in the spatial domain in existing approaches is prone to error, using frequency correspondence is not. Thus, our approach is robust to noise and can easily deal with some of the complex light transport phenomena which cannot be easily handled using current methods. The experimental results are very encouraging.
Jiayuan Zhu, Yee-Hong Yang
PG2
2004 Quadtree-based genetic algorithm and its applications to computer vision
Minglun Gong, Yee-Hong Yang
Pattern Recognit.2
2003 Fast Stereo Matching Using Reliability-Based Dynamic Programming and Consistency Constraints
abstract
A method for solving binocular and multiview stereo matching problems is presented here. A weak consistency constraint is proposed, which expresses the visibility constraint in the image space. It can be proved that the weak consistency constraint holds for scenes that can be represented by a set of 3D points. As well, also proposed is a new reliability measure for dynamic programming techniques, which evaluates the reliability of a given match. A novel reliability-based dynamic programming algorithm is derived accordingly, which can selectively assign disparity values to pixels when the reliabilities of the corresponding matches exceed a given threshold. Consistency constraints and the new reliability-based dynamic programming algorithm can be combined in an iterative approach. The experimental results show that the iterative approach can produce dense (60-90%) and reliable (total error rate of 0.1-1.1%) matching for binocular stereo datasets. It can also generate promising disparity maps for trinocular and multiview stereo datasets.
Minglun Gong, Yee-Hong Yang
ICCV2
2003 Variance Invariant Adaptive Temporal Supersampling for Motion Blurring
abstract
Adaptive temporal sampling, used to create motion blur in distributed ray tracing, generates more sample points in regions with motion blur than in regions without motion blur. When the number of sample points used on stationary objects in regions with motion blur exceeds the number of sample points used in other regions of the image, the variance in the color of the object can differ between the two regions. This paper identifies the cause of this variance discrepancy, and proposes a modification to existing adaptive temporal sampling algorithms which eliminated it. Our results demonstrate that the variance of stationary objects remains approximately the same throughout the entire image and that the proposed modification is capable of improving the running time of existing adaptive temporal sampling algorithms.
Daniel Neilson, Yee-Hong Yang
PG2
2002 Estimating Parameters for Procedural Texturing by Genetic Algorithms
Xuejie Qin, Yee-Hong Yang
Graph. Model.2
2002 Genetic-Based Stereo Algorithm and Disparity Map Evaluation
Minglun Gong, Yee-Hong Yang
Int. J. Comput. Vis.2
2002 Face recognition approach based on rank correlation of Gabor-filtered images
Olugbenga Ayinde, Yee-Hong Yang
Pattern Recognit.2
2002 Region-based face detection
Olugbenga Ayinde, Yee-Hong Yang
Pattern Recognit.2
2001 The Rayset and Its Applications
Minglun Gong, Yee-Hong Yang
Graphics Interface2
2001 Physics-Based Explosion Modeling
Byron Bashforth, Yee-Hong Yang
Graph. Model.2
2001 Layer-Based Morphing
Minglun Gong, Yee-Hong Yang
Graph. Model.2
2001 Multiple Illuminant Direction Detection with Application to Image Synthesis
abstract
Pentland observed (1982, 1984) that the human eye is sensitive to the change of intensities. On an image of a smooth surface, the change of intensities is maximal whenever the illuminant direction is perpendicular to the normal of the surface. This motivates us to introduce the concept of critical points, where the surface normal is perpendicular to some light source direction. Apparently, the illuminant direction has a simple geometric relationship with the corresponding critical points. In this paper, for simplicity reasons, we restrict our discussions to the shading of a Lambertian sphere of known size in a multiple distant light source environment. A novel global representation of the intensity function is derived. Based on this intensity characterization, the least-squares and iteration techniques are used to determine critical points and, thus, the light source directions and their intensities if certain conditions are satisfied. The performance of this new approach is evaluated using both synthetic images and real images. As an application, we use it as a tool to determine light sources in real image synthesis. The experimental results show that this technique can be used to superimpose synthetic objects with a real scene.
Yufei Zhang 0012, Yee-Hong Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Illuminant Direction Determination for Multiple Light Sources
abstract
In this paper, a novel scheme of extracting multiple illuminant directions from an image of a Lambertian sphere of known size is proposed. The illuminant direction detection process is based on the concept of critical points introduced in the paper. We show that the illuminant directions have a close relationship to those critical points and that, by identifying those critical points as many as possible, illuminate direction may be recovered if certain conditions are satisfied. Our preliminary, experimental results show that illumination information can be obtained accurately.
Yufei Zhang 0012, Yee-Hong Yang
CVPR2
1999 A Fast Rule-Based Parameter Free Discrete Hough Transform
abstract
This paper introduces a new discrete Hough transform, DHT, that pre-computes discrete line information (rules) and uses this information to detect line segments in the image. Pre-computing line information removes the need for run-time line calculations and the associated parameters. The proposed approach does not depend on the parameterization of a straight line and is formulated based on the discrete domain. This new DHT is compared with selected existing techniques to demonstrate the large reduction in computation time achieved by this new approach, while not sacrificing accuracy.
Bernie M. A. Genswein, Yee-Hong Yang
Int. J. Pattern Recognit. Artif. Intell.2
1999 Theoretical analysis of illumination in PCA-based vision systems
Yee-Hong Yang
Pattern Recognit.2
1999 An efficient algorithm to compute eigenimages in PCA-based vision systems
Yee-Hong Yang
Pattern Recognit.2
1999 Mosaic image method: a local and global method
Yee-Hong Yang
Pattern Recognit.2
1999 Bounded diffusion for multiscale edge detection using regularized cubic B-spline fitting
abstract
This paper shows that in edge detection the regularization factor alpha is a better scale parameter than the standard deviation (sigma) of the Gaussian pre-filter. The alpha scale space, which exhibits the evolutionary behaviour of an edge in various scales, is the basis for the design of a multiscale edge detector (MRCBS). In MRCBS, the scale is determined adaptively according to the local noise level; the thresholds which control the amount of edge details are adjusted according to the scale; and the anisotropic diffusion is applied in the finest scale to further suppress noise.
Kung-Hao Liang, Tardi Tjahjadi, Yee-Hong Yang
IEEE Trans. Syst. Man Cybern. Part B3
1998 Deformable Object Modeling Using the Time-Dependent Finite Element Method
Yee-Hong Yang
Graph. Model. Image Process.2
1998 Skeletonisation: An electrostatic field-based approach
T. Grigorishin, Gamal H. Abdel-Hamid, Yee-Hong Yang
Pattern Anal. Appl.3
1998 Classifier design with incomplete knowledge
Russell E. Muzzolini, Yee-Hong Yang, Roger A. Pierson
Pattern Recognit.2
1997 Modeling water for computer graphics
David Mould, Yee-Hong Yang
Comput. Graph.2
1997 Towards Developing A Pratical System to Recover Light, Reflectance and Shape
abstract
Due to the complexity of the shape-from-shading problem, most solutions rely on idealistic conditions. Orthographic imaging, a known distant point light source, and known surface reflectance properties are usually assumed. Furthermore, most real surfaces are neither perfectly diffuse (Lambertian) nor ideally specular (mirror-like); however most shape-from-shading algorithms assume Lambertian reflectance. The behavior of shape-from-shading algorithms that rely on idealistic conditions is unpredictable in real imaging situations. In this paper, the LIRAS (LIght, Reflectance, And Shape) Recovery System is proposed. LIRAS is a practical approach to the shape-from-shading problem, as many of these assumptions are relaxed. LIRAS is also a modular system: there is a component that recovers the surface reflectance properties, thus the assumption of Lambertian reflectance is relaxed. Rather than assume a known illuminant direction, a component exists that can recover the light orientation. Once the reflectance map is determined, another LIRAS module can use this information to recover the shape for non-Lambertian surfaces. Each of these modules is described and a discussion of how the components cooperate to recover three-dimensional shape information in real environments is given. Extensive experimental evaluation is conducted using both synthetic and real images and the results are very encouraging. The contributions of this paper include the design and implementation of LIRAS and the extensive quantative and qualitative experimental results, which can provide guidelines for future refinements of other shape recovery systems.
Sanjay Bakshi, Yee-Hong Yang
Int. J. Pattern Recognit. Artif. Intell.2
1997 Roof edge detection using regularized cubic b-spline fitting
Kung-Hao Liang, Tardi Tjahjadi, Yee-Hong Yang
Pattern Recognit.3
1997 Three-Dimensional Surface Reconstruction Using Optical Flow for Medical Imaging
abstract
The recovery of a three-dimensional (3-D) model from a sequence of two-dimensional (2-D) images is very useful in medical image analysis. Image sequences obtained from the relative motion between the object and the camera or the scanner contain more 3-D information than a single image. Methods to visualize the computed tomograms can be divided into two approaches: the surface rendering approach and the volume rendering approach. In this paper, a new surface rendering method using optical flow is proposed. Optical flow is the apparent motion in the image plane produced by the projection of the real 3-D motion onto 2-D image. The 3-D motion of an object can be recovered from the optical-flow field using additional constraints. By extracting the surface information from 3-D motion, it is possible to get an accurate 3-D model of the object. Both synthetic and real image sequences have been used to illustrate the feasibility of the proposed method. The experimental results suggest that the proposed method is suitable for the reconstruction of 3-D models from ultrasound medical images as well as other computed tomograms.
Nan Weng, Yee-Hong Yang, Roger A. Pierson
IEEE Trans. Medical Imaging2
1995 A modified dichromatic reflection model for an analysis of interreflection
abstract
Interreflection changes image intensities in a consistent, non-random way, and if unaccounted for can easily confuse computer vision algorithms. This paper proposes an image model which enables regions of interreflection on objects of inhomogeneous dielectric materials to be identified. The model is based on the dichromatic reflection model and utilises the concept of the one-bounce model of mutual reflection between two object surfaces. One surface is viewed as a source of low-intensity illumination and the second as a function of the light emitted from the first. The proposed model indicates the presence of an additional matte cluster which corresponds to the region of interreflection.
Tardi Tjahjadi, D. Litwin, Yee-Hong Yang
ICIP3
1995 First Sight: A Human Body Outline Labeling System
abstract
First Sight, a vision system in labeling the outline of a moving human body, is proposed in this paper. The emphasis of First Sight is on the analysis of motion information gathered solely from the outline of a moving human object. Two main processes are implemented in First Sight. The first process uses a novel technique to extract the outline of a moving human body from an image sequence. The second process, which employs a new human body model, interprets the outline and produces a labeled two-dimensional human body stick figure for each frame of the image sequence. Extensive knowledge of the structure, shape, and posture of the human body is used in the model. The experimental results of applying the technique on unedited image sequences with self-occlusions and missing boundary lines are encouraging.>
Maylor K. H. Leung, Yee-Hong Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Multiresolution Skeletonization: An Electrostatic Field-Based Approach
abstract
Skeleton representation of an object is believed to be a powerful representation that captures both boundary and region information of the object. The skeleton of a shape is a representation composed of idealized thin lines that preserve the connectivity or topology of the original shape. Although the literature contains a large number of skeletonization algorithms, many open problems remain. A new skeletonization approach that relies on the electrostatic field theory (EFT) is proposed. Many problems associated with existing skeletonization algorithms are solved using the proposed approach. In particular, connectivity, thinness, and other desirable features of a skeleton are guaranteed. Furthermore, the electrostatic field-based approach captures notions of corner detection, multiple scale, thinning, and skeletonization all within one unified framework. Experimental results are very encouraging and are used to illustrate the potential of the proposed approach.>
Gamal H. Abdel-Hamid, Yee-Hong Yang
ICIP (1)2
1994 Shape from Shading for Non-Lambertian Surfaces
abstract
It is known that most real surfaces are neither perfectly diffuse (Lambertian) nor ideally specular (mirror-like); however, most shape-from-shading algorithms assume Lambertian reflectance. It is necessary to develop new techniques to solve the shape-from-shading problem. These techniques must be able to recover the shape of objects whose surfaces are not necessarily Lambertian. A new heuristic-based algorithm, called the general shading logic algorithm, is proposed to recover the shape of objects whose surfaces are non-Lambertian. This algorithm is based on the shading logic algorithm recently proposed by Vega and Yang (see IEEE Transactions on Pattern Analysis and Machine Intelligence, vol.15, no.6, p.592-597, 1993). The proposed algorithm is flexible enough to work with a more general reflectance model. To demonstrate that the proposed algorithm can cope with a wide range of reflectance models, a physically-based model for light reflection is used that can approximate rough surfaces. The model of light reflection used is similar to the Torrance-Sparrow (1967) approach. The general shading logic algorithm has been implemented and evaluated experimentally. The experimental results of the proposed algorithm are very encouraging and the performance is demonstrated by extensive experiments using a wide variety of synthesized and real objects.>
Sanjay Bakshi, Yee-Hong Yang
ICIP (2)2
1994 Three Dimensional Segmentation of Volume Data
abstract
Most of the research in image segmentation has focused on segmenting 2D images. When 2D segmentation techniques are applied to 3D data, the potential increase in information available in the third dimension is not typically used. The 3D multiresolution texture segmentation algorithm (3D MTS) is a proposed approach for incorporating the information in the third dimension by segmenting 3D data into homogeneous volumes. The 3D MTS algorithm is based on previous work which uses texture at multiple resolutions to determine the homogeneity of regions within an image. Experiments were performed using both synthesized and real volume data. The results demonstrate that the proposed approach is robust in the presence of noise and produces accurate segmentation results. The results also show that the 2D MTS algorithm performs well on the given data, however, it is apparent that for more complex 3D textures the 3D MTS algorithm should be able to more accurately identify homogeneous regions. Further experimentation is being performed to help validate this hypothesis.>
Russell E. Muzzolini, Yee-Hong Yang, Roger A. Pierson
ICIP (3)2
1994 Multiresolution Color Image Segmentation
abstract
Image segmentation is the process by which an original image is partitioned into some homogeneous regions. In this paper, a novel multiresolution color image segmentation (MCIS) algorithm which uses Markov random fields (MRF's) is proposed. The proposed approach is a relaxation process that converges to the MAP (maximum a posteriori) estimate of the segmentation. The quadtree structure is used to implement the multiresolution framework, and the simulated annealing technique is employed to control the splitting and merging of nodes so as to minimize an energy function and therefore, maximize the MAP estimate. The multiresolution scheme enables the use of different dissimilarity measures at different resolution levels. Consequently, the proposed algorithm is noise resistant. Since the global clustering information of the image is required in the proposed approach, the scale space filter (SSF) is employed as the first step. The multiresolution approach is used to refine the segmentation. Experimental results of both the synthesized and real images are very encouraging. In order to evaluate experimental results of both synthesized images and real images quantitatively, a new evaluation criterion is proposed and developed.>
Jianqing Liu, Yee-Hong Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Texture characterization using robust statistics
Russell E. Muzzolini, Yee-Hong Yang, Roger A. Pierson
Pattern Recognit.2
1993 Shading Logic: A Heuristic Approach to Recover Shape from Shading
abstract
A heuristic-based algorithm known as the shading logic algorithm is proposed for recovering shape from shading. The heuristics are derived on the basis of a geometrical interpretation of the M.J. Brooks and B.K.P. Horn (1985) algorithm. An experimental evaluation was performed using synthesized objects, in particular, superquadrics. The advantage of using the superquadrics is that the shape of the objects can be varied incrementally and systematically. Despite the fact that the shading logic algorithm is heuristic based, experimental results show that the proposed algorithm has a better performance than the Brooks and Horn algorithm. In addition, the proposed approach does not seem to suffer the stability problem common to most variational-based methods.>
Omar E. Vaga, Yee-Hong Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Multiresolution texture segmentation with application to diagnostic ultrasound images
abstract
A multiresolution texture segmentation (MTS) approach to image segmentation that addresses the issues of texture characterization, image resolution, and time to complete the segmentation is presented. The approach generalizes the conventional simulated annealing method to a multiresolution framework and minimizes an energy function that is dependent on the resolution of the size of the texture blocks in an image. A rigorous experimental procedure is also proposed to demonstrate the advantages of the proposed MTS approach on the accuracy of the segmentation, the efficiency of the algorithm, and the use of varying features at different resolution. Semireal images, created by sampling a series of diagnostic ultrasound images of an ovary in vitro, were tested to produce statistical measures on the performance of the approach. The ultrasound images themselves were then segmented to determine if the approach can achieve accurate results for the intended ultrasound application. Experimental results suggest that the MTS approach converges faster and produces better segmentation results than the single-level approach.
Russell E. Muzzolini, Yee-Hong Yang, Roger A. Pierson
IEEE Trans. Medical Imaging2
1991 Experimental evaluation of motion constraint equations
Darryl L. Willick, Yee-Hong Yang
CVGIP Image Underst.2
1991 Log-Tracker: an Attribute-Based Approach to Tracking Human Body Motion
abstract
Motion provides extra information that can aid in the recognition of objects. One of the most commonly seen objects is, perhaps, the human body. Yet little attention has been paid to the analysis of human motion. One of the key steps required for a successful motion analysis system is the ability to track moving objects. In this paper, we describe a new system called Log-Tracker, which was recently developed for tracking the motion of the different parts of the human body. Occlusion of body parts is termed a forking condition. Two classes of forks as well as the attributes required to classify them are described. Experimental results from two gymnastics sequences indicate that the system is able to track the body parts even when they are occluded for a short period of time. Occlusions that extend for a long period of time still pose problems to Log-Tracker.
Warren Long, Yee-Hong Yang
Int. J. Pattern Recognit. Artif. Intell.2
1991 The background primal sketch: An approach for tracking moving objects
Yee-Hong Yang, Martin D. Levine
Mach. Vis. Appl.1
1990 Dynamic strip algorithm in curve fitting
Maylor K. H. Leung, Yee-Hong Yang
Comput. Vis. Graph. Image Process.2
1990 Generalized Multidimensional Orthogonal Polynomials with Applications to Shape Analysis
abstract
A technique using the generalized multidimensional orthogonal polynomials (GMDOP) for 2-D shape analysis is proposed. In shape analysis, spatial invariances (i.e. translational invariance, scaling invariance, rotational invariance, etc.) are important requirements for a shape analysis algorithm. The described technique provides not only the three invariant properties but also mirror-image rotational invariance and permutational invariance. Experimental results supporting the theory are presented.>
Yee-Hong Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Dynamic two-strip algorithm in curve fitting
Maylor K. H. Leung, Yee-Hong Yang
Pattern Recognit.2
1990 Stationary background generation: An alternative to the difference of two images
Warren Long, Yee-Hong Yang
Pattern Recognit.2
1990 Comparison of two shape-from-shading algorithms
Sudarsan Tandri, Yee-Hong Yang
Pattern Recognit. Lett.2
1988 A new technique for shape analysis using orthogonal polynomials
Yee-Hong Yang
Pattern Recognit. Lett.2
1987 A fast two-dimensional line clipping algorithm via line encoding
Mark S. Sobkow, Paul Pospisil, Yee-Hong Yang
Comput. Graph.3
1987 Human body motion segmentation in a complex scene
Maylor K. H. Leung, Yee-Hong Yang
Pattern Recognit.2
1987 A region based approach for human body motion analysis
Maylor K. H. Leung, Yee-Hong Yang
Pattern Recognit.2
1983 An Evaluation Study of Six Topologies of Parallel Computer Architectures for Scene Matching
Yee-Hong Yang, Tsung-Wei Sze
ICPP1