EDBT 2026 Demo / reviewers in the wild / expert
Tao Yan 0001
dblp:31/1175-1
· DBLP profile ↗
28ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0002-9162-8551ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVGD-Net: A Novel Motion-aware Video Glass Surface Detection MethodabstractGlass surface ubiquitous in both daily life and professional environments presents a potential threat to vision-based systems, such as robot and drone navigation. To solve this challenge, most recent studies have shown significant interest in Video Glass Surface Detection (VGSD). We observe that objects in the reflection (or transmission) layer appear farther from the glass surfaces. Consequently, in video motion scenarios, the notable reflected (or transmitted) objects on the glass surface move slower than objects in non-glass regions within the same spatial plane, and this motion inconsistency can effectively reveal the presence of glass surfaces. Based on this observation, we propose a novel network, named MVGD-Net, for detecting glass surfaces in videos by leveraging motion inconsistency cues. Our MVGD-Net features three novel modules: the Cross-scale Multimodal Fusion Module (CMFM) that integrates extracted spatial features and estimated optical flow maps, the History Guided Attention Module (HGAM) and Temporal Cross Attention Module (TCAM), both of which further enhances temporal features. A Temporal-Spatial Decoder (TSD) is also introduced to fuse the spatial and temporal features for generating the glass region mask. Furthermore, for learning our network, we also propose a large-scale dataset, which comprises 312 diverse glass scenarios with a total of 19,268 frames. Extensive experiments demonstrate that our MVGD-Net outperforms relevant state-of-the-art methods. We will release our code and dataset. Tao Yan 0001 |
AAAI | 3 |
| 2026 | SEP-YOLO: Fourier-Domain Feature Representation for Transparent Object Instance Segmentation
Tao Yan 0001, Jianchao Huang |
ISCAS | 2 |
| 2026 | Teeth-GS: Gaussian Splatting Diffusion with enamel reflectance prior for single-image tooth crown reconstruction
Yanxing Liang, Yinghui Wang 0001, Jinlong Yang 0002, Tao Yan 0001, Jiaxing Shen |
Medical Image Anal. | 5 |
| 2026 | High-resolution image deraining via dual-branch features interaction and fusion
Weilong Huang, Jiaxue Mei, Tao Yan 0001, Yinghui Wang 0001, Xiaojun Chang |
Neural Networks | 3 |
| 2026 | Textureless Surface Feature Point Detection via Micro-Geometry ReconstructionabstractFeature point detection on textureless surfaces remains a fundamental challenge in computer vision due to the absence of discernible color and brightness gradients. From the imaging mechanism perspective, micro-geometry structures of textureless surfaces provide physically stable cues for feature point extraction despite the absence of visual distinctiveness. Therefore, we propose a novel feature point detection method, which reconstructs surface micro-geometry structures from a single RGB image and leverages these micro-geometry structures for feature extraction, without relying on specialized equipment or complex deep learning models. Specifically, our method establishes a novel framework that models the light-surface interaction to analyze phase modulation in reflected light. Then it reconstructs underlying micro-geometry structures through Gabor Kernel-based spectral analysis, enabling accurate quantification of surface height variations from phase information. This information forms the foundation of our proposed Concave-Convex Index (CCI), a robust geometric descriptor that achieves stable feature characterization through geometry-aware measurements. Extensive evaluations on TUM, T-LESS, Shape2.5D datasets and self-collected images, demonstrate our method's superior capability in extracting stably distributed and highly repeatable feature points, even when visible texture or brightness gradients vanish. Our method offers a novel perspective for reliable feature point detection on challenging textureless surfaces across diverse materials and illumination conditions. Yanxing Liang, Yinghui Wang 0001, Tao Yan 0001, Jinlong Yang 0002, Wei Li 0121, Liangyi Huang, Xiaojuan Ning, Temurbek Kuchkorov |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | MDeRainNet: An Efficient Macro-pixel Image Rain Removal NetworkabstractSince raining weather always degrades image quality and poses significant challenges to most computer vision-based intelligent systems, image de-raining has been a hot research topic in computer vision community. Fortunately, in a rainy Light Field (LF) image, background obscured by rain streaks in one sub-view may be visible in the other sub-views, and implicit depth information and recorded 4D structural information may benefit rain streak detection and removal. However, existing LF image rain removal methods either do not fully exploit the global correlations of 4D LF data or only utilize partial sub-views (i.e., under-utilization of the rich angular information), resulting in sub-optimal rain removal performance and no-equally good quality for all de-rained sub-views. In this article, we propose an efficient neural network, called MDeRainNet , for rain streak removal from LF images. The proposed network adopts a multi-scale encoder–decoder architecture, which directly works on Macro-pixel Images (MPIs) for improving the rain removal performance. To fully model the global correlation between the spatial information and the angular information, we propose an Extended Spatial-angular Interaction (ESAI) module to merge the two types of information, in which a simple and effective Transformer-based Spatial-angular Interaction Attention (SAIA) block is also proposed for modeling long-range geometric correlations and making full use of the angular information. Furthermore, to improve the generalization performance of our network on real-world rainy scenes, we propose a novel semi-supervised learning framework for our MDeRainNet , which utilizes multi-level KL loss to bridge the domain gap between features of synthetic and that of real-world rain streaks and introduces colored-residue image-guided contrastive regularization to reconstruct rain-free images. Extensive experiments conducted on both synthetic and real-world Light Field Images (LFIs) demonstrate that our method outperforms the state-of-the-art methods both quantitatively and qualitatively. Tao Yan 0001, Weilong Huang, Weijiang He, Cihang Wei, Xiangjie Zhu, Yinghui Wang 0001, Rynson W. H. Lau |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Multi-scale Attention and Adaptive Self-attention for Occlusion-Aware Depth Estimation in Light Field
Zhineng Zhang, Tao Yan 0001, Jinsheng Liu, Weilong Huang |
CGI (2) | 2 |
| 2025 | DDYOLO: Efficient Rainy Scene Object Detection with Deraining GuidanceabstractRecently, deep learning-based object detection has been widely used in autonomous driving and intelligent security. However, distinct features of targets would be contaminated due to domain shift and unexpected noise in adverse weather conditions, which may lead to poor performance in object detection. Several methods preprocess rain images with complex image deraining network before detection network to improve the detection performance, but fail to balance the time and detection accuracy. In this paper, we propose an end-to-end hybrid network called DDYOLO, which combines an image deraining sub-network and an object detection sub-network to utilize derained features for improving detection results with low time cost. Specifically, to ensure domain consistency between deraining and detection features while reducing redundant computations, both sub-networks share the same backbone network during features encoding. In the decoding process, we propose the Semantic Information Enhancement Block (SIEB) to supplement the missing semantic information layer by layer from the clean image features obtained by our proposed image deraining sub-network to the detection features. Moreover, we propose a rain layer extraction sub-network composed of the External Attention Module (EAM) to exploit the residual background information from degradation layers to assist background restoration and object detection. Extensive results demonstrate that our DDYOLO performs comparably with state-of-the-art methods while reducing inference time for each single-frame. Our code and dataset will be available at https://github.com/YT3DVision/DDYOLO. Tao Yan 0001, Jinsheng Liu, Zhineng Zhang, Weilong Huang |
IJCNN | 2 |
| 2025 | GCGP-YOLO: Global-Local and Channel Grouping Perception Network for Small Object Detection based on YOLOv8abstractDue to the wide field of view, small object sizes, and high object densities in Unmanned Aerial Vehicle (UAV) imagery, conventional object detection networks often struggle to effectively extract and perceive the features of small objects. This leads to misclassification and localization errors for small objects. To address these challenges, we propose GCGP-YOLO, a novel network based on YOLOv8n that improves small object detection performance for aerial image through Global-Local and Channel Grouping Perception. Specifically, we propose a Channel Grouping Perception Module (CGPM) to model the global context information while extracting local features. Subsequently, the implicit semantic relations of objects in the global contextual information are employed to guide the network to enhance the perception of small objects. Meanwhile, by combining large separable kernel convolutions, we propose Long and Short Dependency Pooling (LSDP), a method that captures long-range dependencies without introducing additional computational cost or parameters, thereby enriching the fine-grained information extracted by the network. In addition, we construct Linear Deformable Convolution Block (LDCB) to fit the diverse shapes of small objects, enabling the network to focus more on small objects during multi-scale feature fusion. Finally, the experimental results on the VisDrone2019 and TinyPerson datasets demonstrate that our network outperforms the baseline models with lower computational cost, achieving significant improvements of 8.30% and 5.68% in mAP50, respectively. The code and models are available at https://github.com/YT3DVision/GCGP-YOLO. Jinsheng Liu, Tao Yan 0001, Zhineng Zhang, Weilong Huang |
IJCNN | 2 |
| 2025 | Retinex-Guided Wavelet Diffusion for Low-Light Image Enhancement
Tao Yan 0001 |
PRCV (8) | 2 |
| 2025 | Intelligent Chinese Typesetting Model Based on Information Importance Can Enhance Text ReadabilityabstractIn today’s era of information overload, efficiently extracting valuable information from a large volume of textual data has become a crucial challenge in reading. This study aimed to explore the enhancement of Chinese readability in the digital domain through intelligent typesetting method. This study involved two experiments. The purpose of the first experiment was to achieve the machine learning-based assessment of the importance of individual words in Chinese articles and to automate typesetting based on the importance of words. For the second experiment, readability tests and eye-tracking reading tests were performed and the reading performance and reading attention between intelligent typesetting and general typesetting style was compared. This work proposed three Chinese typesetting methods that distinguished the importance of Chinese text information, based on font size and color. The results showed that first, when reading Chinese text, visual attention was more likely to be drawn to larger font sizes, darker brightness, or warmer-colored characters. Second, intelligent Chinese typography that distinguishes information importance through font size, color brightness, and color hue can enhance Chinese reading comprehension accuracy and subjective evaluation. It was concluded that using the TextRank model to distinguish importance of Chinese vocabulary and intelligent typesetting methods based on visual features of font could obviously improve text readability. Specifically, readability achieved through intelligent typesetting method distinguishing information importance through font color brightness surpasses that of general typesetting significantly. These intelligent typesetting methods can be widely applied in Chinese reading scenarios, such as web pages, e-books, information visualization, and other Chinese reading contexts. Tao Yan 0001, Ruimin Lyu, Yuan Liu 0021 |
Int. J. Hum. Comput. Interact. | 2 |
| 2025 | GhostingNet: A Novel Approach for Glass Surface Detection With Ghosting CuesabstractGhosting effects typically appear on glass surfaces, as each piece of glass has two contact surfaces causing two slightly offset layers of reflections. In this paper, we propose to take advantage of this intrinsic property of glass surfaces and apply it to glass surface detection, with two main technical novelties. First, we formulate a ghosting image formation model to describe the intensity and spatial relations among the main reflections and the background transmission within the glass region. Based on this model, we construct a new Glass Surface Ghosting Dataset (GSGD) to facilitate glass surface detection, with glass images and corresponding ghosting masks and glass surface masks. Second, we propose a novel method, called GhostingNet, for glass surface detection. Our method consists of a Ghosting Effects Detection (GED) module and a Glass Surface Detection (GSD) module. The key component of our GED module is a novel Double Reflection Estimation (DRE) block that models the spatial offsets of reflection layers for ghosting effect detection. The detected ghosting effects are then used to guide the GSD module for glass surface detection. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. We will release our code and dataset. Tao Yan 0001, Ke Xu 0010, Xiangjie Zhu, Helong Li, Benjamin W. Wah, Rynson W. H. Lau |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | NRGlassNet: Glass surface detection from visible and near-infrared image pairs
Tao Yan 0001, Shufan Xu, Helong Li, Xiaojun Chang, Rynson W. H. Lau |
Knowl. Based Syst. | 1 |
| 2024 | GLGFN: Global-Local Grafting Fusion Network for High-Resolution Image DerainingabstractImage deraining is a hot research topic, which aims to remove various rain streaks (raindrops) from rainy images and restore the backgrounds. Though image deraining has been extensively studied in recent years, few methods are able to effectively and efficiently derain real-world high-resolution rainy images. In general, existing image deraining methods are restricted by two main factors while processing high-resolution images. First, the computational complexity and memory usage of existing deep learning-based methods are high when it comes to derain high-resolution images. Second, as the image resolution increases, it is difficult to simultaneously extract and aggregate both global and local features for clean rain removal. In this paper, we propose a novel network, called Global-Local Grafting Fusion Network (GLGFN), for deraining real-world high-resolution images. Our GLGFN utilizes a staggered connection structure to achieve deeper sampling depth while maintaining low computational cost. It adopts the Transformer and CNN based encoders (backbones) to extract global and local features, respectively, and then grafts global features into local features to guide the extraction of rain streaks. In addition, for well fusing global and local features, we also propose a Grafting Fusion Module (GFM), which adopts Cross Sparse Attention (CSA) and Selective Kernel Fusion (SK Fusion) to efficiently aggregate global and local features. Extensive experiments conducted on several high-resolution real rainy datasets have demonstrated the effectiveness and efficiency of our proposed GLGFN. We will release our code and dataset. Tao Yan 0001, Xiangjie Zhu, Weijiang He, Yang Yang 0046, Yinghui Wang 0001, Xiaojun Chang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | YOWOv3: A Lightweight Spatio-Temporal Joint Network for Video Action DetectionabstractSpatio-temporal action detection networks, which need to simultaneously extract and fuse spatial and temporal features, often result in existing models becoming bloated and difficult to run in real-time and deploy on edge devices. This paper introduces an efficient and real-time spatio-temporal action detection model, YOWOv3. This model uses efficient 3D and 2D backbone networks to separately extract spatial and spatial-temporal features from sequential information. A lightweight spatio-temporal feature fusion module, designed by deeply integrating convolution and self-attention mechanisms, further enhances the extraction of spatio-temporal features. We refer to this module as the CFACM (Channel Fusion & Attention Convolution Mix) module. Our approach not only outperforms the latest efficient spatio-temporal action detection models in terms of lightness, reducing the model size by 24% compared to the latter, but also improves the mAP accuracy on the UCF101-24 dataset by 1.35%, while maintaining excellent speed performance, thus achieving a balance between accuracy and speed. Furthermore, existing models often use 3D convolutions to extract temporal information, which may be limited on certain devices, such as Apple’s M series processors. To mitigate the potential issue of 3D convolution operators not being supported during edge deployment of spatio-temporal action detection models, we employ a spatio-temporal shift module containing only 2D convolutions. This enables the model to acquire temporal information and inject the obtained temporal features into multi-level spatio-temporal feature extraction models. This not only liberates the model from the constraints of 3D convolution operations but also enhances the model’s balance between accuracy and speed. This results in state-of-the-art performance in lightweight networks using only 2D convolutions. Anlei Zhu, Yinghui Wang 0001, Jinlong Yang 0002, Tao Yan 0001, Haomiao Ma, Wei Li 0121 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Constructing Interpretable Belief Rule Bases Using a Model-Agnostic Statistical ApproachabstractBelief rule base (BRB) has attracted considerable interest due to its interpretability and exceptional modeling accuracy. Generally, BRB construction relies on prior knowledge or historical data. The limitations of knowledge constrain the knowledge-based BRB and are unsuitable for use in large-scale rule bases. Data-driven techniques excel at extracting model parameters from data, thus significantly improving the accuracy of BRB. However, the previous data-based BRBs neglected the study of interpretability, and some still depend on prior knowledge or introduce additional parameters. All these factors make the BRB highly problem-specific and limit its broad applicability. To address these problems, a model-agnostic statistical BRB (MAS-BRB) modeling approach is proposed in this article. It adopts an MAS methodology for parameter extraction, ensuring that the parameters both fulfill their intended roles within the BRB framework and accurately represent complex, nonlinear data relationships. A comprehensive interpretability analysis of MAS-BRB components further confirms their compliance with established BRB interpretability standards. Experiments conducted on multiple public datasets demonstrate that MAS-BRB not only achieves improved modeling performance but also shows greater effectiveness compared to existing rule-based and traditional machine learning models. Yinghui Wang 0001, Tao Yan 0001, Jinlong Yang 0002, Liangyi Huang |
IEEE Trans. Fuzzy Syst. | 3 |
| 2023 | Rain Removal From Light Field Images With 4D Convolution and Multi-Scale Gaussian ProcessabstractExisting deraining methods focus mainly on a single input image. However, with just a single input image, it is extremely difficult to accurately detect and remove rain streaks, in order to restore a rain-free image. In contrast, a light field image (LFI) embeds abundant 3D structure and texture information of the target scene by recording the direction and position of each incident ray via a plenoptic camera. LFIs are becoming popular in the computer vision and graphics communities. However, making full use of the abundant information available from LFIs, such as 2D array of sub-views and the disparity map of each sub-view, for effective rain removal is still a challenging problem. In this paper, we propose a novel method, 4D-MGP-SRRNet, for rain streak removal from LFIs. Our method takes as input all sub-views of a rainy LFI. To make full use of the LFI, it adopts 4D convolutional layers to simultaneously process all sub-views of the LFI. In the pipeline, the rain detection network, MGPDNet, with a novel Multi-scale Self-guided Gaussian Process (MSGP) module is proposed to detect high-resolution rain streaks from all sub-views of the input LFI at multi-scales. Semi-supervised learning is introduced for MSGP to accurately detect rain streaks by training on both virtual-world rainy LFIs and real-world rainy LFIs at multi-scales via computing pseudo ground truths for real-world rain streaks. We then feed all sub-views subtracting the predicted rain streaks into a 4D convolution-based Depth Estimation Residual Network (DERNet) to estimate the depth maps, which are later converted into fog maps. Finally, all sub-views concatenated with the corresponding rain streaks and fog maps are fed into a powerful rainy LFI restoring model based on the adversarial recurrent neural network to progressively eliminate rain streaks and recover the rain-free LFI. Extensive quantitative and qualitative evaluations conducted on both synthetic LFIs and real-world LFIs demonstrate the effectiveness of our proposed method. Tao Yan 0001, Yang Yang 0046, Rynson W. H. Lau |
IEEE Trans. Image Process. | 1 |
| 2023 | Image defocus deblurring method based on gradient difference of boundary neighborhoodabstractFor static scenes with multiple depth layers, the existing defocused image deblurring methods have the problems of edge ringing artifacts or insufficient deblurring degree due to inaccurate estimation of blur amount, In addition, the prior knowledge in non blind deconvolution is not strong, which leads to image detail recovery challenge. To this end, this paper proposes a blur map estimation method for defocused images based on the gradient difference of the boundary neighborhood, which uses the gradient difference of the boundary neighborhood to accurately obtain the amount of blurring, thus preventing boundary ringing artifacts. Then, the obtained blur map is used for blur detection to determine whether the image needs to be deblurred, thereby improving the efficiency of deblurring without manual intervention and judgment. Finally, a non blind deconvolution algorithm is designed to achieve image deblurring based on the blur amount selection strategy and sparse prior. Experimental results show that our method improves PSNR and SSIM by an average of 4.6% and 7.3%, respectively, compared to existing methods. Experimental results show that our method outperforms existing methods. Compared with existing methods, our method can better solve the problems of boundary ringing artifacts and detail information preservation in defocused image deblurring. Junjie Tao, Yinghui Wang 0001, Haomiao Ma, Tao Yan 0001, Lingyu Ai, Wei Li 0121 |
Virtual Real. Intell. Hardw. | 4 |
| 2022 | Constructing Calligraphy Evaluation Model Based on Writing Movement with LSTM NetworkabstractCalligraphy has a long history as one of the Chinese outstanding traditional arts. However, calligraphy, as an artistic derivative of Chinese characters, suffers from a lack of teachers, a variety of disciplines, and confusing aesthetic standards. With the development of calligraphy, the need for calligraphy evaluation has also gradually increased, but the traditional way to evaluate calligraphy works relies too much on the work of calligraphy and ignores the motion of writing. To address such problems, we construct a multimodal dataset that contains images of calligraphy works, time series data of writing movements and aesthetic evaluation labels. To exploit the time series data of writing movement, we propose a strong benchmark that combines the Long Short-Term Memory network with the K-nearest neighbor algorithm. The proposed model achieved the best accuracy compared with baseline methods. And the evaluation results of our model are close to that of calligraphy experts. This study serves as a guide to the aesthetic evaluation of computational calligraphy. Zhaoyi Wang, Ruimin Lyu, Xinya Liu, Yuefeng Ze, Zhenping Xie, Tao Yan 0001 |
CSCWD | 6 |
| 2022 | Rain Streak Removal From Light Field ImagesabstractRaining is a common weather condition, and may seriously degrade the performances of outdoor computer vision systems, such as surveillance and autonomous navigation. Rain streaks may exhibit diverse appearances in the captured images, depending on their distances from the camera. For example, sparse rain streaks near the camera lens may appear as continuous and translucent strips, while distant densely accumulated rain streaks are more like fog and mist. Existing rain removal methods are mainly based on a single input image. However, on a single image, it is difficult to estimate a reliable depth map for rain removal. A light field image (LFI) records abundant structural and texture information of the target scene by capturing multi-perspective sub-aperture views with a single exposure. With a LFI, it is easier to estimate the depth maps, and rain streak locations across sub-aperture views are highly correlated. We observe that rain streaks usually have different slops and/or chromaic values, compared with the background scene, along the epipolar plane images (EPIs) of an LFI. Thus, we propose to make use of 3D EPIs to detect rain streaks and restore the background. To this end, we propose a novel GAN architecture to remove rain streaks from an LFI. Our method takes as input a 3D EPI, i.e., a stacked of sub-aperture views along the same row of a rainy LFI. It first estimates the disparity maps for the 3D EPI by utilizing an auto-encoder based depth estimation sub-network. The disparity maps concatenated with the input sub-aperture views are then fed into a non-local residual block, and two branched autoencoder sub-networks are used to extract rain-streaks and recover rain-free sub-aperture views. Extensive experiments conducted on both synthetic real-world-like LFIs and real-world LFIs demonstrate the effectiveness of our method. Yuyang Ding, Tao Yan 0001, Fan Zhang 0063, Yuan Liu 0021, Rynson W. H. Lau |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Edge-Preserving Image Filtering Based on Soft ClusteringabstractEdge-preserving image filtering is an essential task in computational photography and imaging. In this paper, we propose a simple yet effective global edge-preserving filter based on soft clustering, and we propose a novel soft clustering algorithm based on a restricted Gaussian mixture model. Given specified parameters, the soft clustering process is firstly performed on the image to derive the partition matrix, from which the affinity matrix is then constructed for filtering. The filtering output is calculated as the weighted average of the pixels in the local window, so the proposed filter could suppress the intensity shift artifacts that impede most global filters. Besides, the weights in the proposed filter are derived by clustering, which properly separates dissimilar pixels, so the proposed filter could handle the halo artifacts that haunt many local filters. Moreover, our filter provides flexible control over the amount of smoothing that is deficient in the deep learning-based filters. Besides the efficacy in smoothing, the proposed filter naturally has low computational complexity. Qualitative and quantitative results suggest that the proposed filter benefits various applications, including edge-preserving smoothing, image enhancing, flash/non-flash fusion, HDR tone mapping, and dehazing. Yang Yang 0046, Hongjun Hui, Lanling Zeng, Yan Zhao 0038, Yongzhao Zhan 0001, Tao Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Generating Stereoscopic Images With Convergence Control Ability From a Light Field Image PairabstractWith the advances in commercial light field cameras, light field image processing has attracted considerable attention from researchers. In this paper, we propose a novel method for generating stereoscopic images from a light field image pair with flexible control over the convergence of the virtual stereo cameras. We have developed a light field image-capturing prototype that consists of two horizontally arranged light field cameras (i.e., with their optical axes being parallel to each other). When using our proposed device for image/video capture, stereo photographers can concentrate on how to capture the desired visual experience without being frequently disturbed by having to manipulate the stereo camera parameters, i.e., the convergence angle of a stereo camera. During postprocessing, our method estimates accurate disparity maps for the light field image pair and then generates the target stereoscopic images that satisfy the desired stereo camera convergence requirements by adopting a novel view synthesis method for light field images. We have conducted extensive experiments to demonstrate the effectiveness of our proposed method. Tao Yan 0001, Yiming Mao 0004, Wenxi Liu, Xiaohua Qian, Rynson W. H. Lau |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Stereoscopic Image Generation From Light Field With Disparity Scaling and Super-ResolutionabstractIn this paper, we propose a novel method to generate stereoscopic images from light-field images with the intended depth range and simultaneously perform image super-resolution. Subject to the small baseline of neighboring subaperture views and low spatial resolution of light-field images captured using compact commercial light-field cameras, the disparity range of any two subaperture views is usually very small. We propose a method to control the disparity range of the target stereoscopic images with linear or nonlinear disparity scaling and properly resolve the disocclusion problem with the aid of a smooth energy term previously used for texture synthesis. The left and right views of the target stereoscopic image are simultaneously generated by a unified optimization framework, which preserves content coherence between the left and right views by a coherence energy term. The disparity range of the target stereoscopic image can be larger than that of the input light field image. This benefits many light field image-based applications, e.g., displaying light field images on various stereo display devices and generating stereoscopic panoramic images from a light field image montage. An extensive experimental evaluation demonstrates the effectiveness of our method. Tao Yan 0001, Jianbo Jiao, Wenxi Liu, Rynson W. H. Lau |
IEEE Trans. Image Process. | 1 |
| 2019 | Deformable Object Tracking With Gated FusionabstractThe tracking-by-detection framework receives growing attention through the integration with the convolutional neural networks (CNNs). Existing tracking-by-detection-based methods, however, fail to track objects with severe appearance variations. This is because the traditional convolutional operation is performed on fixed grids, and thus may not be able to find the correct response while the object is changing pose or under varying environmental conditions. In this paper, we propose a deformable convolution layer to enrich the target appearance representations in the tracking-by-detection framework. We aim to capture the target appearance variations via deformable convolution, which adaptively enhances its original features. In addition, we also propose a gated fusion scheme to control how the variations captured by the deformable convolution affect the original appearance. The enriched feature representation through deformable convolution facilitates the discrimination of the CNN classifier on the target object and background. The extensive experiments on the standard benchmarks show that the proposed tracker performs favorably against the state-of-the-art methods. Wenxi Liu, Yibing Song, Dengsheng Chen, Shengfeng He, Yuanlong Yu 0001, Tao Yan 0001, Gerhard P. Hancke 0002, Rynson W. H. Lau |
IEEE Trans. Image Process. | 6 |
| 2017 | An ensemble classification algorithm for convolutional neural network based on AdaBoostabstractAdaBoost is a classic ensemble learning algorithm with good classifier performance. In the past, it mainly used weak classifier as base classifier, such as KNN. They are simple and easy to train, but the essence of the weak classifier, it is impossible to get very high classification accuracy. In order to improve the correct rate, this paper introduces the AdaBoost ensemble classifier based on convolutional neural network, namely adaBoost-CNN, referred to as ACNN. ACNN design a new training method, it not only gives the weight of base classifier according to the error rate of base classifier in pre-training phase, but also dynamically adjusts this weight and learning coefficient of training sample according to the error rate of each class in ensemble training phase. Finally, through experiments on some public datasets, it was proved that ACNN not only can effectively reduce the classification error rate, but also can solve the problem of class recognition rate imbalanced caused by similar categories or training samples quantity deviation. Li-Fang Chen, Tao Yan 0001, Yun-Hao Zhao, Ye-Jia Fan |
ICIS | 3 |
| 2013 | Consistent stereo image editingabstractStereo images and videos are very popular in recent years, and techniques for processing this media are attracting a lot of attention. In this paper, we extend the shift-map method for stereo image editing. Our method simultaneously processes the left and right images on pixel level using a global optimization algorithm. It enforces photo consistence between the two images and preserves 3D scene structures. It also addresses the occlusion and disocclusion problem, which may enable many stereo image editing functions, such as depth mapping, object depth adjustment and non-homogeneous image resizing. Our experiments show that the proposed method produces high quality results in various editing functions. Tao Yan 0001, Shengfeng He, Rynson W. H. Lau |
ACM Multimedia | 1 |
| 2013 | Seamless stitching of stereo images for generating infinite panoramasabstractA stereo infinite panorama is a panoramic image that may be infinitely extended by continuously stitching together stereo images that depict similar scenes, but are taken from different geographic locations. It can be used to create interesting walkthrough environment. An important issue underlying this application is to seamlessly stitch two stereo images together. Although many methods have been proposed for stitching 2D images, they may not work well on stereo images, due to the difficulty in ensuring disparity consistency. In this paper, we propose a novel method to stitch two stereo images seamlessly. We first apply the graph cut algorithm to compute a seam for stitching, with a novel disparity-aware energy function to both ensure disparity continuity and suppress visual artifacts around the seam. We then apply a modified warping-based disparity scaling algorithm to suppress the seam in depth domain. Experiments show that our stitching method is capable of producing high quality stereo infinite panoramas. Tao Yan 0001, Zhe Huang 0004, Rynson W. H. Lau |
VRST | 1 |
| 2013 | Depth Mapping for Stereoscopic VideosabstractStereoscopic videos have become very popular in recent years. Most of these videos are developed primarily for viewing on large screens located at some distance away from the viewer. If we watch these videos on a small screen located near to us, the depth range of the videos will be seriously reduced, which can significantly degrade the 3D effects of these videos. To address this problem, we propose a linear depth mapping method to adjust the depth range of a stereoscopic video according to the viewing configuration, including pixel density and distance to the screen. Our method tries to minimize the distortion of stereoscopic image contents after depth mapping, by preserving the relationship of neighboring features and preventing line and plane bending. It also considers the depth and motion coherences. While depth coherence ensures smooth changes of the depth field across frames, motion coherence ensures smooth content changes across frames. Our experimental results show that the proposed method can improve the stereoscopic effects while maintaining the quality of the output videos. Tao Yan 0001, Rynson W. H. Lau, Liusheng Huang |
Int. J. Comput. Vis. | 1 |