EDBT 2026 Demo / reviewers in the wild / expert
Qing Wang 0006
dblp:97/6505-6
· DBLP profile ↗
87ranked-venue papers
9as first author
33since 2021 · last 2026
0000-0003-3439-0644ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 43 · 4 first-author · 20 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalizable Occlusion-Aware Human Novel View Synthesis From Image Pairs
Kaijin Zhao, Xin Huang 0021, Guoqing Zhou 0003, Qing Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Zero-Pose-Prior NeRF: Recursive Radiance Field Reconstruction From Unposed and Unordered ImagesabstractThe dependence of neural radiance fields (NeRF) on accurate camera poses has emerged as a critical obstacle to their widespread real-world applications. While recent advances have demonstrated the potential for simultaneously addressing camera registration and scene reconstruction, these methods inherently rely on reasonable initialization derived from pose or scene priors and struggle with complex scenes involving large camera motions, particularly in unordered 360-degree scenes. In this work, we propose Zero-Pose-Prior NeRF to recover radiance fields from unposed and unordered image collections without any prior knowledge. Our key insight is to decompose this complex problem into smaller sub-problems, wherein the sub-problems' camera poses are initially estimated to provide self-bootstrapping priors for the global pose estimation, followed by a recursive registration and reconstruction. To achieve this, we first perform scene partitioning to establish a hierarchical structure that describes registration order from local to global. Thereafter, we devise a conditionally-decoupled positional encoding for NeRFs, which serves as the basic model for camera pose estimation and scene representation. Following this, we develop a recursive registration to recursively estimate the poses of local scenes and register them into a unified global pose space, ultimately enabling the reconstruction of the entire scene. Experiments on real-world scenes show that our approach outperforms the state-of-the-art pose-free methods in terms of accurate camera poses and robust radiance field reconstruction, resulting in high-fidelity view synthesis. Xinxin Liu 0020, Qi Zhang 0029, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006 |
IEEE Trans. Image Process. | 5 |
| 2025 | Bright-NeRF: Brightening Neural Radiance Field with Color Restoration from Low-Light RAW ImagesabstractNeural Radiance Fields (NeRF) have demonstrated prominent performance in novel view synthesis tasks. However, their input heavily relies on image acquisition under normal light conditions, making it challenging to learn accurate scene contents in low-light environments where images typically exhibit significant noise and severe color distortion. To address these challenges, we propose a novel approach, Bright-NeRF, which learns enhanced and high-quality radiance fields from multi-view low-light RAW images in an unsupervised manner. Our method simultaneously achieves color restoration, denoising, and enhanced novel view synthesis. Specifically, we leverage a physically-inspired model of the sensor's response to illumination and introduce a chromatic adaptation loss to constrain the learning of response, enabling consistent color perception of objects regardless of lighting conditions. We further utilize the RAW data's properties to expose the scene's intensity automatically. Additionally, we have collected a multi-view low-light RAW image dataset of real-world scenes to advance research in this field. Experimental results demonstrate that our proposed method significantly outperforms existing 2D and 3D approaches. Our code and dataset will be made publicly available. Xin Huang 0021, Guoqing Zhou 0003, Qifeng Guo, Qing Wang 0006 |
AAAI | 5 |
| 2025 | Material Anything: Generating Materials for Any 3D Object via DiffusionabstractWe present Material Anything, a fully-automated, unified diffusion framework designed to generate physically-based materials for 3D objects. Unlike existing methods that rely on complex pipelines or case-specific optimizations, Material Anything offers a robust, end-to-end solution adaptable to objects under diverse lighting conditions. Our approach leverages a pre-trained image diffusion model, enhanced with a triple-head architecture and rendering loss to improve stability and material quality. Additionally, we introduce confidence masks as a dynamic switcher within the diffusion model, enabling it to effectively handle both textured and texture-less objects across varying lighting conditions. By employing a progressive material generation strategy guided by these confidence masks, along with a UV-space material refiner, our method ensures consistent, UV-ready material outputs. Extensive experiments demonstrate our approach outperforms existing methods across a wide range of object categories and lighting conditions. Xin Huang 0021, Tengfei Wang 0002, Ziwei Liu 0002, Qing Wang 0006 |
CVPR | 4 |
| 2025 | Learning motion-guided salience features for weakly supervised group activity recognition
Zexing Du, Qing Wang 0006 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Phase shift guided dynamic view synthesis from monocular video
Chuyue Zhao, Xin Huang 0021, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006 |
Image Vis. Comput. | 5 |
| 2025 | Generalizable 3D Gaussian Splatting for novel view synthesis
Chuyue Zhao, Xin Huang 0021, Xue Wang 0006, Qing Wang 0006 |
Pattern Recognit. | 5 |
| 2025 | ${\rm{H}}_{2}{\rm{O}}$H2O-NeRF: Radiance Fields Reconstruction for Two-Hand-Held ObjectsabstractOur work aims to reconstruct the appearance and geometry of the two-hand-held object from a sequence of color images. In contrast to traditional single-hand-held manipulation, two-hand-holding allows more flexible interaction, thereby providing back views of the object, which is particularly convenient for reconstruction but generates complex view-dependent occlusions. The recent development of neural rendering provides new potential for hand-held object reconstruction. In this paper, we propose a novel neural representation-based framework to recover radiance fields of the two-hand-held object, named ${\rm{H}}_{2}{\rm{O}}$H2O-NeRF. We first design an object-centric semantic module based on the geometric signed distance function cues to predict 3D object-centric regions and develop the view-dependent visible module based on the image-related cues to label 2D occluded regions. We then combine them to obtain a 2D visible mask that adaptively guides ray sampling on the object for optimization. We also provide a newly collected ${\rm{H}}_{2}{\rm{O}}$H2O dataset to validate the proposed method. Experiments show that our method achieves superior performance on reconstruction completeness and view-consistency synthesis compared to the state-of-the-art methods. Xinxin Liu 0020, Qi Zhang 0029, Xin Huang 0021, Guoqing Zhou 0003, Qing Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human GenerationabstractRecent text-to-3D methods employing diffusion models have made significant advancements in 3D human generation. However, these approaches face challenges due to the limitations of text-to-image diffusion models, which lack an understanding of 3D structures. Consequently, these methods struggle to achieve high-quality human generation, resulting in smooth geometry and cartoon-like appearances. In this paper, we propose HumanNorm, a novel approach for high-quality and realistic 3D human generation. The main idea is to enhance the model's 2D perception of 3D geometry by learning a normal-adapted diffusion model and a normal-aligned diffusion model. The normal-adapted diffusion model can generate high-fidelity normal maps corresponding to user prompts with view-dependent and body-aware text. The normal-aligned diffusion model learns to generate color images aligned with the normal maps, thereby transforming physical geometry details into realistic appearance. Leveraging the proposed normal diffusion model, we devise a progressive geometry generation strategy and a multi-step Score Distillation Sampling (SDS) loss to enhance the performance of 3D human generation. Comprehensive experiments substantiate HumanNorm's ability to generate 3D humans with intricate geometry and realistic appearances. HumanNorm outperforms existing text-to-3D methods in both geometry and texture quality. The project page of HumanNorm is https://humannorm.github.io/. Xin Huang 0021, Ruizhi Shao, Qi Zhang 0029, Hongwen Zhang 0001, Yebin Liu, Qing Wang 0006 |
CVPR | 7 |
| 2024 | Dual-Scale Temporal Dependency Learning for Unsupervised Video Anomaly Detection
Xue Wang 0006, Zexing Du, Qing Wang 0006 |
PRCV (10) | 4 |
| 2024 | A two-stage substation equipment classification method based on dual-scale attentionabstractAbstract Accurate classification of substation equipment images remains challenging due to various factors such as unexpected illumination, viewing angles, scale variations, shadows, surface contaminants, and different elements sharing similar appearances. This paper presents a novel two‐stage substation equipment classification method based on dual‐scale attention. Leveraging the region proposal technique from Faster‐regions with CNN features (RCNN), the input images are initially decomposed into multiple scales to capture latent features. A dual‐scale attention module is introduced to enhance the precision of feature extraction. Furthermore, a two‐stage network is proposed to address the challenge of classifying closely similar substation equipment. A multi‐layer perceptron performs a coarse classification to categorize the equipment into broad categories. Then, a lightweight classifier is employed for fine‐grained subclassification, further distinguishing equipment within the same broad category. To mitigate the issue of limited training data, a specialized dataset is collected and annotated for the substation equipment classification. Experimental results demonstrate that the proposed method achieves remarkable accuracy, recall, and F1‐score surpassing 0.91, outperforming mainstream approaches in terms of recall and F1 scores. Ablation experiments further validate the significant contributions of both the dual‐scale attention and the two‐stage classification module in improving the overall performance of the classification network. Yiyang Yao, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006 |
IET Image Process. | 4 |
| 2024 | Spatially-Varying Illumination-Aware Indoor Harmonization
Zhongyun Hu, Xue Wang 0006, Qing Wang 0006 |
Int. J. Comput. Vis. | 4 |
| 2024 | LTM-NeRF: Embedding 3D Local Tone Mapping in HDR Neural Radiance FieldabstractRecent advances in Neural Radiance Fields (NeRF) have provided a new geometric primitive for novel view synthesis. High Dynamic Range NeRF (HDR NeRF) can render novel views with a higher dynamic range. However, effectively displaying the scene contents of HDR NeRF on diverse devices with limited dynamic range poses a significant challenge. To address this, we present LTM-NeRF, a method designed to recover HDR NeRF and support 3D local tone mapping. LTM-NeRF allows for the synthesis of HDR views, tone-mapped views, and LDR views under different exposure settings, using only the multi-view multi-exposure LDR inputs for supervision. Specifically, we propose a differentiable Camera Response Function (CRF) module for HDR NeRF reconstruction, globally mapping the scene's HDR radiance to LDR pixels. Moreover, we introduce a Neural Exposure Field (NeEF) to represent the spatially varying exposure time of an HDR NeRF to achieve 3D local tone mapping, for compatibility with various displays. Comprehensive experiments demonstrate that our method can not only synthesize HDR views and exposure-varying LDR views accurately but also render locally tone-mapped views naturally. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Sheared Epipolar Focus Spectrum for Dense Light Field ReconstructionabstractThis paper presents a novel technique for the dense reconstruction of light fields (LFs) from sparse input views. Our approach leverages the Epipolar Focus Spectrum (EFS) representation, which models the LF in the transformed spatial-focus domain, avoiding the dependence on the scene depth and providing a high-quality basis for dense LF reconstruction. Previous EFS-based LF reconstruction methods learn the cross-view, occlusion, depth and shearing terms simultaneously, which makes the training difficult due to stability and convergence problems and further results in limited reconstruction performance for challenging scenarios. To address this issue, we conduct a theoretical study on the transformation between the EFSs derived from one LF with sparse and dense angular samplings, and propose that a dense EFS can be decomposed into a linear combination of the EFS of the sparse input, the sheared EFS, and a high-order occlusion term explicitly. The devised learning-based framework with the input of the under-sampled EFS and its sheared version provides high-quality reconstruction results, especially in large disparity areas. Comprehensive experimental evaluations show that our approach outperforms state-of-the-art methods, especially achieves at most dB advantages in reconstructing scenes containing thin structures. Xue Wang 0006, Guoqing Zhou 0003, Hao Zhu 0005, Qing Wang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Learning Semantics-Guided Representations for Scoring Figure SkatingabstractThis paper explores semantic-aware representations for scoring figure skating videos. Most existing approaches to sports video analysis only focus on reasoning action scores based on visual input, limiting their ability to depict high-level semantic representations. Here, we propose a teacher-student-based network with an attention mechanism to realize an adaptive knowledge transfer from the semantic domain to the visual domain, which is termed semantics-guided network (SGN). Specifically, we use a set of learnable atomic queries in the student branch to mimic the semantic-aware distribution in the teacher branch, which is represented by the visual and semantic inputs. In addition, we propose three auxiliary losses to align features in different domains. With aligned feature representations, the adapted teacher is capable of transferring the semantic knowledge to the student. To verify the effectiveness of our method, we collect a new dataset OlympicFS for scoring figure skating. Besides action scores, OlympicFS also provides professional comments on actions for learning semantic representations. By evaluating four challenging datasets, our method achieves state-of-the-art performance. Zexing Du, Di He 0010, Xue Wang 0006, Qing Wang 0006 |
IEEE Trans. Multim. | 4 |
| 2023 | Inverting the Imaging Process by Learning an Implicit Camera ModelabstractRepresenting visual signals with implicit coordinate-based neural networks, as an effective replacement of the traditional discrete signal representation, has gained considerable popularity in computer vision and graphics. In contrast to existing implicit neural representations which focus on modelling the scene only, this paper proposes a novel implicit camera model which represents the physical imaging process of a camera as a deep neural network. We demonstrate the power of this new implicit camera model on two inverse imaging tasks: i) generating all-in-focus photos, and ii) HDR imaging. Specifically, we devise an implicit blur generator and an implicit tone mapper to model the aperture and exposure of the camera's imaging process, respectively. Our implicit camera model is jointly learned together with implicit scene models under multi-focus stack and multi-exposure bracket supervision. We have demonstrated the effectiveness of our new model on a large number of test images and videos, producing accurate and visually appealing all-in-focus and high dynamic range images. In principle, our new implicit neural camera model has the potential to benefit a wide array of other inverse imaging tasks. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
CVPR | 5 |
| 2023 | Local Implicit Ray Function for Generalizable Radiance Field RepresentationabstractWe propose LIRF (Local Implicit Ray Function), a generalizable neural rendering approach for novel view rendering. Current generalizable neural radiance fields (NeRF) methods sample a scene with a single ray per pixel and may therefore render blurred or aliased views when the input views and rendered views capture scene content with different resolutions. To solve this problem, we propose LIRF to aggregate the information from conical frustums to construct a ray. Given 3D positions within conical frustums, LIRF takes 3D coordinates and the features of conical frustums as inputs and predicts a local volumetric radiance field. Since the coordinates are continuous, LIRF renders high-quality novel views at a continuously-valued scale via volume rendering. Besides, we predict the visible weights for each input view via transformer-based feature matching to improve the performance in occluded areas. Experimental results on real-world scenes validate that our method outperforms state-of-the-art methods on novel view rendering of unseen scenes at arbitrary scales. Xin Huang 0021, Qi Zhang 0029, Xiaoyu Li 0002, Xuan Wang 0009, Qing Wang 0006 |
CVPR | 6 |
| 2023 | Wide-Angle Rectification via Content-Aware Conformal MappingabstractDespite the proliferation of ultra wide-angle lenses on smartphone cameras, such lenses often come with severe image distortion (e.g. curved linear structure, unnaturally skewed faces). Most existing rectification methods adopt a global warping transformation to undistort the input wideangle image, yet their performances are not entirely satisfactory, leaving many unwanted residue distortions uncorrected or at the sacrifice of the intended wide FoV (field- of-view). This paper proposes a new method to tackle these challenges. Specifically, we derive a locally-adaptive polardomain conformal mapping to rectify a wide-angle image. Parameters of the mapping are found automatically by analyzing image contents via deep neural networks. Experiments on a large number of photos have confirmed the superior performance of the proposed method compared with all available previous methods. Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
CVPR | 3 |
| 2023 | InterFormer: Human Interaction Understanding with Deformed Transformer
Di He 0010, Zexing Du, Xue Wang 0006, Qing Wang 0006 |
ICIC (5) | 4 |
| 2023 | Perceiving local relative motion and global correlations for weakly supervised group activity recognition
Zexing Du, Xue Wang 0006, Qing Wang 0006 |
Image Vis. Comput. | 3 |
| 2023 | Dilated Transformer with Feature Aggregation Module for Action Segmentation
Zexing Du, Qing Wang 0006 |
Neural Process. Lett. | 2 |
| 2023 | Dense light field reconstruction based on epipolar focus spectrumabstractExisting light field (LF) representations, such as epipolar plane image (EPI) and sub-aperture images, do not consider the structural characteristics across the views, so they usually require additional disparity and spatial structure cues for follow-up tasks. Besides, they have difficulties dealing with occlusions or large disparity scenes. To this end, this paper proposes a novel Epipolar Focus Spectrum (EFS) representation by rearranging the EPI spectrum. Different from the classical EPI representation where an EPI line corresponds to a specific depth, there is a one-to-one mapping from the EFS line to the view. By exploring the EFS sampling task, the analytical function is derived for constructing a non-aliasing EFS. To demonstrate its effectiveness, we develop a trainable EFS-based pipeline for light field reconstruction, where a dense light field can be reconstructed by compensating the missing EFS lines given a sparse light field, yielding promising results with cross-view consistency, especially in the presence of severe occlusion and large disparity. Experimental results on both synthetic and real-world datasets demonstrate the validity and superiority of the proposed method over SOTA methods. Xue Wang 0006, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006 |
Pattern Recognit. | 5 |
| 2023 | Self-Supervised Global Spatio-Temporal Interaction Pre-Training for Group Activity RecognitionabstractThis paper focuses on exploring distinctive spatio-temporal representation in a self-supervised manner for group activity recognition. Firstly, previous networks treat spatial- and temporal-aware information as a whole, limiting their abilities to represent complex spatio-temporal correlations for group activity. Here, we propose the Spatial and Temporal Attention Heads (STAHs) to extract spatial- and temporal-aware representations independently, which generate complementary contexts for boosting group activity understanding. Then, we propose the Global Spatio-Temporal Contrastive (GSTCo) loss to aggregate these two kinds of features. Unlike previous works focusing on the individual temporal consistency while overlooking the correlations between actors, i.e., in a local perspective, we explore the global spatial and temporal dependency. Moreover, GSTCo could effectively avoid the trivial solution faced in contrastive learning by achieving the right balance between spatial and temporal representations. Furthermore, our method imports affordable overhead during pre-training, without additional parameters or computational costs in inference, guaranteeing efficiency. By evaluating on widely-used datasets for group activity recognition, our method achieves good performance. State-of-the-art performance is achieved when applying our pre-trained backbone to existing networks. Extensive experiments verify the generalizability of our method. Zexing Du, Xue Wang 0006, Qing Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Learning Reliable Gradients From Undersampled Circular Light Field for 3D ReconstructionabstractThe paper presents a 3D reconstruction algorithm from an undersampled circular light field (LF). With an ultra-dense angular sampling rate, every scene point captured by a circular LF corresponds to a smooth trajectory in the circular epipolar plane volume (CEPV). Thus per-pixel disparities can be calculated by retrieving the local gradients of the CEPV-trajectories. However, the continuous curve will be broken up into discrete segments in an undersampled circular LF, which leads to a noticeable deterioration of the 3D reconstruction accuracy. We observe that the coherent structure is still embedded in the discrete segments. With less noise and ambiguity, the scene points can be reconstructed using gradients from reliable epipolar plane image (EPI) regions. By analyzing the geometric characteristics of the coherent structure in the CEPV, both the trajectory itself and its gradients could be modeled as 3D predictable series. Thus a mask-guided CNN+LSTM network is proposed to learn the mapping from the CEPV with a lower angular sampling rate to the gradients under a higher angular sampling rate. To segment the reliable regions, the reliable-mask-based loss that assesses the difference between learned gradients and ground truth gradients is added to the loss function. We construct a synthetic circular LF dataset with ground truth for depth and foreground/background segmentation to train the network. Moreover, a real-scene circular LF dataset is collected for performance evaluation. Experimental results on both public and self-constructed datasets demonstrate the superiority of the proposed method over existing state-of-the-art methods. Zhengxi Song, Xue Wang 0006, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Fast and Unsupervised Action Boundary Detection for Action SegmentationabstractTo deal with the great number of untrimmed videos produced every day, we propose an efficient unsupervised action segmentation method by detecting boundaries, named action boundary detection (ABD). In particular, the proposed method has the following advantages: no training stage and low-latency inference. To detect action boundaries, we estimate the similarities across smoothed frames, which inherently have the properties of internal consistency within actions and external discrepancy across actions. Under this circumstance, we successfully transfer the boundary detection task into the change point detection based on the similarity. Then, non-maximum suppression (NMS) is conducted in local windows to select the smallest points as candidate boundaries. In addition, a clustering algorithm is followed to refine the initial proposals. Moreover, we also extend ABD to the online setting, which enables real-time action segmentation in long untrimmed videos. By evaluating on four challenging datasets, our method achieves state-of-the-art performance. Moreover, thanks to the efficiency of ABD, we achieve the best trade-off between the accuracy and the inference time compared with existing unsupervised approaches. Zexing Du, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006 |
CVPR | 4 |
| 2022 | HDR-NeRF: High Dynamic Range Neural Radiance FieldsabstractWe present High Dynamic Range Neural Radiance Fields (HDR-NeRF) to recover an HDR radiance field from a set of low dynamic range (LDR) views with different exposures. Using the HDR-NeRF, we are able to generate both novel HDR views and novel LDR views under different exposures. The key to our method is to model the simplified physical imaging process, which dictates that the radiance of a scene point transforms to a pixel value in the LDR image with two implicit functions: a radiance field and a tone mapper. The radiance field encodes the scene radiance (values vary from 0 to$+\infty$), which outputs the density and radiance of a ray by giving corresponding ray origin and ray direction. The tone mapper models the mapping process that a ray hitting on the camera sensor becomes a pixel value. The color of the ray is predicted by feeding the radiance and the corresponding exposure time into the tone mapper. We use the classic volume rendering technique to project the output radiance, colors and densities into HDR and LDR images, while only the input LDR images are used as the supervision. We collect a new forward-facing HDR dataset to evaluate the proposed method. Experimental results on synthetic and real-world scenes validate that our method can not only accurately control the exposures of synthesized views but also render views with a high dynamic range. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Xuan Wang 0009, Qing Wang 0006 |
CVPR | 6 |
| 2022 | Ray-Space Epipolar Geometry for Light Field CamerasabstractLight field essentially represents rays in space. The epipolar geometry between two light fields is an important relationship that captures ray-ray correspondences and relative configuration of two views. Unfortunately, so far little work has been done in deriving a formal epipolar geometry model that is specifically tailored for light field cameras. This is primarily due to the high-dimensional nature of the ray sampling process with a light field camera. This paper fills in this gap by developing a novel ray-space epipolar geometry which intrinsically encapsulates the complete projective relationship between two light fields, while the generalized epipolar geometry which describes relationship of normalized light fields is the specialization of the proposed model to calibrated cameras. With Plücker parameterization, we propose the ray-space projection model involving a 6×6 ray-space intrinsic matrix for ray sampling of light field camera. Ray-space fundamental matrix and its properties are then derived to constrain ray-ray correspondences for general and special motions. Finally, based on ray-space epipolar geometry, we present two novel algorithms, one for fundamental matrix estimation, and the other for calibration. Experiments on synthetic and real data have validated the effectiveness of ray-space epipolar geometry in solving 3D computer vision tasks with light field cameras. Qi Zhang 0029, Qing Wang 0006, Hongdong Li, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | PNRNet: Physically-Inspired Neural Rendering for Any-to-Any RelightingabstractExisting any-to-any relighting methods suffer from the task-aliasing effects and the loss of local details in the image generation process, such as shading and attached-shadow. In this paper, we present PNRNet, a novel neural architecture that decomposes the any-to-any relighting task into three simpler sub-tasks, i.e. lighting estimation, color temperature transfer, and lighting direction transfer, to avoid the task-aliasing effects. These sub-tasks are easy to learn and can be trained with direct supervisions independently. To better preserve local shading and attached-shadow details, we propose a parallel multi-scale network that incorporates multiple physical attributes to model local illuminations for lighting direction transfer. We also introduce a simple yet effective color temperature transfer network to learn a pixel-level non-linear function which allows color temperature adjustment beyond the predefined color temperatures and generalizes well to real images. Extensive experiments demonstrate that our proposed approach achieves better results quantitatively and qualitatively than prior works. Zhongyun Hu, Ntumba Elie Nsampi, Xue Wang 0006, Qing Wang 0006 |
IEEE Trans. Image Process. | 4 |
| 2021 | Learning Exposure Correction Via Consistency Modeling
Ntumba Elie Nsampi, Zhongyun Hu, Qing Wang 0006 |
BMVC | 3 |
| 2021 | 3D Scene Reconstruction with an Un-calibrated Light Field Camera
Qi Zhang 0029, Hongdong Li, Xue Wang 0006, Qing Wang 0006 |
Int. J. Comput. Vis. | 4 |
| 2021 | Region-based depth feature descriptor for saliency detection on light field
Xue Wang 0006, Yingying Dong, Qi Zhang 0029, Qing Wang 0006 |
Multim. Tools Appl. | 4 |
| 2021 | 4D Light Field Segmentation From Light Field Super-Pixel Hypergraph RepresentationabstractEfficient and accurate segmentation of full 4D light fields is an important task in computer vision and computer graphics. The massive volume and the redundancy of light fields make it an open challenge. In this article, we propose a novel light field hypergraph (LFHG) representation using the light field super-pixel (LFSP) for interactive light field segmentation. The LFSPs not only maintain the light field spatio-angular consistency, but also greatly contribute to the hypergraph coarsening. These advantages make LFSPs useful to improve segmentation performance. Based on the LFHG representation, we present an efficient light field segmentation algorithm via graph-cut optimization. Experimental results on both synthetic and real scene data demonstrate that our method outperforms state-of-the-art methods on the light field segmentation task with respect to both accuracy and efficiency. Xianqiang Lv, Xue Wang 0006, Qing Wang 0006, Jingyi Yu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Revisiting Spatio-Angular Trade-off in Light Field Cameras and Extended Applications in Super-ResolutionabstractLight field cameras (LFCs) have received increasing attention due to their wide-spread applications. However, current LFCs suffer from the well-known spatio-angular trade-off, which is considered an inherent and fundamental limit for LFC designs. In this article, by doing a detailed optical analysis of the sampling process in an LFC, we show that the effective resolution is generally higher than the number of micro-lenses. This contribution makes it theoretically possible to super-resolve a light field. Further optical analysis proves the "2D predictable series" nature of the 4D light field, which provides new insights for analyzing light field using series processing techniques. To model this nature, a specifically designed epipolar plane image (EPI) based CNN-LSTM network is proposed to super-resolve a light field in the spatial and angular dimensions simultaneously. Rather than leveraging semantic information, our network focuses on extracting geometric continuity in the EPI domain. This gives our method an improved generalization ability and makes it applicable to a wide range of previously unseen scenes. Experiments on both synthetic and real light fields demonstrate the improvements over state-of-the-arts, especially in large disparity areas. Hao Zhu 0005, Mantang Guo, Hongdong Li, Qing Wang 0006, Antonio Robles-Kelly |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | DGAN: Disentangled Representation Learning for Anisotropic BRDF ReconstructionabstractAccurate reconstruction of real-world materials' appearance from a very limited number of samples is still a huge challenge in computer vision and graphics. In this paper, we present a novel deep architecture, Disentangled Generative Adversarial Network (DGAN), which performs anisotropic Bidirectional Reflectance Distribution Function (BRDF) reconstruction from single BRDF subspace with the maximum entropy. In contrast to previous approaches that directly map known samples to a full BRDF using a CNN, a disentangled representation learning is applied to guide the reconstruction process. In order to learn different physical factors of the BRDF, the generator of the DGAN mainly consists of a fresnel estimator module (FEM) and a directional module (DM). Considering the fact that the entropy of different BRDF subspace varies, we further divide the BRDF into He-BRDF and Le-BRDF to reconstruct the interior part and the exterior part of the directional factor. Experimental results show that our approach outperforms state-of-the-art methods. Zhongyun Hu, Xue Wang 0006, Qing Wang 0006 |
ICASSP | 3 |
| 2020 | Accurate 3D Reconstruction from Circular Light Field Using CNN-LSTMabstractA light field is formed by densely capturing images on a regular sub-aperture grid. Geometry information endowed in the epipolar plane images(EPI) can only lead to a 2. 5D reconstruction. In order to obtain a full 360°view of an object, we focus on light fields captured by a circularly moving camera, resulting in circular light fields (or Cir-LFs in short). Compared with traditional EPIs, Circular EPIs(CEPIs) provide unique advantages, such as that corresponding points forming a 3D sinusoid like curve instead of a 2D straight line and geometry information encoded sequentially in multiple adjacent views along the curve. However, current reconstruction methods only focus on the 2D projection of 3D curve, leading to distortions in the reconstructed upper and lower surfaces. We propose to analyze 3D features contained in the 3D CEPI volume and we develop a deep CNN-LSTM network to model the gradient map in the CEPI volume. Additionally, a large scale Cir-LF dataset is constructed for research purpose. Experiments on both synthetic and real scenes demonstrate the effectiveness and generaliability of the proposed method. Zhengxi Song, Hao Zhu 0005, Xue Wang 0006, Hongdong Li, Qing Wang 0006 |
ICME | 6 |
| 2020 | 4D Light Field Superpixel and SegmentationabstractSuperpixel segmentation of 2D images has been widely used in many computer vision tasks. Previous algorithms model the color, position, or higher spectral information for segmenting a 2D image. However, limited to the Gaussian imaging principle in a traditional camera, where each pixel is formed by summing lots of light rays from different angles, there is not a thorough segmentation solution to eliminate the ambiguity in defocus and occlusion boundary areas. In this paper, we consider the essential element of image pixel, i.e., rays in light space, and propose light field superpixel (LFSP) to eliminate the ambiguity. The LFSP is first defined mathematically and then two evaluation metrics, named LFSP self-similarity and effective label ratio, are proposed to evaluate the refocus-invariant and full-sliced properties of segmentation. By building a clique system containing 80 neighbors in light field, a robust refocus-invariant LFSP segmentation algorithm is developed. Experimental results on both synthetic and real light field datasets demonstrate the advantages over the current state of the art in terms of traditional evaluation metrics. Additionally, the LFSP self-similarity evaluations under different light field refocus levels show the refocus-invariance of the proposed algorithm. The full-sliced property of the proposed LFSP algorithm is verified by comparing it with the classical supervoxel algorithms. Finally, an LFSP-based application is demonstrated to show the effectiveness of LFSP in light field editing. Hao Zhu 0005, Qi Zhang 0029, Qing Wang 0006, Hongdong Li |
IEEE Trans. Image Process. | 3 |
| 2019 | Ray-Space Projection Model for Light Field CameraabstractLight field essentially represents the collection of rays in space. The rays captured by multiple light field cameras form subsets of full rays in 3D space and can be transformed to each other. However, most previous approaches model the projection from an arbitrary point in 3D space to corresponding pixel on the sensor. There are few models on describing the ray sampling and transformation among multiple light field cameras. In the paper, we propose a novel ray-space projection model to transform sets of rays captured by multiple light field cameras in term of the Plucker coordinates. We first derive a 6×6 ray-space intrinsic matrix based on multi-projection-center (MPC) model. A homogeneous ray-space projection matrix and a fundamental matrix are then proposed to establish ray-ray correspondences among multiple light fields. Finally, based on the ray-space projection matrix, a novel camera calibration method is proposed to verify the proposed model. A linear constraint and a ray-ray cost function are established for linear initial solution and non-linear optimization respectively. Experimental results on both synthetic and real light field data have verified the effectiveness and robustness of the proposed model. Qi Zhang 0029, Jinbo Ling, Qing Wang 0006, Jingyi Yu 0001 |
CVPR | 3 |
| 2019 | A Generic Multi-Projection-Center Model and Calibration Method for Light Field CamerasabstractLight field cameras can capture both spatial and angular information of light rays, enabling 3D reconstruction by a single exposure. The geometry of 3D reconstruction is affected by intrinsic parameters of a light field camera significantly. In the paper, we propose a multi-projection-center (MPC) model with 6 intrinsic parameters to characterize light field cameras based on traditional two-parallel-plane (TPP) representation. The MPC model can generally parameterize light field in different imaging formations, including conventional and focused light field cameras. By the constraints of 4D ray and 3D geometry, a 3D projective transformation is deduced to describe the relationship between geometric structure and the MPC coordinates. Based on the MPC model and projective transformation, we propose a calibration algorithm to verify our light field camera model. Our calibration method includes a close-form solution and a non-linear optimization by minimizing re-projection errors. Experimental results on both simulated and real scene data have verified the performance of our algorithm. Qi Zhang 0029, Chunping Zhang, Jinbo Ling, Qing Wang 0006, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Dense Light Field Reconstruction from Sparse Sampling Using Residual Network
Mantang Guo, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006 |
ACCV (6) | 4 |
| 2018 | Common Self-polar Triangle of Concentric Conics for Light Field Camera Calibration
Qi Zhang 0029, Qing Wang 0006 |
ACCV (6) | 2 |
| 2017 | 4D Light Field Superpixel and SegmentationabstractSuperpixel segmentation of 2D image has been widely used in many computer vision tasks. However, limited to the Gaussian imaging principle, there is not a thorough segmentation solution to the ambiguity in defocus and occlusion boundary areas. In this paper, we consider the essential element of image pixel, i.e., rays in the light space and propose light field superpixel (LFSP) segmentation to eliminate the ambiguity. The LFSP is first defined mathematically and then a refocus-invariant metric named LFSP self-similarity is proposed to evaluate the segmentation performance. By building a clique system containing 80 neighbors in light field, a robust refocus-invariant LFSP segmentation algorithm is developed. Experimental results on both synthetic and real light field datasets demonstrate the advantages over the state-of-the-arts in terms of traditional evaluation metrics. Additionally the LFSP self-similarity evaluation under different light field refocus levels shows the refocus-invariance of the proposed algorithm. Hao Zhu 0005, Qi Zhang 0029, Qing Wang 0006 |
CVPR | 3 |
| 2017 | High angular resolution light field reconstruction with coded-aperture maskabstractIn the past decade, light field imaging has greatly extended the imaging capabilities of traditional photography. However, the applications of light field imaging are limited by the aliasing artifacts due to the plenoptic sampling trade-off between angular and spatial domains. We propose to use a coded aperture light field camera instead of the traditional one, which can get more angular information without losing spatial resolution. To that end, we exploit a theoretical model to explain the relationship between light field and the raw data captured by the sensor. Then, we design a mask to code the rays using compressive sensing. Last, the sparse characteristic of light field in gradient domain and the corresponding optimization methods are utilized to reconstruct the high angular resolution light field. Experimental results on synthetic data and real data demonstrate that our system can obtain high angular resolution light field by producing a low-aliasing refocused image and high PSNR multi-view images. Wanxin Qu, Guoqing Zhou 0003, Hao Zhu 0005, Zhaolin Xiao, Qing Wang 0006, René Vidal |
ICIP | 5 |
| 2017 | Extending the FOV from disparity and color consistencies in multiview light fieldsabstractLight field, which is captured by a plenoptic camera, is always limited in its narrow field of view (FOV) by the physical size of the aperture. To break through the restriction, we propose to extend the FOV using multiview light fields. A series of light fields are acquired by translating the camera at isometric spatial positions. In contrast to previous methods, our algorithm is the first that achieves light field registration and rendering based on epipolar plane image (EPI) properties, including disparity and color consistencies. Furthermore, the aliasing caused by the under-sampling in the angular space is eliminated by synthesizing novel views in the EPI space. Experimental results on the real scene data have demonstrated the effectiveness of our algorithm. Zhao Ren, Qi Zhang 0029, Hao Zhu 0005, Qing Wang 0006 |
ICIP | 4 |
| 2017 | Robust outlier removal using penalized linear regression in multiview geometry
Guoqing Zhou 0003, Qing Wang 0006, Zhaolin Xiao |
Neurocomputing | 2 |
| 2017 | Light field imaging: models, calibrations, reconstructions, and applicationsabstractLight field imaging is an emerging technology in computational photography areas. Based on innovative designs of the imaging model and the optical path, light field cameras not only record the spatial intensity of threedimensional (3D) objects, but also capture the angular information of the physical world, which provides new ways to address various problems in computer vision, such as 3D reconstruction, saliency detection, and object recognition. In this paper, three key aspects of light field cameras, i.e., model, calibration, and reconstruction, are reviewed extensively. Furthermore, light field based applications on informatics, physics, medicine, and biology are exhibited. Finally, open issues in light field imaging and long-term application prospects in other natural sciences are discussed. Hao Zhu 0005, Qing Wang 0006, Jingyi Yu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2017 | Motion-Based Temporal Alignment of Independently Moving CamerasabstractThis paper presents a method to establish a nonlinear temporal correspondence between two video sequences captured by cameras independently moving in a dynamic 3D scene. We assume that the 3D spatial poses of the cameras are known for each frame. With predefined trajectory basis, the coefficients of the reconstructed trajectory of a moving scene point reflect the rhythm in motion. A robust rank constraint from the coefficient matrices is exploited to measure the spatiotemporal alignment quality for every feasible pair of video fragments. Point correspondences across sequences are not required or even it is possible that different points are tracked in different sequences, only if they satisfy the assumption that every 3D point tracked in the observed sequence can be described as a linear combination of a subset of the 3D points tracked in the reference sequence. Synchronization is then performed using a graph-based search algorithm to find the globally optimal path that minimizes both spatial and temporal misalignments. Our algorithm can use both complete and incomplete feature trajectories along time, and is robust to mild outliers. We verify the robustness and performance of the proposed approach on synthetic data as well as on challenging real video sequences. Xue Wang 0006, Jianbo Shi, Hyun Soo Park, Qing Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Aliasing Detection and Reduction Scheme on Angularly Undersampled Light FieldsabstractWhen using plenoptic camera for digital refocusing, angular undersampling can cause severe (angular) aliasing artifacts. Previous approaches have focused on avoiding aliasing by pre-processing the acquired light field via prefiltering, demosaicing, reparameterization, and so on. In this paper, we present a different solution that first detects and then removes angular aliasing at the light field refocusing stage. Different from previous frequency domain aliasing analysis, we carry out a spatial domain analysis to reveal whether the angular aliasing would occur and uncover where in the image it would occur. The spatial analysis also facilitates easy separation of the aliasing versus non-aliasing regions and angular aliasing removal. Experiments on both synthetic scene and real light field data sets (camera array and Lytro camera) demonstrate that our approach has a number of advantages over the classical prefiltering and depth-dependent light field rendering techniques. Zhaolin Xiao, Qing Wang 0006, Guoqing Zhou 0003, Jingyi Yu 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Dense Depth-Map Estimation and Geometry Inference from Light Fields via Global Optimization
Lipeng Si, Qing Wang 0006 |
ACCV (3) | 2 |
| 2016 | LFHOG: A discriminative descriptor for live face detection from light field imageabstractHow to avoid the invading of the attack in the biometric system, such as 2D printed photos, gradually becomes an important research hotspot. In this paper, we present a novel descriptor in light field to tackle the issue. Based on the angular and spatial information in light field, the proposed light field histogram of gradient (LFHoG) descriptor is derived from three directions, including vertical, horizontal and depth. Different with traditional HoG in 2D image, the gradient in depth direction is distinctive in light field. To validate the effectiveness of the proposed LFHoG descriptor, experiments have been carried out on light field datasets taken by a Lytro camera. The descriptor can achieve 99.75% accuracy on the user collected dataset, which proves the correctness and effectiveness of the LFHoG descriptor. Hao Zhu 0005, Qing Wang 0006 |
ICIP | 3 |
| 2016 | Rectifying projective distortion in 4D light fieldabstractThe accuracy of calibration will significantly affect the post processing capability of light field imaging. The geometry of the reconstructed scene is related to the parameters of light field closely, involving the accuracy of decoded rays and ambiguities from ray correspondences. Through exploring the ray correspondence, we derive a transformation matrix to describe the projective distortion on reconstructed scene in 4D light field. Based on our derivation, we simplify the light field camera geometry as a 4-parameter model and calibrate its intrinsic parameters, including a linear initialization and a nonlinear refine process. The proposed light field calibration can simply be implemented with a parallel bi-planar board. Experiments on both simulation and real scene data validate the performance of the calibration. Chunping Zhang, Qing Wang 0006 |
ICIP | 3 |
| 2016 | Decoding and calibration method on focused plenoptic cameraabstractThe ability of light gathering of plenoptic camera opens up new opportunities for a wide range of computer vision applications. An efficient and accurate method to calibrate plenoptic camera is crucial for its development. This paper describes a 10-intrinsic-parameter model for focused plenoptic camera with misalignment. By exploiting the relationship between the raw image features and the depth–scale information in the scene, we propose to estimate the intrinsic parameters from raw images directly, with a parallel biplanar board which provides depth prior. The proposed method enables an accurate decoding of light field on both angular and positional information, and guarantees a unique solution for the 10 intrinsic parameters in geometry. Experiments on both simulation and real scene data validate the performance of the proposed calibration method. Chunping Zhang, Qing Wang 0006 |
Comput. Vis. Media | 3 |
| 2016 | Accurate disparity estimation in light field using ground control pointsabstractThe recent development of light field cameras has received growing interest, as their rich angular information has potential benefits for many computer vision tasks. In this paper, we introduce a novel method to obtain a dense disparity map by use of ground control points (GCPs) in the light field. Previous work optimizes the disparity map by local estimation which includes both reliable points and unreliable points. To reduce the negative effect of the unreliable points, we predict the disparity at non-GCPs from GCPs. Our method performs more robustly in shadow areas than previous methods based on GCP work, since we combine color information and local disparity. Experiments and comparisons on a public dataset demonstrate the effectiveness of our proposed method. Hao Zhu 0005, Qing Wang 0006 |
Comput. Vis. Media | 2 |
| 2014 | Aliasing Detection and Reduction in Plenoptic ImagingabstractWhen using plenoptic camera for digital refocusing, angular undersampling can cause severe (angular) aliasing artifacts. Previous approaches have focused on avoiding aliasing by pre-processing the acquired light field via prefiltering, demosaicing, reparameterization, etc. In this paper, we present a different solution that first detects and then removes aliasing at the light field refocusing stage. Different from previous frequency domain aliasing analysis, we carry out a spatial domain analysis to reveal whether the aliasing would occur and uncover where in the image it would occur. The spatial analysis also facilitates easy separation of the aliasing vs. non-aliasing regions and aliasing removal. Experiments on both synthetic scene and real light field camera array data sets demonstrate that our approach has a number of advantages over the classical prefiltering and depth-dependent light field rendering techniques. Zhaolin Xiao, Qing Wang 0006, Guoqing Zhou 0003, Jingyi Yu 0001 |
CVPR | 2 |
| 2014 | Reconstructing scene depth and appearance behind foreground occlusion using camera arrayabstractForeground occlusion is a significant challenge in 3D reconstruction. In the paper, we first characterize the differences between multiview reconstruction with and without foreground occlusion. Considering both scene depth and appearance are unknown, we propose a generalized model for scene reconstruction. Then, we propose an iterative reconstruction approach in the global optimization framework, which is well performed on the camera array system. Even when all views are partially occluded, our approach can recover accurate depth map as well as scene appearance. Experimental results have indicated that our approach is more robust to foreground occlusions and outperforms state-of-the-art approaches. Zhaolin Xiao, Qing Wang 0006, Lipeng Si, Guoqing Zhou 0003 |
ICIP | 2 |
| 2014 | Sparse representation with geometric configuration constraint for line segment matching
Qing Wang 0006, Tingwang Chen, Lipeng Si |
Neurocomputing | 1 |
| 2013 | Enhanced Continuous Tabu Search for Parameter Estimation in Multiview GeometryabstractOptimization using the L_infty norm has been becoming an effective way to solve parameter estimation problems in multiview geometry. But the computational cost increases rapidly with the size of measurement data. Although some strategies have been presented to improve the efficiency of L_infty optimization, it is still an open issue. In the paper, we propose a novel approach under the framework of enhanced continuous tabu search (ECTS) for generic parameter estimation in multiview geometry. ECTS is an optimization method in the domain of artificial intelligence, which has an interesting ability of covering a wide solution space by promoting the search far away from current solution and consecutively decreasing the possibility of trapping in the local minima. Taking the triangulation as an example, we propose the corresponding ways in the key steps of ECTS, diversification and intensification. We also present theoretical proof to guarantee the global convergence of search with probability one. Experimental results have validated that the ECTS based approach can obtain global optimum efficiently, especially for large scale dimension of parameter. Potentially, the novel ECTS based algorithm can be applied in many applications of multiview geometry. Guoqing Zhou 0003, Qing Wang 0006 |
ICCV | 2 |
| 2012 | Robust Visual Tracking Using Dynamic Classifier Selection with Sparse Representation of Label Noise
Yuefeng Chen, Qing Wang 0006 |
ACCV (3) | 2 |
| 2012 | Spatio-Temporal Clustering Model for Multi-object Tracking through Occlusions
Qing Wang 0006 |
ACCV (3) | 2 |
| 2010 | 3D Line Segment Detection for Unorganized Point Clouds from Multi-view Stereo
Tingwang Chen, Qing Wang 0006 |
ACCV (2) | 2 |
| 2010 | A Comprehensive Evaluation on Non-deterministic Motion EstimationabstractWhen computing optical flow with region-based matching, very few of them can be reliably obtained, especially for the high-contrast areas or those with little texture. Instead of using a single pixel from the reference frame, non-deterministic motion utilizes multiple pixels within a neighborhood to represent the corresponding pixel in the current frame. Although remarkable improvement has been made with this method, the weight associated to each reference pixel is quite sensitive to the selection of its standard deviation. To address this issue, a dual probability is presented in this paper. Intuitively, it enhances those weights of pixels that are more similar to its counterpart in the current frame, while suppressing the rest of them. Experimental results show that the proposed method is effective to deal with intense motion and occlusion, especially in the case of reducing the adverse impact of noise. Changzhu Wu, Qing Wang 0006 |
ICPR | 2 |
| 2009 | Distance-Based Multiple Paths Quantization of Vocabulary Tree for Object and Scene Retrieval
Heng Yang 0003, Qing Wang 0006, Ellen Yi-Luen Do |
ACCV (1) | 2 |
| 2009 | Indexing Large Visual Vocabulary by Randomized Dimensions Hashing for High Quantization Accuracy: Improving the Object Retrieval Quality
Heng Yang 0003, Qing Wang 0006, Zhoucan He |
CAIP | 2 |
| 2009 | Grouping and Organizing Unordered Images for Multi-view Feature CorrespondencesabstractHandling numerous unordered images for scene reconstruction and categorization attracts increasing interests for commercial and scientific efforts. In this paper, we address the issue of efficient organization of content-related images from plenty of input images on several scenes with contaminated ones. First a robust view-similarity measure is proposed and the images can be categorized effectively without any constraints; then two speedup strategies, seed growing based grouping and tentative feature matching, are presented respectively. The experimental results on two image dataset demonstrate that the proposed method can efficiently and effectively organize unordered views without any geometric constraints, and can further provide nice data for 3D modeling. Zhoucan He, Qing Wang 0006 |
ICIG | 2 |
| 2009 | Efficient Scene Image Clustering for Internet CollectionsabstractThis paper proposes an efficient approach to find clusters of spatially related scene images collected from the website. Our method firstly builds a guide table, in which the ranked results are given according to the relevance scores of image pairs obtained by the image retrieval methods. Then the image clusters are generated by repeatedly choosing a seed image and performing query expansion directed by the guide table. In the query process, feature matching is performed by using an affine invariant constraint which is presented to effectively reject outliers of the image feature correspondences. The proposed image clustering approach has been tested on the Bell Tower dataset consisting of more than 1K images which are collected from the photo-sharing website Flickr.com. The experimental results demonstrate the efficiency and effectiveness of our method. Heng Yang 0003, Qing Wang 0006, Zhoucan He |
ICIG | 2 |
| 2009 | Joint image registration and super-resolution reconstruction based on regularized total least normabstractAccurate registration of the low resolution (LR) images is a critical step in image super resolution reconstruction (SRR). Conventional algorithms always use invariable motion parameters derived from registration algorithms, and carry on SRR without considering the registration errors in the disjointed method. In this paper we propose a new method that performs joint image registration and SRR based on regularized total least norm (RTLN), updating the motion parameters and HR image simultaneously. Not only translation but also rotation motion are considered, which makes the motion model more universal. Experimental results have shown that our approach is more effective and efficient than traditional ones. Qing Wang 0006, Xiaoli Song |
ICIP | 1 |
| 2009 | Multiple unordered wide-baseline image matching and groupingabstractThis paper focuses on the multi-view feature matching problem from unordered image sets. Firstly, an efficient and effective high dimensional feature matching algorithm is proposed, so called ELSH (extended local sensitive hash), which can significantly improve matching accuracy at fast speed. Secondly, a novel unsupervised image grouping strategy is proposed to cluster the unordered images into content-related group, which does not normally require any other constraints. Extensive experimental results have shown that our method can obtain better performance than the classical algorithms in tackling multi-view matching problem. Zhoucan He, Qing Wang 0006, Heng Yang 0003 |
ICME | 2 |
| 2008 | Indexing Sub-Vector Distance for High-Dimensional Feature MatchingabstractHigh-dimensional feature matching based on nearest neighbors search is a core part of many image-matching based problems in computer vision which are solved by local invariant features. In this paper, we propose a new indexing structure for the high-dimensional feature matching, which is based on the distance of the sub-vectors. In addition, we employ an effective image-similarity measure of two images based on the exponential distribution of the Euclidean distance between matched feature vectors. Experimental results have demonstrated the efficiency and effectiveness of the proposed methods in extensive image matching and image retrieval applications. 1 Heng Yang 0003, Qing Wang 0006, Zhoucan He |
BMVC | 2 |
| 2008 | MAP Model for Large-scale 3D Reconstruction and Coarse Matching for Unordered Wide-baseline PhotosabstractIn this paper we presented a novel idea for large-scale 3D scene reconstruction and annealing based image grouping algorithm for unordered wide-baseline photos. Firstly, an alternative maximum a posterior (MAP) model which can easily incorporate image clustering prior knowledge is proposed. Second, an efficient annealing clustering algorithm is developed for organizing photos into clusters by calculating matching number of invariant features. Thirdly, we analyze the time complexity and efficiency of the proposed approach. Finally a series of experiments are performed on the real image data and synthetic data. The experimental result shows that the MAP model and relative annealing algorithm are efficient enough to tackle the large-scale 3D reconstruction problem, and it can be extended to solve other similar SFM parameters estimation problem as well. 1 Xiuyuan Zeng, Qing Wang 0006, Jiong Xu |
BMVC | 2 |
| 2008 | Robust Wide Baseline Feature Point Matching Based on Scale Invariant Feature Descriptor
Sicong Yue, Qing Wang 0006, Rongchun Zhao |
ICIC (1) | 2 |
| 2008 | A novel local feature descriptor for image matchingabstractImage matching is a fundamental task of many problems in computer vision. This paper presents a novel local feature descriptor based on the gradient distance and orientation histogram (GDOH), which can be used for reliably matching between different views of a scene for wide baseline. The proposed descriptor is invariant to image scale, rotation, illumination and partial viewpoint changes. At present, the SIFT descriptor is generally considered as the most appealing descriptor for practical uses, but the high dimensionality is a drawback of SIFT in the feature matching step. The purpose of GDOH is to reduce the dimensional size of the descriptor, yet still maintain distinctness and robustness as much as SIFT. The experimental results show that the proposed descriptor can result in effectiveness and efficiency in image matching and image retrieval application. Heng Yang 0003, Qing Wang 0006 |
ICME | 2 |
| 2008 | Motion estimation approach based on dual-tree complex waveletsabstractThe loss of information due to occlusion and other complications has been one of the main bottlenecks in the field of motion estimation. In this paper, we propose a novel motion estimation algorithm based on dual-tree complex wavelet (DT-CWT), which utilizes its approximate shift invariance and directional selectivity. Subbands within different orientations and levels are individually treated to overcome the issue of error propagation induced by frequently used coarse-to-fine searching strategy, as these subbands are approximately independent from each other. Then, the frame next to the current one is introduced to supplement the lost information due to the occlusion and to improve the estimation accuracy within boundary areas. Experimental results show that the proposed method is effective with robustness to intense motion, scene change and occlusion. Changzhu Wu, Qing Wang 0006 |
ICPR | 2 |
| 2008 | Randomized sub-vectors hashing for high-dimensional image feature matchingabstractHigh-dimensional image feature matching is an important part of many image matching based problems in computer vision which are solved by local invariant features. In this paper, we propose a new indexing/searching method based on Randomized Sub-Vectors Hashing (called RSVH) for high-dimensional image feature matching. The essential of the proposed idea is that the feature vectors are considered similar (measured by Euclidean distance) when the L2 norms of their corresponding randomized sub-vectors are approximately same respectively. Experimental results have demonstrated that our algorithm can perform much better than the famous BBF (Best-Bin-First) and LSH (Locality Sensitive Hashing) algorithms in extensive image matching and image retrieval applications. Heng Yang 0003, Qing Wang 0006, Zhoucan He |
ACM Multimedia | 2 |
| 2007 | A Novel Macroblock Layer Rate Control for H.264/AVCabstractA novel macroblock layer rate control algorithm for H.264/ AVC is proposed in this paper. To solve the issues of rate control model in H.264/AVC, we presented a new coding complexity measure based on integer transform coefficients, so-called mean absolute transform quantized distortion (MATQD). Then, we proposed four effective models respectively to improve our rate control performance, including MATQD prediction model, target bits allocation model, header bits prediction model for a marcoblock and the quadratic R-D model based on MATQD. Experimental results have shown that our method has more precise rate control performance than that of the standard rate control algorithm JVT-G012 in JM8.5 with the average absolute rate error reducing 0.396 bps. Furthermore, the proposed method can produce much smoother bit fluctuation with the average rate deviation reducing 21.74% and obtain much better average image quality with average PSNR improving 0.336 dB (the best gain is up to 0.77 dB). Heng Yang 0003, Qing Wang 0006 |
ICME | 2 |
| 2006 | QP_TR Trust Region Blob Tracking Through Scale-SpaceabstractA new approach of tracking objects in image sequences is proposed, in which the constant changes of the size and orientation of the target can be precisely described. For each incoming frame, a probability distribution image of the target is created, where the target's area turns into a blob. The scale of this blob can be determined based on the local maxima of differential scale-space filters. We employ the QP_TR trust region algorithm to search the local maxima of orientational multi-scale normalized Laplacian filter of the probability distribution image to locate the target as well as to determine its scale and orientation. Based on the tracking results of sequence examples, the new method is proven to be capable of describing the target more accurately and thus achieves much better tracking precision. Jingping Jia, Qing Wang 0006, Yanmei Chai, Rongchun Zhao |
ICIP | 2 |
| 2006 | A New Oriented Adaptive Cross Search Algorithm for Block Matching Motion EstimationabstractBlock-matching motion estimation plays an important role in video coding and faster, more robust and more effective search algorithms are needed. Recently, a great number of fast block matching algorithms (BMAs) have been proposed in the literature based on the discovery of the center-biased characteristics of motion-vector distribution. In this paper, a novel oriented adaptive cross search (OACS) algorithm is proposed, where small cross, large cross and T-shape search patterns are defined and utilized adaptively. In accordance with the adaptive tracing of orientation change or optimal point, three kinds of key points are defined to decide which kind or which oriented pattern may be chosen for the next step. Experimental results on the benchmarks have shown that the OACS algorithm can provide average speed ups of 74.65%, 39.78%, 42.44%, and 7.84% over DS, SDS, CDS, and SCDS, respectively. Finally, the mean absolute distortion and PSNR of luminance component are close to the results of other fast BMAs and the similar search accuracy can be maintained as expected Heng Yang 0003, Qing Wang 0006 |
ICME | 2 |
| 2005 | Texture analysis and retrieval using fractal signature and B-spline wavelet transform with second order derivativeabstractIn the paper, we proposed a novel over-complete B-spline wavelet transform and fractal signature for texture image analysis and retrieval. Traditionally, discrete wavelet frame took the first order derivative of smoothing function into account, which is equivalent to Canny edge detection. The second order derivative spline wavelet has the ability to detect the variation of the edge width, and the finite impulse response is conducted in the paper. Additionally, statistical features in wavelet domain and fractal signature are utilized in the retrieval. Experimental results have shown that the proposed method is reasonable to describe the essences of the textures and can reach the highest retrieval rate (76.54%) comparing with Gabor Filter and first order derivative over-complete wavelet transform. Qing Wang 0006, David Dagan Feng |
ICIP (1) | 1 |
| 2003 | Efficient Evaluation in XML to XML Transformations
Qing Wang 0006, Junmei Zhou, Hongwei Wu, Yizhong Wu, Aoying Zhou |
APWeb | 1 |
| 2003 | FIXT: A Flexible Index for XML Transformation
Jianchang Xiao, Qing Wang 0006, Aoying Zhou |
APWeb | 2 |
| 2003 | TREX: DTD-Conforming XML to XML Transformations
Aoying Zhou, Qing Wang 0006, Zhimao Guo, Xueqing Gong, Shihui Zheng, Hongwei Wu, Jianchang Xiao, Kun Yue, Wenfei Fan |
SIGMOD Conference | 2 |
| 2003 | Peer-Serv: A Framework of Web Services in Peer-to-Peer Environment
Qing Wang 0006, Junmei Zhou, Aoying Zhou |
WAIM | 1 |
| 2003 | UD(k, l)-Index: An Efficient Approximate Index for XML Data
Hongwei Wu, Qing Wang 0006, Jeffrey Xu Yu, Aoying Zhou, Shuigeng Zhou |
WAIM | 2 |
| 2003 | Hierarchical content classification and script determination for automatic document image processing
Zheru Chi, Qing Wang 0006, Wan-Chi Siu |
Pattern Recognit. | 2 |
| 2002 | A fast 2D entropic thresholding method by wavelet decompositionabstractCompared with ID grayscale histogram analysis, 2D entropic thresholding makes use of local average as well as pixel gray level. However, it is time consuming to search the threshold vector in the 2D histogram. In the paper, a fast algorithm using wavelet decomposition is proposed, with which a set of candidates of the vector was first obtained in the decomposed histogram. The optimal threshold vector is then obtained without exhaustive searching. Experimental results have shown that our algorithm not only finds the threshold vector as well as Brink's method (1992) but also saves computation costs, using up only 0.53% of the processing time taken by exhaustive searching. Qing Wang 0006, Qiurang Wang, David Dagan Feng, Rongchun Zhao, Zheru Chi |
ICIP (3) | 1 |
| 2002 | Image Thresholding by Maximizing the Index of Nonfuzziness of the 2-D Grayscale Histogram
Qing Wang 0006, Zheru Chi, Rongchun Zhao |
Comput. Vis. Image Underst. | 1 |
| 2001 | Match Between Normalization Schemes and Feature Sets for Handwritten Chinese Character RecognitionabstractBecause of the large number of Chinese characters and many different writing styles involved, the recognition of handwritten Chinese characters remains a very challenging task. It is well recognized that a good feature set plays a key role in a successful recognition system. Shape normalization is as well an essential step toward achieving translation, scale, and rotation invariance in recognition. Many shape normalization methods and different feature sets have been proposed in the literature. We first review five commonly used shape normalization schemes and then discuss various feature extraction techniques usually used in handwritten Chinese character recognition. Based on numerous experiments conducted on 3,755 handwritten Chinese characters (GB2312-80), we discuss the matches made between the normalization schemes and the feature sets and suggest the best match between them in terms of classification performance. The nearest neighbor classifier was adopted in our experiments with templates obtained by using the K-means clustering algorithm. Qing Wang 0006, Zheru Chi, David Dagan Feng, Rongchun Zhao |
ICDAR | 1 |
| 2001 | Handwritten Chinese Character Segmentation Using a Two-Stage ApproachabstractCorrect segmentation of handwritten Chinese characters is crucial to the successful recognition. However, because of the many difficulties involved, little work has been done in this area. In this paper, a two-stage approach is addressed to segment unconstrained handwritten Chinese character strings. A string is first coarsely segmented according to the background skeleton and vertical projection after a proper image preprocessing. At the fine segmentation stage that follows, the strokes that may contain segmentation points are first identified. The feature points are then extracted from candidate strokes and taken as segmentation point candidates through each of which a segmentation path may be formed. Geometric features are extracted and fuzzy decision rules learned from examples are used to evaluate the segmentation paths. By using this two-stage segmentation approach, we can achieve both good performance and efficiency in segmenting unconstrained handwritten Chinese characters. Shuyan Zhao, Zheru Chi, Qing Wang 0006 |
ICDAR | 4 |
| 2000 | Hidden Markov Random Field Based Approach for Off-Line Handwritten Chinese Character RecognitionabstractThis paper presents a hidden Markov mesh random field (HMMRF) based approach for off-line handwritten Chinese characters recognition using statistical observation sequences embedded in the strokes of a character. Due to a large set of Chinese characters and many different writing styles, the recognition of handwritten Chinese characters is very challenging. In our approach, the binary image is first normalized by a nonlinear shape normalization scheme to adjust the width, length, and the correlation of strokes. Two types of stroke-based features are then extracted to represent the observation sequence. The estimation of model parameters and state sequence decoding algorithms are also discussed in the paper. Experimental results on 470 isolated handwritten Chinese characters demonstrate the effectiveness of our approach. Qing Wang 0006, Rongchun Zhao, Zheru Chi, David Dagan Feng |
ICPR | 1 |