Hui Zhang 0062

dblp:181/2846-62 · DBLP profile ↗
← Back
36ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-1681-7926ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies
abstract
Large Vision-Language Models (LVLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they still face critical challenges in modeling long-range dependencies under the usage of Rotary Positional Encoding (ROPE). Although it can facilitate precise modeling of token positions, it induces progressive attention decay as token distance increases, especially with progressive attention decay over distant token pairs, which severely impairs the model's ability to remember global context. To alleviate this issue, we propose inference-only Three-step Decay Resilience Strategies (T-DRS), comprising (1) Semantic-Driven DRS (SD-DRS), amplifying semantically meaningful but distant signals via content-aware residuals, (2) Distance-aware Control DRS (DC-DRS), which can purify attention by smoothly modulating weights based on positional distances, suppressing noise while preserving locality, and (3) re-Reinforce Distant DRS (reRD-DRS), consolidating the remaining informative remote dependencies to maintain global coherence. Together, the T-DRS recover suppressed long-range token pairs without harming local inductive biases. Extensive experiments on Vision Question Answering (VQA) benchmarks demonstrate that T-DRS can consistently improve performance in an inference-only manner.
Yujian Lee, Zailong Chen, Hui Zhang 0062
AAAI5
2026 Describing-Verifying-Scoring: A Hierarchical Reasoning Framework for Zero-Shot Composed Image Retrieval
abstract
Zero-Shot Composed Image Retrieval (ZS-CIR) aims to identify target images using a composed query of a reference image and modification text without labeled triplets. While recent advances leverage Multimodal Large Language Models (MLLMs) for intent reasoning, they often suffer from hallucination-induced inaccuracies where misaligned descriptions degrade retrieval reliability, and insufficient reasoning due to shallow prompting strategies. To address these challenges, we propose DVSCIR, a novel training-free framework featuring a hierarchical Describing-Verifying-Scoring pipeline with MLLM. Specifically, the Describing stage generates an initial candidate caption, followed by a Verifying stage that rectifies potential hallucinations to ensure description accuracy. The Scoring stage performs a fine-grained re-ranking to identify the optimal match. Within each stage, a hierarchical Chain-of-Thought (CoT) process tailored for ZS-CIR guides the MLLM from low-level perception to deep intentional reasoning via sequential steps within structured sections. This progression ensures robust cross-modal correspondence through a hierarchical refinement of the retrieval process. Extensive experiments across four benchmarks demonstrate that DVSCIR achieves state-of-the-art performance, validating its effectiveness in ZS-CIR.
Guquan Jing, Yujian Lee, Hui Zhang 0062
ICMR4
2025 Face Relighting with Ratio Function for Explicit Geometric Representation
abstract
This paper addresses the problem of face relighting under varying illumination conditions. Lighting is a fundamental element in portrait photography that shapes the mood, geometry, and overall realism of the captured characters. Most previous studies have mainly treated relighting as a 2D generation task without incorporating the geometric features of the characters. In contrast, inspired by ratio image-based methods, this paper proposes to disentangle shadow and brightness variations through geometric information and utilizes generative adversarial networks (GANs) to obtain relighted images with brightness consistency. We design a novel relighting-ratio function that integrates the Cook-Torrance reflectance model to more explicitly represent the face geometry than previous ratio image-based methods. This relighting-ratio function is derived from an image rendering formula that quantizes variables such as albedo that are affected by the lighting direction, while systematically excluding variables such as normal and viewpoint that are not affected by lighting. We conduct quantitative and qualitative experiments on the Multi-PIE and CelebA-HQ datasets and show that the proposed method outperforms existing SOTA methods using lighting directions.
Yiyang Hu, Zequn Zhang, Hui Zhang 0062, Guquan Jing
ICASSP3
2025 ESTI: An Efficient Spatial-Temporal Interaction Network For Video-Based Person Re-Identification
abstract
Video-based person re-identification (Re-ID) aims to identify the target pedestrian from video sequences. However, redundant information exist in input frames. Extracting spatial-temporal features in whole adjacent frames can introduce additional computational overhead. Furthermore, this process leads to the loss of critical spatial and temporal details, causing suboptimal representations. To mitigate these issues, we propose an Efficient Spatial-Temporal Interaction (ESTI) network, which processes half of the input sequence separately through spatial and temporal branches, extracting high-level discriminative features across multiple layers and avoiding redundancy computations. In particular, we propose a Feature Enhancement Module (FEM) for the spatial branch to focus on enhancing spatial dependencies adaptively, and a Temporal Interaction Module (TIM) for temporal branch to capture temporal correlations effectively. Spatial-temporal interaction is performed at the final layer to generate distinctive representations. Extensive experiments on three challenging video Re-ID datasets show that our ESTI achieves competitive results while maintaining low computational complexity.
Guquan Jing, Yiyang Hu, Yujian Lee, Hui Zhang 0062
ICME5
2025 ContrastiveGaussian: High-Fidelity 3D Generation with Contrastive Learning and Gaussian Splatting
abstract
Creating 3D content from single-view images is a challenging problem that has attracted considerable attention in recent years. Current approaches typically utilize score distillation sampling (SDS) from pre-trained 2D diffusion models to generate multi-view 3D representations. Although some methods have made notable progress by balancing generation speed and model quality, their performance is often limited by the visual inconsistencies of the diffusion model outputs. In this work, we propose ContrastiveGaussian, which integrates contrastive learning into the generative process. By using a perceptual loss, we effectively differentiate between positive and negative samples, leveraging the visual inconsistencies to improve 3D generation quality. To further enhance sample differentiation and improve contrastive learning, we incorporate a super-resolution model and introduce another Quantity-Aware Triplet Loss to address varying sample distributions during training. Our experiments demonstrate that our approach achieves superior texture fidelity and improved geometric consistency. Code will be available at https://github.com/YaNLlan-Ljb/ContrastiveGaussian.
Junbang Liu, Enpei Huang, Dongxing Mao, Hui Zhang 0062, Xinyuan Song 0002, Yongxin Ni
ICME4
2025 Fast and Physically-based Neural Explicit Surface for Relightable Human Avatars
abstract
Efficiently modeling relightable human avatars from sparse-view videos is crucial for AR/VR applications. Current methods use neural implicit representations to capture dynamic geometry and reflectance, which incur high costs due to the need for dense sampling in volume rendering. To overcome these challenges, we introduce Physically-based Neural Explicit Surface (PhyNES), which employs compact neural material maps based on the Neural Explicit Surface (NES) representation. PhyNES organizes human models in a compact 2D space, enhancing material disentanglement efficiency. By connecting Signed Distance Fields to explicit surfaces, PhyNES enables efficient geometry inference around a parameterized human shape model. This approach models dynamic geometry, texture, and material maps as 2D neural representations, enabling efficient rasterization. PhyNES effectively captures physical surface attributes under varying illumination, enabling real-time physically-based rendering. Experiments show that PhyNES achieves relighting quality comparable to SOTA methods while significantly improving rendering speed, memory efficiency, and reconstruction quality.
Jie Chen 0026, Hui Zhang 0062
ICME4
2025 Contextual Reasoning for Robust Composed Image Retrieval with Vision-Language Models
abstract
Composed Image Retrieval (CIR) combines a reference image with modification text for precise and flexible searches. However, existing methods face two key challenges: first, the limited information in modification text hampers the model's ability to understand user intent, leading to reduced accuracy and diversity; second, reliance on unidirectional constraints overlooks the complementary role of reference and target captions. In this paper, we propose CR-CIR a novel framework that leverages Contextual Reasoning and vision-language models to enhance CIR. Specifically, we use a VLM (e.g., BLIP2) to address the scarcity of textual annotations in existing datasets by generating descriptive captions for both reference and target images. In addition, we enhance the modification text with contextual information using a VLM (e.g., MiniCPM), enriching the model's understanding of user intent. Then our method incorporates a Dual Reasoning Modification Module, which imposes bidirectional constraints by integrating both image and text modalities. Additionally, we introduce a Modality Shift Regularization Loss that assumes symmetry and correlation between text and image domain transformations in the latent space. This new loss function enforces consistent modality shifts, significantly enhancing the model's interpretative and generalization abilities. Experimental results on benchmark CIR datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. Our code and dataset will be available at https://github.com/kola1124/CR-CIR.git.
Yujian Lee, Xubo Liu 0001, Hui Zhang 0062, Zailong Chen, Yiyang Hu, Guquan Jing, Yunting Lai
ICMR4
2025 Text-Guided Realistic Single Image Relighting with Wavelet Mamba Diffusion Network
Yunting Lai, Hui Zhang 0062, Yiyang Hu, Guquan Jing
ICMR2
2025 ShadowAdapter: Adapting Segment Anything Model with Auto-Prompt for shadow detection
Leiping Jie, Hui Zhang 0062
Expert Syst. Appl.2
2025 3D-Aided Pedestrian Representation Learning for Video-Based Person Re-Identification
abstract
Video-based person re-identification (Re-ID) aims to match the target pedestrian from video sequences. Recent methods perform frame-level feature extraction followed by temporal aggregation to obtain video representations. However, they pay insufficient attention to the quality of frame-level features, which suffer from issues including multi-frame misalignment, partial occlusion and appearance confusion. People live in a 3D space. 3D pedestrian representations can provide rich geometric information and shape cues that offer promising solutions to these challenges in video-based Re-ID. To mitigate these issues, this paper proposes a 3D-Aid Pedestrian Representation Learning (3DAPRL) network, which introduces 3D modality to video-based Re-ID. Specifically, two novel modules are designed,i.e., the Cross-Modal Fusion (CMF) module and the Shape-aware Spatial-Temporal Interaction (SSTI) module, to enhance pedestrian representation learning. The CMF module generates discriminative fusion representations by utilizing 3D pedestrian data, while the SSTI module learns spatial-temporal 3D shape representation which are distinguishable for finding the target pedestrian in video scenarios. Both features generated from the CMF and SSTI modules contribute to the final video representation. Extensive experiments on four challenging video-based Re-ID datasets demonstrate that our 3DAPRL network reaches better performance than state-of-the-arts methods.
Guquan Jing, Yujian Lee, Yiyang Hu, Hui Zhang 0062
IEEE Trans. Circuits Syst. Video Technol.5
2023 DEdgeNet: Extrinsic Calibration of Camera and LiDAR with Depth-discontinuous Edges
abstract
This paper addresses the problem of calibrating extrinsic parameter matrix between an RGB camera and a LiDAR. Multimodal sensing systems are essential for fully autonomous navigation platforms. A key pre-requisite for such a system is calibration between different sensors. As the two most widely equipped sensors, calibration between RGB cameras and LiDARs remains challenging. Existing methods address this problem without using explicit geometric priors. In this paper, we propose a novel real-time network that utilizes depth-discontinuous edges extracted from a single image to calibrate cameras and LiDARs. Our network consists of two key components: (1) a self-supervised edge extraction network named DEdgeNet, which detects depth-discontinuous edges from a single image and extracts corresponding features; (2) prediction of the extrinsic parameter matrix between the camera and the LiDAR by matching fixed features in RGB images and updating depth features in a coarse-to-fine frame. Specifically, considering that edges are rich and common in natural scenes, DEdgeNet simplifies RGB image encoding and extracts fixed edges for feature matching. We conducted extensive experiments on the KITTI-odometry dataset. The results show that our method achieves an average rotation error of 0.028° and an average translation error of 0.247 cm, which demonstrates the superiority of our method.
Yiyang Hu, Leiping Jie, Hui Zhang 0062
ICRA4
2023 Linear Auto-calibration of Pan-Tilt-Zoom Cameras With Rotation Center Offset
abstract
This paper addresses the linear auto-calibration problem of a pan-tilt-zoom (PTZ) camera. Unlike existing methods, we take full advantage of the offset of the camera center from the rotation center, which is usually non-negligible in bullet-type PTZ cameras. Without any prior assumption, we propose a linear method to recover all intrinsic parameters. First, we successively acquired at least four images using the zoom and rotation capabilities of the PTZ camera. Second, using the homography of two images at the same location but different scales, the principal point and zoom scalar can be linearly recovered. Finally, based on the unknown offset of the camera center and rotation center, we propose a linear method to solve the scale factor in the Kruppa equation and recover the remaining camera intrinsic parameters, namely focal lengths and skew. Synthetic and real experiments demonstrate the feasibility of our approach.
Yu Liu 0046, Hui Zhang 0062
ICRA2
2023 RMLANet: Random Multi-Level Attention Network for Shadow Detection and Removal
abstract
This paper addresses the problem of shadow detection and shadow removal from a single image. Despite awareness of utilizing both local and global contexts, previous works only aggregate features level by level in a coarse-to-fine manner. To overcome this problem, we present RMLANet, a novel Random Multi-Level Attention Network. To be specific, we first design a shuffled multi-level feature aggregation module to fuse the multi-level features and the guiding features using the self-attention mechanism. Nevertheless, the computational complexity of dense self-attention is unaffordable when processing high-resolution inputs. We argue that dense attention between any pixel pair is unnecessary due to the local consistency in images. Then we further propose a sparse attention mechanism to reduce the number of attention pairs, which greatly reduces the computational complexity. Through extensive experiments on four shadow detection and three shadow removal benchmark datasets, our proposed RMLANet achieves superior performance over current state-of-the-art approaches for both shadow detection and shadow removal. Codes are publicly available athttps://github.com/LeipingJie/RMLANet.
Leiping Jie, Hui Zhang 0062
IEEE Trans. Circuits Syst. Video Technol.2
2022 MGRLN-Net: Mask-Guided Residual Learning Network for Joint Single-Image Shadow Detection and Removal
Leiping Jie, Hui Zhang 0062
ACCV (3)2
2022 Camera Auto-calibration from the Steiner Conic of the Fundamental Matrix
Yu Liu 0046, Hui Zhang 0062
ECCV (2)2
2022 PSP-MVSNet: Deep Patch-Based Similarity Perceptual for Multi-view Stereo Depth Inference
Leiping Jie, Hui Zhang 0062
ICANN (1)2
2022 A Fast and Efficient Network for Single Image Shadow Detection
abstract
Shadows in images can degrade the performance of many applications. In this paper, we propose a novel multi-level feature-aware network, called TransShadow, which uses Transformer to capture both local and global context from a single image for shadow detection. Specifically, we design a multi-level feature-aware module, where multi-level features are selected and processed by the Transformer to distinguish shadowed and non-shadowed regions. To further utilize the remaining feature levels, progressive upsampling with skip connections is proposed to fuse more information for shadow detection. Experimental results show that our approach achieves comparative performance as the state-of-the-art method on benchmark datasets SBU and ISTD with the smallest model size and fastest inference speed. More importantly, our model shows the best generalization performance on the benchmark dataset UCF.
Leiping Jie, Hui Zhang 0062
ICASSP2
2022 RMLANet: Random Multi-Level Attention Network for Shadow Detection
abstract
This paper addresses the problem of shadow detection from a single image. Previous approaches have shown that exploiting both global and local contexts in deep convolutional neural network layers can greatly improve performance. However, multi-level contexts remain underexplored. To achieve this, we propose RMLANet, a novel Random Multi-Level Attention Network. Specifically, we leverage shuffled multi-level features simultaneously with guiding features, and employ the transformer to capture global context. Furthermore, to reduce the computational and memory overhead caused by the self-attention mechanism in the vanilla transformer, we propose a random sampling strategy to reduce the number of inputs to the transformer. This is motivated by observing local consistency in images, which suggests that dense attention is unnecessary. Extensive experimental results demonstrate that our method outperforms current state-of-the-art methods on three widely used benchmark datasets SBU, ISTD and UCF.
Leiping Jie, Hui Zhang 0062
ICME2
2022 Coverage hole detection in WSN with force-directed algorithm and transfer learning
Yue-Hui Lai, Se-Hang Cheong, Hui Zhang 0062, Yain-Whar Si
Appl. Intell.3
2018 Intrinsic Calibration of a Camera to a Line-Structured Light Using a Single View of Two Spheres
Yu Liu 0046, Xiaoyong Zhou, Qingqing Ma, Hui Zhang 0062
ACIVS5
2016 Homography Estimation from the Common Self-Polar Triangle of Separate Ellipses
abstract
How to avoid ambiguity is a challenging problem for conic-based homography estimation. In this paper, we address the problem of homography estimation from two separate ellipses. We find that any two ellipses have a unique common self-polar triangle, which can provide three line correspondences. Furthermore, by investigating the location features of the common self-polar triangle, we show that one vertex of the triangle lies outside of both ellipses, while the other two vertices lies inside the ellipses separately. Accordingly, one more line correspondence can be obtained from the intersections of the conics and the common self-polar triangle. Therefore, four line correspondences can be obtained based on the common self-polar triangle, which can provide enough constraints for the homography estimation. The main contributions in this paper include: (1) A new discovery on the location features of the common self-polar triangle of separate ellipses. (2) A novel approach for homography estimation. Simulate experiments and real experiments are conducted to demonstrate the feasibility and accuracy of our approach.
Haifei Huang, Hui Zhang 0062, Yiu-Ming Cheung
CVPR2
2016 The common self-polar triangle of separate circles: Properties and applications to camera calibration
abstract
This paper investigates the properties of the common self-polar triangle of separate coplanar circles and applies them to camera calibration. We find that any two separate circles have a unique common self-polar triangle. In particular, we show that one vertex of the common self-polar triangle lies on the line at infinity. Given three separate circles, the line at infinity can be recovered using the vertices of the common self-polar triangles. Accordingly, the vanishing line of the support plane can be obtained in their images. This allows recovering the imaged circular points, which provides good constraints on the image of the absolute conic. Compared to previous calibration methods using separate circles, our approach can avoid solving quartic equation, which often causes numerical instability. In the application, we test one calibration algorithm and accurate results are achieved.
Haifei Huang, Hui Zhang 0062, Yiu-Ming Cheung
ICIP2
2015 The common self-polar triangle of concentric circles and its application to camera calibration
abstract
In projective geometry, the common self-polar triangle has often been used to discuss the position relationship of two planar conics. However, there are few researches on the properties of the common self-polar triangle, especially when the two planar conics are special conics. In this paper, we explore the properties of the common self-polar triangle, when the two conics happen to be concentric circles. We show there exist infinite many common self-polar triangles of two concentric circles, and provide a method to locate the vertices of these triangles. By investigating all these triangles, we find that they encode two important properties. The first one is all triangles share one common vertex, and the opposite side of the common vertex lies on the same line, which are the circle center and the line at the infinity of the support plane. The second is all triangles are right triangles. Based on these two properties, the imaged circle center and the varnishing line of support plane can be recovered simultaneously, and many conjugate pairs on vanishing line can be obtained. These allow to induce good constraints on the image of absolute conic. We evaluate two calibration algorithms, whereby accurate results are achieved. The main contribution of this paper is that we initiate a new perspective to look into circle-based camera calibration problem. We believe that other calibration methods using different circle patterns can benefit from this perspective, especially for the patterns which involve more than two circles.
Haifei Huang, Hui Zhang 0062, Yiu-Ming Cheung
CVPR2
2014 Camera Calibration Based on the Common Self-polar Triangle of Sphere Images
Haifei Huang, Hui Zhang 0062, Yiu-Ming Cheung
ACCV (2)2
2012 Self-calibration and Motion Recovery from Silhouettes with Two Mirrors
Hui Zhang 0062, Ling Shao 0001, Kwan-Yee Kenneth Wong
ACCV (4)1
2012 One shot learning gesture recognition with Kinect sensor
abstract
Gestures are both natural and intuitive for Human-Computer-Interaction (HCI) and the one-shot learning scenario is one of the real world situations in terms of gesture recognition problems. In this demo, we present a hand gesture recognition system using the Kinect sensor, which addresses the problem of one-shot learning gesture recognition with a user-defined training and testing system. Such a system can behave like a remote control where the user can allocate a specific function using a prefered gesture by performing it only once. To adopt the gesture recognition framework, the system first automatically segments an action sequence into atomic tokens, and then adopts the Extended-Motion-History-Image (Extended-MHI) for motion feature representation. We evaluate the performance of our system quantitatively in Chalearn Gesture Challenge, and apply it to a virtual one shot learning gesture recognition system.
Di Wu 0009, Fan Zhu 0001, Ling Shao 0001, Hui Zhang 0062
ACM Multimedia4
2011 A Generalized Coding Artifacts and Noise Removal Algorithm for Digitally Compressed Video Signals
Ling Shao 0001, Hui Zhang 0062, Yan Liu 0004
MMM (1)2
2011 Transform based spatio-temporal descriptors for human action recognition
Ling Shao 0001, Ruoyun Gao, Yan Liu 0004, Hui Zhang 0062
Neurocomputing4
2009 Self-Calibration of Turntable Sequences from Silhouettes
abstract
This paper addresses the problem of recovering both the intrinsic and extrinsic parameters of a camera from the silhouettes of an object in a turntable sequence. Previous silhouette-based approaches have exploited correspondences induced by epipolar tangents to estimate the image invariants under turntable motion and achieved a weak calibration of the cameras. It is known that the fundamental matrix relating any two views in a turntable sequence can be expressed explicitly in terms of the image invariants, the rotation angle, and a fixed scalar. It will be shown that the imaged circular points for the turntable plane can also be formulated in terms of the same image invariants and fixed scalar. This allows the imaged circular points to be recovered directly from the estimated image invariants, and provide constraints for the estimation of the imaged absolute conic. The camera calibration matrix can thus be recovered. A robust method for estimating the fixed scalar from image triplets is introduced, and a method for recovering the rotation angles using the estimated imaged circular points and epipoles is presented. Using the estimated camera intrinsics and extrinsics, a Euclidean reconstruction can be obtained. Experimental results on real data sequences are presented, which demonstrate the high precision achieved by the proposed method.
Hui Zhang 0062, Kwan-Yee Kenneth Wong
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Motion Recovery for Uncalibrated Turntable Sequences Using Silhouettes and a Single Point
Hui Zhang 0062, Ling Shao 0001, Kwan-Yee Kenneth Wong
ACIVS1
2008 1D Camera Geometry and Its Application to the Self-Calibration of Circular Motion Sequences
abstract
This paper proposes a novel method for robustly recovering the camera geometry of an uncalibrated image sequence taken under circular motion. Under circular motion, all the camera centers lie on a circle and the mapping from the plane containing this circle to the horizon line observed in the image can be modelled as a 1D projection. A 2 x 2 homography is introduced in this paper to relate the projections of the camera centers in two 1D views. It is shown that the two imaged circular points of the motion plane and the rotation angle between the two views can be derived directly from such a homography. This way of recovering the imaged circular points and rotation angles is intrinsically a multiple view approach, as all the sequence geometry embedded in the epipoles is exploited in the estimation of the homography for each view pair. This results in a more robust method compared to those computing the rotation angles using adjacent views only. The proposed method has been applied to self-calibrate turntable sequences using either point features or silhouettes, and highly accurate results have been achieved.
Kwan-Yee Kenneth Wong, Guoqiang Zhang 0003, Chen Liang 0004, Hui Zhang 0062
IEEE Trans. Pattern Anal. Mach. Intell.4
2008 An Overview and Performance Evaluation of Classification-Based Least Squares Trained Filters
abstract
An overview of the classification-based least squares trained filters on picture quality improvement algorithms is presented. For each algorithm, the training process is unique and individually selected classification methods are proposed. Objective evaluation is carried out to single out the optimal classification method for each application. To optimize combined video processing algorithms, integrated solutions are benchmarked against cascaded filters. The results show that the performance of integrated designs is superior to that of cascaded filters when the combined applications have conflicting demands in the frequency spectrum.
Ling Shao 0001, Hui Zhang 0062, Gerard de Haan
IEEE Trans. Image Process.2
2007 Camera Calibration from Images of Spheres
abstract
This paper introduces a novel approach for solving the problem of camera calibration from spheres. By exploiting the relationship between the dual images of spheres and the dual image of the absolute conic (IAC), it is shown that the common pole and polar with regard to the conic images of two spheres are also the pole and polar with regard to the IAC. This provides two constraints for estimating the IAC and, hence, allows a camera to be calibrated from an image of at least three spheres. Experimental results show the feasibility of the proposed approach.
Hui Zhang 0062, Kwan-Yee Kenneth Wong, Guoqiang Zhang 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 1D Camera Geometry and Its Application to Circular Motion Estimation
abstract
This paper describes a new and robust method for estimating circular motion geometry from an uncalibrated image sequence. Under circular motion, all the camera centers lie on a circle, and the mapping of the plane containing this circle to the horizon line in the image can be modelled as a 1D projection. A 2×2 homography is introduced in this paper to relate the projections of the camera centers in two 1D views. It is shown that the two imaged circular points and the rotation angle between the two views can be derived directly from the eigenvectors and eigenvalues of such a homography respectively. The proposed 1D geometry can be nicely applied to circular motion estimation using either point correspondences or silhouettes. The method introduced here is intrinsically a multiple view approach as all the sequence geometry embedded in the epipoles is exploited in the computation of the homography for a view pair. This results in a robust method which gives accurate estimated rotation angles and imaged circular points. Experimental results are presented to demonstrate the simplicity and applicability of the new method.
Guoqiang Zhang 0003, Hui Zhang 0062, Kwan-Yee Kenneth Wong
BMVC2
2005 Auto-Calibration and Motion Recovery from Silhouettes for Turntable Sequences
abstract
This paper addresses the problem of structure and motion from silhou-ettes for turntable sequences. Previous works have exploited corresponding points induced by epipolar tangencies to estimate the image invariants un-der turntable motion and recover the epipolar geometry. In these approaches, however, camera intrinsics are needed in order to obtain Euclidean motion and reconstruction. This paper proposes a novel approach to precisely esti-mate the image invariants and the rotation angles in the absence of the camera intrinsics, and to perform auto-calibration. By exploiting a special parame-terization of the epipoles, it is shown that the imaged circular points can be formulated in terms of the image invariants. A fixed scalar κ, introduced to account for the different scales in the homogeneous representations of the image invariants used in the parameterizations, is found crucial in both cali-bration and motion estimation. Given the image invariants, namely the hori-zon, the imaged rotation axis and its orthogonal vanishing point, this scalar can be determined from the epipoles in an image triplet. A robust method for estimating κ is proposed and the rotation angles can be recovered using this estimated value of κ. All the estimated variables are then refined using bundle-adjustment and auto-calibration is performed using the imaged circu-lar points, the imaged rotation axis and the associated vanishing point. This allows the recovery of the full camera positions and orientations, and hence Euclidean reconstruction. Experimental results demonstrate the simplicity of this novel approach and the high precision in the estimated motion and reconstruction. 1
Hui Zhang 0062, Guoqiang Zhang 0003, Kwan-Yee Kenneth Wong
BMVC1
2005 Camera calibration with spheres: linear approaches
abstract
This paper addresses the problem of camera calibration from spheres. By studying the relationship between the dual images of spheres and that of the absolute conic, a linear solution has been derived from a recently proposed non-linear semi-definite approach. However, experiments show that this approach is quite sensitive to noise. In order to overcome this problem, a second approach has been proposed, where the orthogonal calibration relationship is obtained by regarding any two spheres as a surface of revolution. This allows a camera to be fully calibrated from an image of three spheres. Besides, a conic homography is derived from the imaged spheres, and from its eigenvectors the orthogonal invariants can be computed directly. Experiments on synthetic and real data show the practicality of such an approach.
Hui Zhang 0062, Guoqiang Zhang 0003, Kwan-Yee Kenneth Wong
ICIP (2)1