Xueying Qin

dblp:94/6767 · DBLP profile ↗
← Back
84ranked-venue papers
5as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 65 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 19 · 10 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 IntentionAR: An Intention-Driven Camera-Projector System for AR Assembly Guidance
abstract
Augmented reality (AR) assembly guidance systems can help users quickly master the assembly process of unfamiliar objects. However, it is very difficult for practical use without a thorough understanding of users' intentions. Furthermore, existing approaches still struggle to handle complex hand-part interactions and nonlinear assembly steps, and achieving intuitive and real-time guidance remains challenging. To address these issues, we propose an intention-driven camera-projector AR assembly guidance system (IntentionAR) that integrates online user intention inference with a finite-state assembly machine. The intention module recognizes interaction actions and infers the target object, enabling feedback before assembly begins, while the finite state machine (FSM) manages step progression. Using a camera-projector device, the system highlights candidate parts and, based on inferred intention, flags correct/incorrect selections to provide real-time guidance. As the assembly progresses, the models used for spatial registration and visualization switch dynamically to accommodate the nonlinear workflow. Experiments and user studies show that the system delivers a more robust AR interaction experience and improves assembly efficiency. A free copy of this paper and all supplemental materials, as well as project assets and source code, will be available at the project website.
Xin Cao 0010, Mingyu Ma 0013, Shanhao Yang, Kang Xie, Xibin Song, Fan Zhong 0001, Hai-Ning Liang, Xueying Qin
IEEE Trans. Vis. Comput. Graph.9
2025 One-shot 3D Object Canonicalization based on Geometric and Semantic Consistency
abstract
3D object canonicalization is a fundamental task, essential for various downstream tasks. Existing methods rely on either cumbersome manual processes or priors learned from extensive, per- category training samples. Real- world datasets, however, often exhibit long- tail distributions, challenging existing learning- based methods, especially in categories with limited samples. We address this by introducing the first one- shot category- level object canonicalization framework that operates under arbitrary poses, requiring only a single canonical model as a reference (the "prior model") for each category. To canonicalize any object, our framework first extracts semantic cues with large language models (LLMs) and vision- language models (VLMs) to establish correspondences with the prior model. We introduce a novel joint energy function to enforce geometric and semantic consistency, aligning object orientations precisely despite significant shape variations. Moreover, we adopt a support- plane strategy to reduce search space for initial poses and utilize a semantic relationship map to select the canonical pose from multiple hypotheses. Extensive experiments on multiple datasets demonstrate that our framework achieves state- of- the- art performance and validates key design choices. Using our framework, we create the Canonical Objaverse Dataset (COD), canonicalizing 32K samples in the Objaverse-LVIS dataset, underscoring the effectiveness of our framework on handling large- scale datasets. Project page at https://github.com/JinLi998/CanonObjaverseDataset
Wenzheng Chen, Qiyu Dai, Qingzhe Gao, Xueying Qin, Baoquan Chen
CVPR6
2025 Prior-free 3D Object Tracking
abstract
In this paper, we introduce a novel, truly prior-free 3D object tracking method that operates without given any model or training priors. Unlike existing methods that typically require pre-defined 3D models or specific training datasets as priors, which limit their applicability, our method is free from these constraints. Our method consists of a geometry generation module and a pose optimization module. Its core idea is to enable these two modules to automatically and iteratively enhance each other, thereby gradually building all the necessary information for the tracking task. We thus call the method as Bidirectional Iterative Tracking(BIT). The geometry generation module starts without priors and gradually generates high-precision mesh models for tracking, while the pose optimization module generates additional data during object tracking to further refine the generated models. Moreover, the generated 3D models can be stored and easily reused, allowing for seamless integration into various other tracking systems, not just our methods. Experimental results demonstrate that BIT outperforms many existing methods, even those that extensively utilize prior knowledge, while BIT does not rely on such information. Additionally, the generated 3D models deliver results comparable to actual 3D models, highlighting their superior and innovative qualities. The code is available at https://github.com/songxiuqiang/BIT.git.
Xiuqiang Song, Zhengxian Zhang, Fan Zhong 0001, Guofeng Zhang 0001, Xueying Qin
CVPR7
2025 A Visual Servo System for Robotic on-Orbit Servicing Based on 3D Perception of Non-Cooperative Satellite
abstract
The 3D perception of satellites, including both their shape and pose, is a key foundation for robotic on-orbit servicing. However, the demanding space environment-such as intense and dim illumination-presents significant challenges. Previous non-cooperative methods focus on specific geometric features like solar panel brackets or docking rings, overlooking the satellite's overall shape and increasing the risk of collisions during grasping. Additionally, satellites are often weakly textured, limiting the accuracy of 3D perception. To address these issues, we propose, for the first time, a 3D perceptionbased visual servo system of non-cooperative satellites. This system combines reconstruction and tracking to enhance shape perception and pose estimation accuracy in orbital conditions. Specifically, we employ an alternating iterative strategy to simultaneously reconstruct and track the satellite and introduce a novel constraint to fuse different cues under extreme conditions. Further, we develop a simulation environment platform, a dualarm microgravity grasping system, and an online monitoring module to enhance system capabilities for on-orbit servicing. Synthetic and real-world datasets from the simulation environment are also created for experimental validation. Results show that each module of our system achieves state-of-the-art performance.
Panpan Zhao, Yeheng Chen, Xiuqiang Song, Wenxuan Chen, Wenjuan Du, Xiangxu Meng, Xueying Qin
ICRA13
2025 RBTFusion: Region-Based Tracking for Real-Time Reconstruction from RGB-D Data
Shiyi Lu, Panpan Zhao, Xin Cao 0010, Xiaolong Yuan, Xueying Qin
ICXR6
2025 Robotic Arm Augmented Reality Control System
Mingyu Ma 0013, Xueying Qin
ICXR2
2025 3DOPT-SLAM: Robust SLAM via Joint Optimization with 3D Object Pose Tracking
Shanhao Yang, Xueying Qin, Xiangxu Meng
ICXR5
2025 DRMTrack: An Extended Distributed Millimeter-Wave Radar Framework for Indoor Multitarget Human Trajectory Tracking
abstract
Most researches on indoor target trajectory tracking using millimeter-wave radar have faced challenges such as interference from multipath effects, which led to corrupted point clouds, and large tracking errors in multi-target scenarios. This study aims to improve the accuracy of human multiple-object tracking, addressing the practical challenges of target trajectory tracking in indoor environments. We proposed an extended Distributed Radar Multi-Target Tracking (DRMTrack) framework that enabled the fusion of point clouds from multiple radar nodes. Additionally, by exploiting the spatial distribution characteristics of the target point clouds for clustering, tracking filtering and integrating historical data for tracking association, the framework enhances multiple-object trajectory tracking performance. Experimental results demonstrate that for various trajectory paths, the minimum position tracking error for multi-target is 5.2 cm. In the five-target tracking scenario, the 90th percentile tracking error is 18.8 cm, representing a 27.7% improvement in accuracy compared to single-radar tracking. The DRMTrack system effectively reduces interference from clutter point clouds, enhances multiple-object tracking precision, and supports real-time computation. This system is suitable for motion monitoring in indoor environments such as homes and hospital rooms, effectively fulfilling the needs of real-time, non-contact health management.
Guangqiang He, Weijie Wu, Hao Zhang 0118, Peng Wang 0115, Pang Wu, Zhenfeng Li, Xianxiang Chen, Lidong Du, Xueying Qin, Zhen Fang 0003, Yirong Wu
IEEE Internet Things J.9
2025 Relax! The Semilenient Core of Choreographic Programming (Functional Pearl)
abstract
The past few years have seen a surge of interest in choreographic programming, a programming paradigm for concurrent and distributed systems. The paradigm allows programmers to implement a distributed interaction protocol with a single high-level program, called a choreography, and then mechanically project it into correct implementations of its participating processes. A choreography can be expressed as a λ -term parameterized by constructors for creating data “at” a process and for communicating data between processes. Through this lens, recent work has shown how one can add choreographies to mainstream languages like Java, or even embed choreographies as a DSL in languages like Haskell and Rust. These new choreographic languages allow programmers to write in applicative style (like in functional programming) and write higher-order choreographies for better modularity. But the semantics of functional choreographic languages is not well-understood. Whereas typical λ -calculi can have their operational semantics defined with just a few rules, existing models for choreographic λ -calculi have dozens of complex rules and no clear or agreed-upon evaluation strategy . We show that functional choreographic programming is simple. Beginning with the Chor λ model from previous work, we strip away inessential features to produce a “core” model called λ χ . We discover that underneath Chor λ ’s apparently ad-hoc semantics lies a close connection to non-strict λ -calculi; we call the resulting evaluation strategy semilenient . Then, inspired by previous non-strict calculi, we develop a notion of choreographic evaluation contexts and a special commute rule to simplify and explain the unusual semantics of functional choreographic languages. The extra structure leads us to a presentation of λ χ with just ten rules, and a discovery of three missing rules in previous presentations of Chor λ . We also show how the extra structure comes with nice properties, which we use to simplify the correspondence proof between choreographies and their projections. Our model serves as both a principled foundation for functional choreographic languages and a good entry point for newcomers.
Dan Plyukhin, Xueying Qin, Fabrizio Montesi
Proc. ACM Program. Lang.2
2024 Semi-Decoupled 6D Pose Estimation via Multi-Modal Feature Fusion
abstract
The existing methods for 6D pose estimation based on RGB-D employ RGB images and observed point cloud derived from depth maps as input, then concurrently predicting both rotation and translation. However, rotation and translation possess distinct characteristics and scale ranges, and their simultaneous prediction can lead to mutual influence in the network parameter space. Additionally, the observed point cloud are susceptible to systematic noise and partial data loss, presenting challenges for the network to capture comprehensive object features. To address these issues, we propose the Semi-Decoupled 6D pose estimation via multi-modal feature fusion (SD6D). SD6D comprises a Multi-Modal Fusion Module and a Semi-Decoupled Prediction Module. The former dynamically fuses different modal data (RGB, depth, CAD model) based on their inter-modality correlations, aiding in establishing 2D-3D correspondences and addressing issues stemming from systematic noise and partial data loss. The latter semi-decouples the prediction of rotation and translation, predicting them separately based on their distinct characteristics. We conducted experiments on two popular benchmark datasets, which prove the superiority of our method.
Zhenhu Zhang, Xin Cao 0010, Xueying Qin, Ruofeng Tong 0001
ICASSP4
2024 CSS-Net: Domain Generalization in Category-level Pose Estimation via Corresponding Structural Superpoints
abstract
Category-level pose estimation is crucial for estimating the pose and size of unseen objects. Previous methods, mainly trained and tested on data with the same distribution, are limited in their ability to generalize to unseen domain data. For instance, when applied to new scenes or categories, frequent data collection and network training can be cumbersome. To address this issue, we propose a domain generalization method in category-level pose estimation based on structural superpoints, which is trained solely on simulated data and can generalize to unseen domain distributions in real datasets. Specifically, by extracting superpoints for structural correspondence in a self-supervised manner, our method achieves cross-domain data and cross-instance shape generalization. Accordingly, we designed a network and loss function, CoupleLoss, for regressing pose and size. Furthermore, we validated the effectiveness of our method on the wild6D and real275 datasets, achieving state-of-the-art results.
Xibin Song, Changhe Tu, Xueying Qin
ICME5
2024 Implicit Coarse-to-Fine 3D Perception for Category-level Object Pose Estimation from Monocular RGB Image
abstract
Category-level object pose estimation demonstrates robust generalization capabilities that benefit robotics applications. However, exclusive reliance on RGB images without leveraging any 3D information introduces ambiguity in the translation and size of objects, leading to suboptimal performance. In this paper, we propose a framework for category-level pose estimation from a single RGB image in an end-to-end manner, i.e., Feature Auxiliary Perception Network (FAP-Net). To address inaccurate pose estimation caused by the inherent ambiguity of RGB images, we design a coarse-to-fine approach that first harnesses geometry supervision to facilitate coarse 3D feature perception and subsequently refines the features based on pose and size constraints. Experimental results on REAL275 and CAMERA25 demonstrate that FAP-Net achieves significant improvements (14.7% on 10°10cm and 11.4% on IoU50 on the real-scene REAL275 dataset) over the state-of-the-art and real-time inference (42 FPS).
Xibin Song, Yeheng Chen, Xueying Qin
ICRA6
2024 Shoggoth: A Formal Foundation for Strategic Rewriting
abstract
Rewriting is a versatile and powerful technique used in many domains. Strategic rewriting allows programmers to control the application of rewrite rules by composing individual rewrite rules into complex rewrite strategies. These strategies are semantically complex, as they may be nondeterministic, they may raise errors that trigger backtracking, and they may not terminate. Given such semantic complexity, it is necessary to establish a formal understanding of rewrite strategies and to enable reasoning about them in order to answer questions like: How do we know that a rewrite strategy terminates? How do we know that a rewrite strategy does not fail because we compose two incompatible rewrites? How do we know that a desired property holds after applying a rewrite strategy? In this paper, we introduce Shoggoth: a formal foundation for understanding, analysing and reasoning about strategic rewriting that is capable of answering these questions. We provide a denotational semantics of System S, a core language for strategic rewriting, and prove its equivalence to our big-step operational semantics, which extends existing work by explicitly accounting for divergence. We further define a location-based weakest precondition calculus to enable formal reasoning about rewriting strategies, and we prove this calculus sound with respect to the denotational semantics. We show how this calculus can be used in practice to reason about properties of rewriting strategies, including termination, that they are well-composed, and that desired postconditions hold. The semantics and calculus are formalised in Isabelle/HOL and all proofs are mechanised.
Xueying Qin, Liam O'Connor, Rob J. van Glabbeek, Peter Höfner, Ohad Kammar, Michel Steuwer
Proc. ACM Program. Lang.1
2024 Corr-Track: Category-Level 6D Pose Tracking with Soft-Correspondence Matrix Estimation
abstract
Category-level pose tracking methods can continuously track the pose of objects without requiring any prior knowledge of the specific shape of the tracked instance. This makes them advantageous in augmented reality and virtual reality applications. The key challenge is how to train neural networks to accurately predict the poses of objects they have never seen before and exhibit strong generalization performance. We propose a novel category-level 6D pose tracking method Corr-Track, which is capable of accurately tracking objects belonging to the same category from depth video streams. Our approach utilizes direct soft correspondence constraints to train a neural network, which estimates bidirectional soft correspondences between sparsely sampled point clouds of objects in two frames. We first introduce a soft correspondence matrix for pose tracking tasks and establish effective constraints through direct spatial point-to-point correspondence representations in the sparse point cloud correspondence matrix. We propose the "point cloud expansion" strategy to address the "point cloud shrinkage" problem resulting from soft correspondences. This strategy ensures that the corresponding point cloud accurately reproduces the shape of the target point cloud, leading to precise pose tracking results. We evaluated our approach on the NOCS-REAL275 and Wild6D dataset and observed superior performance compared to previous methods. Additionally, we conducted cross-category experiments that further demonstrated its generalization capability.
Xin Cao 0010, Panpan Zhao, Xueying Qin
IEEE Trans. Vis. Comput. Graph.5
2023 Online Hand-Eye Calibration with Decoupling by 3D Textureless Object Tracking
abstract
Hand-eye calibration estimates the pose of a camera relative to a robot, which is a fundamental problem for visually guided robots, especially for dynamic object grasping. Most methods use 2D fiducial markers with distinctive visual features and require pre-calibration for accurate calibration, which can not work online. In this paper, we propose a novel hand-eye calibration method based on the natural 3D object, which can work online and automatically even if the object is textureless or weakly textured. We first propose a Pose Refinement Network (PR-Net) to improve the accuracy of 3D object tracking. Then we build a 3D convergence point constraint based on the multi-view information with the accurate object pose to adjust the object position. Finally, we optimize the hand-eye pose by the closed-loop constraint with the optimized object position, solving the problem that is easy to fall into a local minimum. The experiments show that the average error of our hand-eye calibration method is 1.20 degrees and 23.18 mm. The results achieve state-of-the-art by using the working object to realize the online hand-eye calibration.
Kang Xie, Wenxuan Chen, Xin Cao 0010, Jiankai Qian, Xueying Qin
ICRA8
2023 Minilag Filter for Jitter Elimination of Pose Trajectory in AR Environment
abstract
In AR applications, the jitter of virtual objects can weaken the sense of integration with the real environment. This jitter is often caused by noise in the pose obtained by 3D tracking or localization methods, especially in monocular vision systems without IMU support. Filtering the pose is an effective method to eliminate jitter, however, it can also cause significant lag in the filtered pose, seriously degrading the AR experience. Existing filters struggle to simultaneously reduce jitter while maintaining low lag. In this paper, we propose a novel Minilag filter, which achieves excellent pose smoothing while significantly reducing the lag through backtracking update and compensation strategies, and has excellent real-time performance. We represent the rotation in the pose in the Lie algebra and filter it in locally Euclidean space, ensuring that the filtering of rotation is consistent with that of vectors. We also analyze the noise distribution and characteristics in the tracked pose, providing a theoretical basis for setting filter parameters. We evaluated the proposed filter using both objective mathematical metrics and a user study, and the experimental results demonstrate that our method achieves state-of-the-art performance.
Xiuqiang Song, Weijian Xie, Nan Wang 0020, Fan Zhong 0001, Guofeng Zhang 0001, Xueying Qin
ISMAR7
2023 Accompany Children's Learning for You: An Intelligent Companion Learning System
abstract
Abstract Nowadays, parents attach importance to their children's primary education but often lack time and correct pedagogical principles to accompany their children's learning. Besides, existing learning systems cannot perceive children's emotional changes. They may also cause children's self‐control and cognitive problems due to smart devices such as mobile phones and tablets. To tackle these issues, we propose an intelligent companion learning system to accompany children in learning English words, namely theIntelligent Augmented Reality Educator (IARE). The IARE realizes the perception and feedback of children's engagement through the intelligent agent (IA) module, and presents the humanized interaction based on projective Augmented Reality (AR). Specifically, IA perceives the children's learning engagement change and spelling status in real‐time through our online lightweight temporal multiple instance attention module and character recognition module, based on which analyses the performance of the individual learning process and gives appropriate feedback and guidance. We allow children to interact with physical letters, thus avoiding the excessive interference of electronic devices. To test the efficacy of our system, we conduct a pilot study with 14 English learning children. The results show that our system can significantly improve children's intrinsic motivation and self‐efficacy.
Jiankai Qian, Xinbo Jiang, Jiayao Ma 0001, Zhenzhen Gao, Xueying Qin
Comput. Graph. Forum6
2023 3D Object Tracking for Rough Models
abstract
Abstract Visual monocular 6D pose tracking methods for textureless or weakly‐textured objects heavily rely on contour constraints established by the precise 3D model. However, precise models are not always available in reality, and rough models can potentially degrade tracking performance and impede the widespread usage of 3D object tracking. To address this new problem, we propose a novel tracking method that handles rough models. We reshape the rough contour through the probability map, which can avoid explicitly processing the 3D rough model itself. We further emphasize the inner region information of the object, where the points are sampled to provide color constrains. To sufficiently satisfy the assumption of small displacement between frames, the 2D translation of the object is pre‐searched for a better initial pose. Finally, we combine constraints from both the contour and inner region to optimize the object pose. Experimental results demonstrate that the proposed method achieves state‐of‐the‐art performance on both roughly and precisely modeled objects. Particularly for the highly rough model, the accuracy is significantly improved (40.4% v.s. 16.9%).
Xiuqiang Song, Weijian Xie, Nan Wang 0020, Fan Zhong 0001, Guofeng Zhang 0001, Xueying Qin
Comput. Graph. Forum7
2023 Exploring Gaze-assisted and Hand-based Region Selection in Augmented Reality
abstract
Region selection is a fundamental task in interactive systems. In 2D user interfaces, users typically use a rectangle selection tool to formulate a region using a mouse or touchpad. Region selection in 3D spaces, especially in Augmented Reality (AR) Head-Mounted Displays (HMDs) is different and challenging because users need to select an intended region via freehand mid-air gestures or eye-based actions that are touchless interactions. In this work, we aim to fill in the gap in the design of region selection techniques in AR HMDs. We first analyzed and discretized the interaction procedure of region selection and explored design possibilities for each step. We then developed four techniques for region selection in AR HMDs, which leveraged users' hand and gaze for unimodal or multimodal interaction. The techniques were evaluated via a user study with a controlled region selection task. The findings led to three design recommendations and two proof-of-concept application examples.
Rongkai Shi, Yushi Wei, Xueying Qin, Pan Hui 0001, Hai-Ning Liang
Proc. ACM Hum. Comput. Interact.3
2023 Guided Linear Upsampling
abstract
Guided upsampling is an effective approach for accelerating high-resolution image processing. In this paper, we propose a simple yet effective guided upsampling method. Each pixel in the high-resolution image is represented as a linear interpolation of two low-resolution pixels, whose indices and weights are optimized to minimize the upsampling error. The downsampling can be jointly optimized in order to prevent missing small isolated regions. Our method can be derived from the color line model and local color transformations. Compared to previous methods, our method can better preserve detail effects while suppressing artifacts such as bleeding and blurring. It is efficient, easy to implement, and free of sensitive parameters. We evaluate the proposed method with a wide range of image operators, and show its advantages through quantitative and qualitative analysis. We demonstrate the advantages of our method for both interactive image editing and real-time high-resolution video processing. In particular, for interactive editing, the joint optimization can be precomputed, thus allowing for instant feedback without hardware acceleration.
Shuangbing Song, Fan Zhong 0001, Tianju Wang, Xueying Qin, Changhe Tu
ACM Trans. Graph.4
2022 BCOT: A Markerless High-Precision 3D Object Tracking Benchmark
abstract
Template-based 3D object tracking still lacks a high-precision benchmark of real scenes due to the difficulty of annotating the accurate 3D poses of real moving video objects without using markers. In this paper, we present a multi-view approach to estimate the accurate 3D poses of real moving objects, and then use binocular data to construct a new benchmark for monocular textureless 3D object tracking. The proposed method requires no markers, and the cameras only need to be synchronous, relatively fixed as cross-view and calibrated. Based on our object-centered model, we jointly optimize the object pose by minimizing shape reprojection constraints in all views, which greatly improves the accuracy compared with the single-view approach, and is even more accurate than the depth-based method. Our new benchmark dataset contains 20 textureless objects, 22 scenes, 404 video sequences and 126K images captured in real scenes. The annotation error is guaranteed to be less than 2mm, according to both theoretical analysis and validation experiments. We reevaluate the state-of-the-art 3D object tracking methods with our dataset, reporting their performance ranking in real scenes. Our BCOT benchmark and code can be found at https://ar3dv.github.io/BCOT-Benchmark/.
Bin Wang 0035, Shiqiang Zhu, Xin Cao 0010, Fan Zhong 0001, Wenxuan Chen, Jason Gu, Xueying Qin
CVPR9
2022 Large-Displacement 3D Object Tracking with Hybrid Non-local Optimization
Xuhui Tian, Xinran Lin, Fan Zhong 0001, Xueying Qin
ECCV (22)4
2022 A hierarchical model for learning to understand head gesture videos
Songhua Xu, Xueying Qin
Pattern Recognit.3
2022 Pixel-Wise Weighted Region-Based 3D Object Tracking Using Contour Constraints
abstract
Region-based methods are currently achieving state-of-the-art performance for monocular 3D object tracking. However, they are still prone to fail in cases of partial occlusions and ambiguous colors. We propose a novel region-based method to tackle these problems. The key idea is to derive a pixel-wise weighted region-based cost function using contour constraints. First, we propose a novel region-based cost function using search lines around the object contour, which is more efficient than previous region-based cost functions using signed distance transform, and in the meantime can deal with partial occlusions and ambiguous colors more effectively. Second, we propose an optimal searching strategy to search the object contour points in cluttered scenes, and then use the object contour points to detect partial occlusions and ambiguous colors. Third, we propose a pixel-wise weight function based on color and distance constraints of the object contour points, and integrate it into the proposed region-based cost function to reduce the negative impact of partial occlusions and ambiguous colors. We verify the effectiveness and efficiency of our method on challenging public datasets. Experiments demonstrate that our method outperforms the recent state-of-the-art region-based methods in complex scenarios, especially in the presence of partial occlusions and ambiguous colors.
Fan Zhong 0001, Xueying Qin
IEEE Trans. Vis. Comput. Graph.3
2021 MFE: Multi-scale Feature Enhancement for Object Detection
Zhenhu Zhang, Xueying Qin, Fan Zhong 0001
BMVC2
2021 Hierarchical Temporal Multi-Instance Learning for Video-based Student Learning Engagement Assessment
abstract
Video-based automatic assessment of a student's learning engagement on the fly can provide immense values for delivering personalized instructional services, a vehicle particularly important for massive online education. To train such an assessor, a major challenge lies in the collection of sufficient labels at the appropriate temporal granularity since a learner's engagement status may continuously change throughout a study session. Supplying labels at either frame or clip level incurs a high annotation cost. To overcome such a challenge, this paper proposes a novel hierarchical multiple instance learning (MIL) solution, which only requires labels anchored on full-length videos to learn to assess student engagement at an arbitrary temporal granularity and for an arbitrary duration in a study session. The hierarchical model mainly comprises a bottom module and a top module, respectively dedicated to learning the latent relationship between a clip and its constituent frames and that between a video and its constituent clips, with the constraints on the training stage that the average engagements of local clips is that of the video label. To verify the effectiveness of our method, we compare the performance of the proposed approach with that of several state-of-the-art peer solutions through extensive experiments.
Jiayao Ma 0001, Xinbo Jiang, Songhua Xu, Xueying Qin
IJCAI4
2021 Fast 3D texture-less object tracking with geometric contour and local region
Xiuqiang Song, Fan Zhong 0001, Xueying Qin
Comput. Graph.4
2021 3D Object Tracking with Adaptively Weighted Local Bundles
Jia-Chen Li, Fan Zhong 0001, Songhua Xu, Xueying Qin
J. Comput. Sci. Technol.4
2021 ReLoc: Indoor Visual Localization with Hierarchical Sitemap and View Synthesis
Hui-Xuan Wang, Jing-Liang Peng, Shi-Yi Lu, Xin Cao 0010, Xueying Qin, Changhe Tu
J. Comput. Sci. Technol.5
2020 Illumination Harmonization with Gray Mean Scale
Shuangbing Song, Fan Zhong 0001, Xueying Qin, Changhe Tu
CGI3
2020 An Occlusion-aware Edge-Based Method for Monocular 3D Object Tracking using Edge Confidence
abstract
Abstract We propose an edge‐based method for 6DOF pose tracking of rigid objects using a monocular RGB camera. One of the critical problem for edge‐based methods is to search the object contour points in the image corresponding to the known 3D model points. However, previous methods often produce false object contour points in case of cluttered backgrounds and partial occlusions. In this paper, we propose a novel edge‐based 3D objects tracking method to tackle this problem. To search the object contour points, foreground and background clutter points are first filtered out using edge color cue, then object contour points are searched by maximizing their edge confidence which combines edge color and distance cues. Furthermore, the edge confidence is integrated into the edge‐based energy function to reduce the influence of false contour points caused by cluttered backgrounds and partial occlusions. We also extend our method to multi‐object tracking which can handle mutual occlusions. We compare our method with the recent state‐of‐art methods on challenging public datasets. Experiments demonstrate that our method improves robustness and accuracy against cluttered backgrounds and partial occlusions.
Fan Zhong 0001, Yuqing Sun 0001, Xueying Qin
Comput. Graph. Forum4
2020 Achieving high-performance the functional way: a functional pearl on expressing high-performance optimizations as rewrite strategies
abstract
Optimizing programs to run efficiently on modern parallel hardware is hard but crucial for many applications. The predominantly used imperative languages - like C or OpenCL - force the programmer to intertwine the code describing functionality and optimizations. This results in a portability nightmare that is particularly problematic given the accelerating trend towards specialized hardware devices to further increase efficiency. Many emerging DSLs used in performance demanding domains such as deep learning or high-performance image processing attempt to simplify or even fully automate the optimization process. Using a high-level - often functional - language, programmers focus on describing functionality in a declarative way. In some systems such as Halide or TVM, a separate schedule specifies how the program should be optimized. Unfortunately, these schedules are not written in well-defined programming languages. Instead, they are implemented as a set of ad-hoc predefined APIs that the compiler writers have exposed. In this functional pearl, we show how to employ functional programming techniques to solve this challenge with elegance. We present two functional languages that work together - each addressing a separate concern. RISE is a functional language for expressing computations using well known functional data-parallel patterns. ELEVATE is a functional language for describing optimization strategies. A high-level RISE program is transformed into a low-level form using optimization strategies written in ELEVATE . From the rewritten low-level program high-performance parallel code is automatically generated. In contrast to existing high-performance domain-specific systems with scheduling APIs, in our approach programmers are not restricted to a set of built-in operations and optimizations but freely define their own computational patterns in RISE and optimization strategies in ELEVATE in a composable and reusable way. We show how our holistic functional approach achieves competitive performance with the state-of-the-art imperative systems Halide and TVM.
Bastian Hagedorn, Johannes Lenfers, Thomas Koehler 0005, Xueying Qin, Sergei Gorlatch, Michel Steuwer
Proc. ACM Program. Lang.4
2019 Robust edge-based 3D object tracking with direction-based pose validation
Bin Wang 0035, Fan Zhong 0001, Xueying Qin
Multim. Tools Appl.3
2019 Deeply Supervised Depth Map Super-Resolution as Novel View Synthesis
abstract
Deep convolutional neural network (DCNN) has been successfully applied to depth map super-resolution and outperforms existing methods by a wide margin. However, there still exist two major issues with these DCNN-based depth map super-resolution methods that hinder the performance: 1) the low-resolution depth maps either need to be up-sampled before feeding into the network or substantial deconvolution has to be used and 2) the supervision (high-resolution depth maps) is only applied at the end of the network, thus it is difficult to handle large up-sampling factors, such as x8 and x16. In this paper, we propose a new framework to tackle the above problems. First, we propose to represent the task of depth map superresolution as a series of novel view synthesis sub-tasks. The novel view synthesis sub-task aims at generating (synthesizing) a depth map from a different camera pose, which could be learned in parallel. Second, to handle large up-sampling factors, we present a deeply supervised network structure to enforce strong supervision in each stage of the network. Third, a multiscale fusion strategy is proposed to effectively exploit the feature maps at different scales and handle the blocking effect. In this way, our proposed framework could deal with challenging depth map super-resolution efficiently under large up-sampling factors (e.g., x8 and x16). Our method only uses the low-resolution depth map as input, and the support of color image is not needed, which greatly reduces the restriction of our method. Extensive experiments on various benchmarking data sets demonstrate the superiority of our method over current state-of-the-art depth map super-resolution methods.
Xibin Song, Yuchao Dai, Xueying Qin
IEEE Trans. Circuits Syst. Video Technol.3
2018 Sparsely Grouped Multi-Task Generative Adversarial Networks for Facial Attribute Manipulation
abstract
Recently, Image-to-Image Translation (IIT) has achieved great progress in image style transfer and semantic context manipulation for images. However, existing approaches require exhaustively labelling training data, which is labor demanding, difficult to scale up, and hard to adapt to a new domain. To overcome such a key limitation, we propose Sparsely Grouped Generative Adversarial Networks (SG-GAN) as a novel approach that can translate images in sparsely grouped datasets where only a few train samples are labelled. Using a one-input multi-output architecture, SG-GAN is well-suited for tackling multi-task learning and sparsely grouped learning tasks. The new model is able to translate images among multiple groups using only a single trained model. To experimentally validate the advantages of the new model, we apply the proposed method to tackle a series of attribute manipulation tasks for facial images as a case study. Experimental results show that SG-GAN can achieve comparable results with state-of-the-art methods on adequately labelled datasets while attaining a superior image translation quality on sparsely grouped datasets~\footnoteCode is available at https://github.com/zhangqianhui/SGGAN-tensorflow..
Jichao Zhang, Yezhi Shu, Songhua Xu, Gongze Cao, Fan Zhong 0001, Meng Liu 0006, Xueying Qin
ACM Multimedia7
2018 Active Assembly Guidance with Online Video Parsing
abstract
In this paper, we introduce an online video-based system that actively assists users in assembly tasks. The system guides and monitors the assembly process by providing instructions and feedback on possibly erroneous operations, enabling easy and effective guidance in AR/MR applications. The core of our system is an online video-based assembly parsing method that can understand the assembly process, which is known to be extremely hard previously. Our method exploits the availability of the participating parts to significantly alleviate the problem, reducing the recognition task to an identification problem, within a constrained search space. To further constrain the search space, and understand the observed assembly activity, we introduce a tree-based global-inference technique. Our key idea is to incorporate part-interaction rules as powerful constraints which significantly regularize the search space and correctly parse the assembly video at interactive rates. Complex examples demonstrate the effectiveness of our method.
Bin Wang 0035, Andrei Sharf, Yangyan Li, Fan Zhong 0001, Xueying Qin, Daniel Cohen-Or, Baoquan Chen
VR6
2018 Accurate and fast 3D head pose estimation with noisy RGBD images
Fan Zhong 0001, Xueying Qin
Multim. Tools Appl.4
2018 Modeling deviations of rgb-d cameras for accurate depth map and color image registration
Xibin Song, Jianmin Zheng, Fan Zhong 0001, Xueying Qin
Multim. Tools Appl.4
2017 ST-GAN: Unsupervised Facial Image Semantic Transformation Using Generative Adversarial Networks
abstract
Image semantic transformation aims to convert one image into another image with different semantic features (e.g., face pose, hairstyle). The previous methods, which learn the mapping function from one image domain to the other, require supervised information directly or indirectly. In this paper, we propose an unsupervised image semantic transformation method called semantic transformation generative adversarial networks (ST-GAN), and experimentally verify it on face dataset. We further improve ST-GAN with the Wasserstein distance to generate more realistic images and propose a method called local mutual information maximization to obtain a more explicit semantic transformation. ST-GAN has the ability to map the image semantic features into the latent vector and then perform transformation by controlling the latent vector.
Jichao Zhang, Fan Zhong 0001, Gongze Cao, Xueying Qin
ACML4
2017 Pose optimization in edge distance field for textureless 3D object tracking
abstract
This paper presents a monocular model-based 3D tracking approach for textureless objects. Instead of explicitly searching for 3D-2D correspondences as previous methods, which unavoidably generates individual outlier matches, we aim to minimize the holistic distance between the predicted object contour and the query image edges. We propose a method that can directly solve 3D pose parameters in unsegmented edge distance field. We derive the differentials of edge matching distance with respect to the pose parameters, and search the optimal 3D pose parameters using standard gradient-based non-linear optimization techniques. To avoid being trapped in local minima and to deal with potential large inter-frame motions, a particle filtering process with a first order autoregressive state dynamics is exploited. Occlusions are handled by a robust estimator. The effectiveness of our approach is demonstrated using comparative experiments on real image sequences with occlusions, large motions and cluttered backgrounds.
Bin Wang 0035, Fan Zhong 0001, Xueying Qin
CGI3
2017 Realistic image composite with best-buddy prior of natural image patches
abstract
Realistic image composite requires the appearance of foreground and background layers to be consistent. This is difficult to achieve because the foreground and the background may be taken from very different environments. This paper proposes a novel composite adjustment method that can harmonize appearance of different composite layers. We introduce the Best-Buddy Prior (BBP), which is a novel compact representations of the joint co-occurrence distribution of natural image patches. BBP can be learned from unlabelled images given only the unsupervised regional segmentation. The most-probable adjustment of foreground can be estimated efficiently in the BBP space as the shift vector to the local maximum of density function. Both qualitative and quantitative evaluations show that our method outperforms previous composite adjustment methods.
Yuan Wang 0085, Fan Zhong 0001, Xueying Qin
ICIP4
2016 Deep Depth Super-Resolution: Learning Depth Super-Resolution Using Deep Convolutional Neural Network
Xibin Song, Yuchao Dai, Xueying Qin
ACCV (4)3
2016 Edge-guided depth map enhancement
abstract
Low-cost depth sensing devices, such as Microsoft Kinect, can only produce noisy depth maps that are mis-aligned with color images, and even contain many holes. Even though the coupled high quality color images contain rich information which can be exploited to enhance the depth maps, the redundant color edges often introduce incorrect depth edges in the result depth map, since color images contain more textures than depth maps. To solve this problem, we propose a novel approach which generates accurate color-consistent depth edges by employing both color and depth images. First, Edges of raw depth maps are extracted using image pyramid strategy. Then, the redundant edges in color images are removed according to the raw depth edges, and, accurate color-consistent depth edges are generated by combining raw depth edges with current color edges. Finally, constraints extracted from both raw depth and color images and the generated depth edges are fused in a MRF optimization framework to obtain the enhanced depth map, which is accurately aligned with coupled color image. As experimentally demonstrated, the proposed method achieves outstanding performance when compared with previous approaches.
Xibin Song, Fan Zhong 0001, Xueying Qin
ICPR5
2016 Action recognition based on global optimal similarity measuring
Xinbo Jiang, Fan Zhong 0001, Qunsheng Peng 0001, Xueying Qin
Multim. Tools Appl.4
2015 Enlarging Image by Constrained Least Square Approach with Shape Preserving
Fan Zhang 0045, Xin Zhang 0079, Xueying Qin, Caiming Zhang 0001
J. Comput. Sci. Technol.3
2015 Visual Tracking via Sparse and Local Linear Coding
abstract
The state search is an important component of any object tracking algorithm. Numerous algorithms have been proposed, but stochastic sampling methods (e.g., particle filters) are arguably one of the most effective approaches. However, the discretization of the state space complicates the search for the precise object location. In this paper, we propose a novel tracking algorithm that extends the state space of particle observations from discrete to continuous. The solution is determined accurately via iterative linear coding between two convex hulls. The algorithm is modeled by an optimal function, which can be efficiently solved by either convex sparse coding or locality constrained linear coding. The algorithm is also very flexible and can be combined with many generic object representations. Thus, we first use sparse representation to achieve an efficient searching mechanism of the algorithm and demonstrate its accuracy. Next, two other object representation models, i.e., least soft-threshold squares and adaptive structural local sparse appearance, are implemented with improved accuracy to demonstrate the flexibility of our algorithm. Qualitative and quantitative experimental results demonstrate that the proposed tracking algorithm performs favorably against the state-of-the-art methods in dynamic scenes.
Xueying Qin, Fan Zhong 0001, Hongbo Li 0012, Qunsheng Peng 0001, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.2
2015 Global optimal searching for textureless 3D object tracking
Bin Wang 0035, Fan Zhong 0001, Xueying Qin, Baoquan Chen
Vis. Comput.4
2014 Slippage-free background replacement for hand-held video
abstract
We introduce a method for replacing the background in a video of a moving foreground subject, when both the source video capturing the subject, and the target video capturing the new background scene, are natural videos, casually captured using a freely moving hand-held camera. We assume that the foreground subject has already been extracted, and focus on the challenging task of generating a video with a new background, such that the new background motion appears compatible with the original one. Failure to match the motion results in disturbing slippage or moonwalk artifacts, where the subject's feet appear to slide or slip over the ground. While matching the motion across the entire frame is impossible for scenes with differing geometry, we aim to match the local motion of the ground in the vicinity of the subject. This is achieved by reordering and warping the available target background frames in a manner that optimizes a suitably designed objective function.
Fan Zhong 0001, Xueying Qin, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen
ACM Trans. Graph.3
2014 Online robust action recognition based on a hierarchical model
Xinbo Jiang, Fan Zhong 0001, Qunsheng Peng 0001, Xueying Qin
Vis. Comput.4
2014 Estimation of Kinect depth confidence through self-training
Xibin Song, Fan Zhong 0001, Yanke Wang, Xueying Qin
Vis. Comput.4
2014 Depth map enhancement based on color and depth consistency
Yanke Wang, Fan Zhong 0001, Qunsheng Peng 0001, Xueying Qin
Vis. Comput.4
2013 Visual Tracking via Subspace Motion Model
abstract
The art of visual tracking has been widely studied in the past decades [6]. While most of researches focus on exploring new methods to represent object appearance, little attention has been paid on the description of object motion. In this paper we propose a novel motion model for visual tracking, and in comparison with previous methods, it can better parameterize instantaneous image motion caused by both object and camera movements. Our approach is inspired by the subspace theory of image motion, that is, for a rigid object imaged by a projective camera, the displacements matrix of its trajectories over a short period of time should approximately lie in a low-dimensional subspace with a certain rank upper bound [2, 5]. We adopt this subspace as the state transition space in particle filtering (PF) [3]. This differs from affine model in two ways: first, the dimension number as well as the sampling weight for each dimension at each moment can be determined by the rank of the subspace automatically; second, the subspace motion model can naturally represent the disparity brought by object or camera rotation. We will show that when compared with the affine model, the subspace motion model is superior in accuracy. Figure 1 illustrates the procedures of our method. To estimate the motion model, some 2D feature points of the object are first tracked by the standard KLT approach [4]. Assuming that k successive frames have been tracked before the current frame It , then the displacements matrix can be built as:
Fan Zhong 0001, Qunsheng Peng 0001, Xueying Qin
BMVC5
2013 Robust Action Recognition Based on a Hierarchical Model
abstract
With the strong demand for human machine interaction, action recognition has attracted more and more attention in recent years. Traditional video-based approaches are very sensitive to background activity, and also lack the ability to discriminate complex 3D motion. With the emergence and development of commercial depth cameras, action recognition based on 3D skeleton joints is becoming more and more popular. However, a skeleton-based approach is still very challenging because of the large variation in human actions and temporal dynamics. In this paper, we propose a hierarchical model for action recognition. To handle confusing motions in a large feature space, a motion-based grouping method is first proposed, which can efficiently assign each video a group label, and then for each group, a pre-trained classifier is used for frame-labeling. Unlike previous methods, we adopt a bottom-up approach that first performs action recognition for each frame. The final action label is obtained by fusing the classification to its frames, with the effect of each frame being adaptively adjusted based on its local properties. The proposed method is evaluated using two challenge datasets captured by a Kinect. Experiments show that our method can perform more robustly than state-of-the-art approaches.
Xinbo Jiang, Fan Zhong 0001, Qunsheng Peng 0001, Xueying Qin
CW4
2013 Lighting Simulation of Augmented Outdoor Scene Based on a Legacy Photograph
abstract
Abstract We propose a novel approach to simulate the illumination of augmented outdoor scene based on a legacy photograph. Unlike previous works which only take surface radiosity or lighting related prior information as the basis of illumination estimation, our method integrates both of these two items. By adopting spherical harmonics, we deduce a linear model with only six illumination parameters. The illumination of an outdoor scene is finally calculated by solving a linear least square problem with the color constraint of the sunlight and the skylight. A high quality environment map is then set up, leading to realistic rendering results. We also explore the problem of shadow casting between real and virtual objects without knowing the geometry of objects which cast shadows. An efficient method is proposed to project complex shadows (such as tree's shadows) on the ground of the real scene to the surface of the virtual object with texture mapping. Finally, we present an unified scheme for image composition of a real outdoor scene with virtual objects ensuring their illumination consistency and shadow consistency. Experiments demonstrate the effectiveness and flexibility of our method.
Guanyu Xing, Xuehong Zhou, Qunsheng Peng 0001, Yanli Liu 0002, Xueying Qin
Comput. Graph. Forum5
2013 Online illumination estimation of outdoor scenes based on videos containing no shadow area
Guanyu Xing, Xuehong Zhou, Yanli Liu 0002, Xueying Qin, Qunsheng Peng 0001
Sci. China Inf. Sci.4
2013 Cylindrical panoramic mosaicing from a pipeline video through MRF based optimization
Chuan Niu, Fan Zhong 0001, Songhua Xu, Chenglei Yang, Xueying Qin
Vis. Comput.5
2013 Basis image decomposition of outdoor time-lapse videos
Fan Zhong 0001, Lili Lin, Guanyu Xing, Qunsheng Peng 0001, Xueying Qin
Vis. Comput.6
2012 Visual Tracking in Continuous Appearance Space via Sparse Coding
Fan Zhong 0001, Qunsheng Peng 0001, Xueying Qin
ACCV (3)5
2012 Decomposition Equation of Basis Images with Consideration of Global Illumination
Xueying Qin, Lili Lin, Fan Zhong 0001, Guanyu Xing, Qunsheng Peng 0001
CVM1
2012 A practical approach for real-time illumination estimation of outdoor videos
Guanyu Xing, Yanli Liu 0002, Xueying Qin, Qunsheng Peng 0001
Comput. Graph.3
2012 Discontinuity-aware video object cutout
abstract
Existing video object cutout systems can only deal with limited cases. They usually require detailed user interactions to segment real-life videos, which often suffer from both inseparable statistics (similar appearance between foreground and background) and temporal discontinuities (e.g. large movements, newly-exposed regions following disocclusion or topology change). In this paper, we present an efficient video cutout system to meet this challenge. A novel directional classifier is proposed to handle temporal discontinuities robustly, and then multiple classifiers are incorporated to cover a variety of cases. The outputs of these classifiers are integrated via another classifier, which is learnt from real examples. The foreground matte is solved by a coherent matting procedure, and remaining errors can be removed easily by additive spatio-temporal local editing. Experiments demonstrate that our system performs more robustly and more intelligently than existing systems in dealing with various input types, thus saving a lot of user labor and time.
Fan Zhong 0001, Xueying Qin, Qunsheng Peng 0001, Xiangxu Meng
ACM Trans. Graph.2
2011 Creating Cylindrical Panoramic Mosaic from a Pipeline Video
abstract
In geological engineering, stratum structure detection is a fundamental problem in project planning and implementation. One of the most commonly employed detection technologies is to take videos of borehole using a forward moving camera. Following this approach, the problem of stratum structure detection is transformed into the problem of constructing a panoramic image from the taken video sequences, which are typically in low quality. In this paper, we propose a novel method to create a panoramic image of the borehole from the video sequence without camera calibration and tracking. To stitch together pixels of neighboring frame images, our camera model is designed with a focal length changing feature, along with a small rotation freedom in the two-dimensional image space. Essentially, our camera model assumes target objects lie on a cylindrical wall and the camera moves forward along the central axis of the cylindrical wall. Our method robustly resolves these two degrees-of-freedoms through KLT feature tracking and constructs a panoramic image by stitching strips. Experiment results show that our method could efficiently generate high-quality panoramas for very long video sequences.
Chuan Niu, Fan Zhong 0001, Songhua Xu, Chenglei Yang, Xueying Qin
CAD/Graphics5
2011 On-line Illumination Estimation of Outdoor Scenes Based on Area Selection for Augmented Reality
abstract
In augmented reality, consistent illumination plays an important role when integrating a virtual object into a video of real scene. In this paper, we propose a novel image based framework to estimate-on-line the dynamic ally changing illumination parameters of outdoor video sequences captured by a fixed camera. Unlike previous approaches which either request to know the scene geometry or involve huge storage to preserve time dependent basis images or statistic parameters, our approach requires very simple interaction at the initialization stage by a few brushes to select areas with specified surface normal, which are used to calculate the sunlight parameters. An optimization procedure is also applied, ensuring the robustness and precision of our estimation. Experimental results demonstrate the effectiveness and flexibility of the proposed approach.
Guanyu Xing, Yanli Liu 0002, Xueying Qin, Qunsheng Peng 0001
CAD/Graphics3
2011 A Local Behavior Model for Small Pedestrian Groups
abstract
Simulating the local behavior of small pedestrian groups among a crowd is an emerging problem. Small groups are commonly found in general crowds and their behavior are significantly distinguished from that of individual pedestrians. In this paper, we propose a novel approach to simulate the walking behavior of such small groups. We construct dynamic group formations which not only facilitate easy communication between members within the same group but also adapt to the environment constraint. Moreover, an extended collision avoidance algorithm based on Optimal Reciprocal Collision Avoidance (ORCA) is proposed taking into account the factors of group cohesion and reaction flexibility so as to generate diverse interaction behavior between walkers. Several examples demonstrate the efficiency of the proposed algorithm.
Yijiang Zhang, Julien Pettré, Xueying Qin, Stéphane Donikian, Qunsheng Peng 0001
CAD/Graphics3
2011 Online inserting virtual characters into dynamic video scenes
abstract
ABSTRACT The seamless integration of virtual characters into dynamic scenes captured by video is a challenging problem. In order to achieve consistent composite results, both the virtual and real characters must share the same geometrical constraints and their interactions must follow some common sense. One essential question is how to detect the motion of real objects—such as real characters moving in the video—and how to steer virtual characters accordingly to avoid unrealistic collisions. We propose an online solution. First, by analysis of the input video, the motion states of the real pedestrians are recovered into a common world 3D coordinate system. Meanwhile, a simplified accuracy measurement is defined to represent the confidence of the motion estimate. Then, under the constraints imposed by the real dynamic objects, the motion of virtual characters are accommodated by a uniform steering model. The final step is to merge virtual objects back to the real video scene by taking into account visibility and occlusion constraints between real foreground objects and virtual ones. Several examples demonstrate the efficiency of the proposed algorithm. Copyright © 2011 John Wiley & Sons, Ltd.
Yijiang Zhang, Julien Pettré, Jan Ondrej, Xueying Qin, Qunsheng Peng 0001, Stéphane Donikian
Comput. Animat. Virtual Worlds4
2011 Robust image segmentation against complex color distribution
Fan Zhong 0001, Xueying Qin, Qunsheng Peng 0001
Vis. Comput.2
2010 Transductive segmentation of live video with non-stationary background
abstract
Online foreground extraction is very difficult due to the complexity of real scenes. Almost all the previous methods assume that the background is stationary, which not only incur unreliable result due to background activities like dynamic shadow, moving background objects etc., but also makes them hard to be extended to the case of non-stationary background. In this paper we assume that the background is continuous instead of stationary, and present a transductive video segmentation method that can handle dynamic scenes captured by a hand-held moving camera. The segmentation is propagated based on local color models and temporal prior, as well as a dynamic global color model (DGKDE) in the case of occlusion. A novel local color modeling method, FLKDE, is proposed to model both local color distribution and temporal prior at each pixel. FLKDE can be learned additively to reach real-time speed. Finally, a very fast geodesic-based method is adopted to solve for the segmentation. Experiments show that our method can generate good quality segmentation for wide variety of scenes, and can reach 15~25 fps for 640 × 480 size of input image sequences.
Fan Zhong 0001, Xueying Qin, Qunsheng Peng 0001
CVPR2
2010 A new approach to outdoor illumination estimation based on statistical analysis for augmented reality
abstract
Abstract Illumination consistency plays an important role in realistic rendering of virtual characters which are integrated into a live video of real scene. This paper proposes a novel method for estimating the illumination conditions of outdoor videos captured by a fixed viewpoint. We first derive an analytical model which relates the statistics of an image to the lighting parameters of the scene adhering to the basic illumination model. Exploiting this model, we then develop a framework to estimate the lighting conditions of live videos. In order to apply the above approach to scenes containing dynamic objects such as intrusive pedestrians and swaying trees, we enforce two constraints, namely spatial and temporal illumination coherence, to refine the solution. Our approach requires no geometric information of the scenes and is sufficient for real‐time performance. Experiments show that with the lighting parameters recovered by our method, virtual characters can be seamlessly integrated into the live video. Copyright © 2010 John Wiley & Sons, Ltd.
Yanli Liu 0002, Xueying Qin, Guanyu Xing, Qunsheng Peng 0001
Comput. Animat. Virtual Worlds2
2009 Confidence-Based Color Modeling for Online Video Segmentation
Fan Zhong 0001, Xueying Qin, Jiazhou Chen 0002, Wei Hua 0002, Qunsheng Peng 0001
ACCV (2)2
2009 Locally Developable Constraint for Document Surface Reconstruction
abstract
This article presents a global optimization approach to reconstruct surfaces from a single document image. Instead of assuming globally developable in previous works which restricted the surface to be cylindrical, conical, etc, we use a free form parametric model which is simple yet expressive enough in the reconstruction task. We then apply developable constraints locally on sample points from extracted feature curves. And in order to further achieve better stability, the developable constraint we put on it is not so strong. Instead of isometric or conformal constraints used frequently in surface parametrization tasks, we use orthogonality. We show that even this is enough to reconstruct a wide class of document surfaces even with an uncalibrated camera.
Yuanlong Shao, Xinguo Liu, Xueying Qin, Hujun Bao
ICDAR3
2009 Light source estimation of outdoor scenes for mixed reality
Yanli Liu 0006, Xueying Qin, Songhua Xu, Eihachiro Nakamae, Qunsheng Peng 0001
Vis. Comput.2
2009 Video stabilization based on a 3D perspective camera model
Guofeng Zhang 0001, Wei Hua 0002, Xueying Qin, Yuanlong Shao, Hujun Bao
Vis. Comput.3
2008 A Two-Stage Audio Retrieval Method for Searching Unannotated Audio Clips
abstract
Traditional audio retrieval systems deal principally with audio clips having text descriptions. To retrieve unannotated audio clips is cumbersome because of the immaturity of content-based analysis and retrieval techniques. In this paper, we propose a two-stage audio retrieval method, consisting of a first stage of text-based retrieval and a second stage of content-based retrieval. This new retrieval method can be employed to retrieve audio clips from an audio collection having only partial text annotations, which is true of many online audio datasets. We have developed a prototype audio retrieval system based on our algorithm and carefully evaluated its performance. The results demonstrate the effectiveness of our new audio retrieval method. Our method can be generalized and applied to other kinds of non-textual data such as images and videos.
Songhua Xu, Suchao Chen, Kevin Y. Yip, Francis C. M. Lau 0001, Xueying Qin
ISM5
2007 Robust Metric Reconstruction from Challenging Video Sequences
abstract
Although camera self-calibration and metric reconstruction have been extensively studied during the past decades, automatic metric reconstruction from long video sequences with varying focal length is still very challenging. Several critical issues in practical implementations are not adequately addressed. For example, how to select the initial frames for initializing the projective reconstruction? What criteria should be used? How to handle the large zooming problem? How to choose an appropriate moment for upgrading the projective reconstruction to a metric one? This paper gives a careful investigation of all these issues. Practical and effective approaches are proposed. In particular, we show that existing image-based distance is not an adequate measurement for selecting the initial frames. We propose a novel measurement to take into account the zoom degree, the self-calibration quality, as well as image-based distance. We then introduce a new strategy to decide when to upgrade the projective reconstruction to a metric one. Finally, to alleviate the heavy computational cost in the bundle adjustment, a local on-demand approach is proposed. Our method is also extensively compared with the state-of-the-art commercial software to evidence its robustness and stability.
Guofeng Zhang 0001, Xueying Qin, Wei Hua 0002, Tien-Tsin Wong, Pheng-Ann Heng, Hujun Bao
CVPR2
2007 Stereoscopic Video Synthesis from a Monocular Video
abstract
This paper presents an automatic and robust approach to synthesize stereoscopic videos from ordinary monocular videos acquired by commodity video cameras. Instead of recovering the depth map, the proposed method synthesizes the binocular parallax in stereoscopic video directly from the motion parallax in monocular video. The synthesis is formulated as an optimization problem via introducing a cost function of the stereoscopic effects, the similarity, and the smoothness constraints. The optimization selects the most suitable frames in the input video for generating the stereoscopic video frames. With the optimized selection, convincing and smooth stereoscopic video can be synthesized even by simple constant-depth warping. No user interaction is required. We demonstrate the visually plausible results obtained given the input clips acquired by ordinary handheld video camera.
Guofeng Zhang 0001, Wei Hua 0002, Xueying Qin, Tien-Tsin Wong, Hujun Bao
IEEE Trans. Vis. Comput. Graph.3
2006 As-consistent-As-possible compositing of virtual objects and video sequences
abstract
Abstract We present an efficient approach that merges the virtual objects into video sequences taken by a freely moving camera in a realistic manner. The composition is visually and geometrically consistent through three main steps. First, a robust camera tracking algorithm based on key frames is proposed, which precisely recovers the focal length with a novel multi‐frame strategy. Next, the concerned 3D models of the real scenes are reconstructed by means of an extended multi‐baseline algorithm. Finally, the virtual objects in the form of 3D models are integrated into the real scenes, with special cares on the interaction consistency including shadow casting, occlusions, and object animation. A variety of experiments have been implemented, which demonstrate the robustness and efficiency of our approach. Copyright © 2006 John Wiley & Sons, Ltd.
Guofeng Zhang 0001, Xueying Qin, Xiaobo An, Wei Chen 0001, Hujun Bao
Comput. Animat. Virtual Worlds2
2005 Oriented Poisson matting
abstract
Matting is of great importance in image editing, which evaluates the opacity value with dependence on the provided foreground and background information. How to keep the subtle details is the main focus in soft matting. In this paper, we extend the previous Poisson matting approach in three points. An acquisition of subtle details is proposed, which is employed for the input to matting procedure. The modification to Poisson equation from divergence-based to eigenvector-based makes the matting more faithful to the details. Moreover, the construction of simultaneous equations is presented. By these improvements, Poisson matting could be finished in one procedure and extract the delicate matte. We name the proposed method as Oriented Poisson Matting. The demonstrations show that ours outperforms previous Poisson matting.
Zhenlong Du, Hai Lin 0003, Xueying Qin, Hujun Bao
ICIP (2)3
2004 Anti-Aliased Rendering of Water Surface
Xueying Qin, Eihachiro Nakamae, Wei Hua 0002, Yasuo Nagai, Qunsheng Peng 0001
J. Comput. Sci. Technol.1
2003 Fast Photo-Realistic Rendering of Trees in Daylight
abstract
Abstract We propose a fast approach for photo‐realistic rendering of trees under various kinds of daylight, which is particularlyuseful for the environmental assessment of landscapes. In our approach the 3D tree models are transformedto a quasi‐3D tree database registering geometrical and shading information of tree surfaces, i.e. their normalvectors, relative depth, and shadowing of direct sunlight and skylight, by using a combination of 2D buffers.Thus the rendering speed of quasi‐3D trees depends on their display sizes only, regardless of the complexity oftheir original 3D tree models. By utilizing a two‐step shadowing algorithm, our proposed method can create highquality forest scenes illuminated by both sunlight and skylight at a low cost. It can generate both umbrae andpenumbrae on a tree cast by other trees and any other objects such as buildings or clouds. Transparency, specularreflection and inter‐reflection of leaves, which influence the delicate shading effects of trees, can also be simulatedwith verisimilitude. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Three dimensional Graphics and Realism
Xueying Qin, Eihachiro Nakamae, Katsumi Tadamura, Yasuo Nagai
Comput. Graph. Forum1
2001 Rendering optimal solar shadows with plural sunlight depth buffers
Katsumi Tadamura, Xueying Qin, Guofang Jiao, Eihachiro Nakamae
Vis. Comput.2
1999 Rendering Optimal Solar Shadows using Plural Sunlight Depth Buffers
abstract
We propose a novel method, based on the two-pass z-buffer algorithm, to calculate shadows with sufficient precision and efficiency for rendering a daytime landscape with solar penumbrae. The feature of the proposed method is that the precision of the shadows can be preserved to be superior than that of any visible surface by using the appropriate number of plural shadow buffers; it gives a fairly satisfactory trade-off between computation time and quality of shadows. By using the shadow buffers, before calculating solar illuminance we can divide visible surfaces into three types of fragments, that is, bright regions, umbra regions, and the remaining penumbra regions. Thus, we can efficiently render a daytime landscape with sufficient precision of penumbrae because the algorithm can set the calculation points of solar illuminance variably, with their number depending on the type of region they belong to.
Katsumi Tadamura, Xueying Qin, Guofang Jiao, Eihachiro Nakamae
Computer Graphics International2
1999 Creating a Precise Panorama from Panned Video Sequence Images
abstract
Here we present an approach to offer a precise panoramic image composited from panned video sequence images, which are taken by a video camera on a tripod. The varying exposure levels due to the panned view fields taken by a camera with auto iris are adjusted. Any geometrical distortions in each frame, due to camera lens and/or CCD and perspective projection, are calibrated. The interlaced video sequence images commonly used for video systems are deinterlaced in order to create a high quality panoramic image.
Xueying Qin, Eihachiro Nakamae, Katsumi Tadamura
PG1
1999 Computer-generated still images composited with panned/zoomed landscape video sequences
Eihachiro Nakamae, Xueying Qin, Guofang Jiao, Przemyslaw Rokita, Katsumi Tadamura
Vis. Comput.2
1998 Computer Generated Still Images Composited with Panned Landscape Video Sequences
abstract
We present a novel approach to offer at a reasonable cost, panned landscape video sequence images well matched with photorealistic computer generated still images of large scale construction projects, such as bridges and electric power transmission towers. Our newly developed compositing algorithms can be used with the equipment which consists of a regular camcorder sold on the market, a video editing system, and our own software for rendering day-time scenes under various weather conditions. Other types of equipment such as an elaborate camera controller or special reference points in the landscapes, are unnecessary. The proposed algorithms provide the following features: (1) the computer generated still images are well harmonized with panned video sequence images at any time of day under various weather conditions, in terms of both optical and geometrical accuracy; (2) a precise panorama composited from video sequence images is created, in which varying exposure levels due to changing camera view fields can be coped with and any geometrical distortion is eliminated.
Eihachiro Nakamae, Xueying Qin, Guofang Jiao, Przemyslaw Rokita, Katsumi Tadamura, Yuji Usagawa
MMM2