EDBT 2026 Demo / reviewers in the wild / expert
Zhengyou Zhang
dblp:73/5043
· DBLP profile ↗
196ranked-venue papers
59as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 133 · 36 first-author · 3 since 2021Artificial intelligence and machine learning · 100 · 48 first-author · 5 since 2021Systems, architecture and hardware · 10 · 5 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMsabstractLarge language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess LLMs’ proficiency in following instructions on diverse real-world tasks. We construct a hierarchical task tree encompassing seven major areas covering over 200 categories and over 800 tasks, which covers diverse capabilities such as question answering, reasoning, multi-turn dialogue, and text generation, to evaluate LLMs in a comprehensive and in-depth manner. We also design detailed evaluation standards and processes to facilitate consistent, unbiased judgments from human evaluators. A test set of over 3,000 instances is released, spanning different difficulty levels and knowledge domains. Our work provides a standardized methodology to evaluate human alignment in LLMs for both English and Chinese. We also analyze the feasibility of automating parts of evaluation with a strong LLM (GPT-4). Our framework supports a thorough assessment of LLMs as they are integrated into real-world applications. We have made publicly available the task tree, TencentLLMEval dataset, and evaluation methodology which have been demonstrated as effective in assessing the performance of Tencent Hunyuan LLMs. By doing so, we aim to facilitate the benchmarking of advances in the development of safe and human-aligned LLMs. Shuyi Xie, Wenlin Yao, Yong Dai 0001, Zishan Xu, Fan Lin, Donglin Zhou, Lifeng Jin, Xinhua Feng, Pengzhi Wei, Zhichao Hu, Dong Yu 0001, Zhengyou Zhang |
ACM Trans. Intell. Syst. Technol. | 14 |
| 2024 | Max: A Wheeled-Legged Quadruped Robot for Multimodal Agile LocomotionabstractTo enrich legged robots with fast energy-efficient mobility on even terrain, wheeled-legged robots have emerged as a valued robot form in robotics research. This paper describes the complete development of a new wheeled-legged quadruped robot named Max, ranging from its mechanical design over system architecture to core algorithms implemented for it to realize various motion behaviors. Instead of attaching wheels to the distal ends of legs as in the existing wheeled-legged robot designs, this robot has wheels installed on the knees with a special switching mechanism to convert a leg between the legged and wheeled locomotion modes. This design keeps the wheeled leg lightweight, enabling the robot to preserve the motion agility as a quadruped robot while gaining the energy-efficiency as a four-wheel or even two-wheel mobile robot. An online locomotion generation method is proposed to compute the 6-D body trajectory of the robot in walking on the perceived terrain, while dynamic movements such as leaps and flips are generated by a unified trajectory optimizer, which is also used to generate the transition motions of the robot to transform into the wheeled mode. The diverse mobility of the proposed robot Max is verified with extensive experiments.Note to Practitioners—Empowering robots with all-terrain mobility is a fundamental open problem in developing a new generation of robots. To this end, combinations of wheels and legs have been explored for robots to possess both traversability on uneven terrains and efficiency on even terrains. This paper proposes a new wheeled-legged quadruped robot with focuses on the integrated design of wheeled legs, system architecture, and core algorithms implemented for various legged and wheeled locomotion behaviors. To embed wheels without adding additional motors and keep the light weight of original legs, a special switching mechanism is designed and integrated at the knee joints where wheels are installed. Algorithms for generating quadrupedal walk according to online perceived terrain information as well as other dynamic legged and wheeled motions are discussed and demonstrated. The system architecture for allocating all vision and motion algorithms is also presented. This work is intended to provide a whole picture of developing this new robot including both hardware and software aspects. Qinqin Zhou 0002, Xinyang Jiang, Wanchao Chi, Shenghao Zhang 0001, Jingfan Zhang, Rui Wang 0193, Jingchen Li 0001, Shuai Wang 0007, Lingzhu Xiang, Yu Zheng 0001, Zhengyou Zhang |
IEEE Trans Autom. Sci. Eng. | 17 |
| 2024 | Enabling Versatility and Dexterity of the Dual-Arm Manipulators: A General Framework Toward Universal Cooperative ManipulationabstractGrasping and manipulating various kinds of objects cooperatively is the core skill of a dual-arm robot when deployed as an autonomous agent in a human-centered environment. This requires fully exploiting the robot's versatility and dexterity. In this work, we propose a general framework for dual-arm manipulators that contains two correlative modules. The learning-based dexterity-reachability-aware perception module deals with vision-based bimanual grasping. It employs an end-to-end evaluation network and probabilistic modeling of the robot's reachability to deliver feasible and dexterity-optimum grasp pairs for unseen objects. The optimization-based versatility-oriented control module addresses the online cooperative manipulation control by using a hierarchical quadratic programming formulation. Self-collision avoidance and dual-arm manipulability ellipsoid tracking with high reliability and fidelity are simultaneously achieved based on a learned lightweight distance proxy function and a speed-level tracking technique on Riemannian manifold. Intrinsic system safety is guaranteed, and a novel interface for skill transfer is enabled. A long-horizon rearrangement experiment, a bimanual turnover manipulation, and multiple comparative performance evaluation verify the effectiveness of the proposed framework. Zhehua Zhou, Yang Yang 0031, Guangyao Zhai, Marion Leibold, Fenglei Ni, Zhengyou Zhang, Martin Buss, Yu Zheng 0001 |
IEEE Trans. Robotics | 8 |
| 2022 | Real-time Inertial Parameter Identification of Floating-Base Robots Through Iterative Primitive Shape DivisionabstractDynamic models play a key role in robot motion generation and control and the identification of inertial parameters is a critical component for obtaining an accurate dynamic model of a robot. This paper presents a novel iterative primitive shape division method for the inertia parameter identification of floating-base robots. Describing a robot by a set of primitive shapes with uniform mass distributions, the method iteratively divides the primitive shapes into smaller ones and refines their masses, which quickly converges to yielding the true inertia parameters of the robot. This method guarantees the physical consistency of the obtained parameters, possesses a high computational efficiency for online deployment, and works without contact force measurement. Furthermore, it can be used to estimate the position and magnitude of an external load applied to the robot. Simulations and experiments on a quadruped robot have been conducted to verify the effectiveness and efficiency of the proposed method. Jiafeng Xu, Yu Zheng 0001, Xinyang Jiang, Lingzhu Xiang, Zhengyou Zhang |
ICRA | 6 |
| 2022 | RECCraft System: Towards Reliable and Efficient Collective Robotic ConstructionabstractThis research presents a novel Collective Robotic Construction (CRC) system named RECCraft. The RECCraft hardware system is composed of the mobile manipulation vehicles, the cubic blocks, and the folding ramp blocks. Solid connection and easy removal of the blocks are achieved by an electropermanent magnet and silicon steel sheets. With one degree of freedom (DOF) lifting manipulator, the robot can carry a block 3.7 times its volume. An active folding ramp block can provide a robust passage to the upper level for the robot. Our study focuses on systemic improvement of the construction speed and reliability of the robotic construction system. Visual perception system realized by Apritag is adopted, featured by convenient deployment and high precision, to provide a reliable guarantee for robotic construction. RL-based planner provides end-to-end solution for planning tasks of building multi-layer constructions, which is validated by simulation platform and real prototype. Compared with construction speed of existing robotic construction systems, our proposed RECCraft system achieves state-of-the-art level. The robot builds a 2-layer construction by RL-based planner in 4 minutes and 16 seconds, which achieves construction volumetric throughput of 6.7×105mm3/s. Qiwei Xu, Yizheng Zhang, Shenghao Zhang 0001, Zhuoxing Wu, Xiong Li 0001, Jiahong Chen, Zengjun Zhao, Luyang Tang, Zhengyou Zhang, Lei Han 0001 |
IROS | 12 |
| 2022 | An Adaptive Approach to Whole-Body Balance Control of Wheel-Bipedal Robot OllieabstractThe wheel-bipedal robot has the advantages of both wheeled robots and legged robots, but as a cost, it is more challenging to perform flexible movements in various surroundings while keeping it balanced. The inaccurate dynamics of the robot makes the balance problem even more intractable. To solve this problem, the robot Ollie is used as a testbed. The whole-body control (WBC) framework is adopted to enhance the dexterity of the robot with multiple degrees of freedom in the task space. Moreover, a learning-based adaptive technique is applied to assist the WBC such that the balance controller can be designed in the absence of the accurate dynamics. Physical experiments demonstrate that the robot can manage various actions, with the help of the combination of the WBC and the learning-based adaptive technique. Jingfan Zhang, Shuai Wang 0007, Jie Lai, Zhenshan Bing, Yu Zheng 0001, Zhengyou Zhang |
IROS | 8 |
| 2022 | Explainable Hierarchical Imitation Learning for Robotic Drink PouringabstractTo accurately pour drinks into various containers is an essential skill for service robots. However, drink pouring is a dynamic process and difficult to model. Traditional deep imitation learning techniques for implementing autonomous robotic pouring have an inherent black-box effect and require a large amount of demonstration data for model training. To address these issues, an Explainable Hierarchical Imitation Learning (EHIL) method is proposed in this paper such that a robot can learn high-level general knowledge and execute low-level actions across multiple drink pouring scenarios. Moreover, with the EHIL method, a logical graph can be constructed for task execution, through which the decision-making process for action generation can be made explainable to users and the causes of failure can be traced out. Based on the logical graph, the framework is manipulable to achieve different targets while the adaptability to unseen scenarios can be achieved in an explainable manner. A series of experiments have been conducted to verify the effectiveness of the proposed method. Results indicate that EHIL outperforms the traditional behavior cloning method in terms of success rate, adaptability, manipulability, and explainability. Note to Practitioners—Pouring liquids is a common activity in people’s daily lives and all wet-lab industries. Drink pouring dynamic control is difficult to model, while the accurate perception of flow is challenging. To enable the robot to learn under unknown dynamics via observing the human demonstration, deep imitation learning can be used. To address the limitations of traditional deep neural networks, an Explainable Hierarchical Imitation Learning (EHIL) method is proposed in this paper. The proposed method enables the robot to learn a sequence of reasonable pouring phases for performing the task rather than simply execute the task via traditional behavior cloning. In this way, explainability and safety can be ensured. Manipulability can be achieved by reconstructing the logical graph. The target of this research is to obtain pouring dynamics via the learning method and realize the precise and quick pouring of drink from the source containers to various targeted containers with reliable performance, adaptability, manipulability, and explainability. Dandan Zhang 0001, Qiang Li 0001, Yu Zheng 0001, Lei Wei 0002, Zhengyou Zhang |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2022 | High-Fidelity 3D Digital Human Head Creation from RGB-D SelfiesabstractWe present a fully automatic system that can produce high-fidelity, photo-realistic three-dimensional (3D) digital human heads with a consumer RGB-D selfie camera. The system only needs the user to take a short selfie RGB-D video while rotating his/her head and can produce a high-quality head reconstruction in less than 30 s. Our main contribution is a new facial geometry modeling and reflectance synthesis procedure that significantly improves the state of the art. Specifically, given the input video a two-stage frame selection procedure is first employed to select a few high-quality frames for reconstruction. Then a differentiable renderer-based 3D Morphable Model (3DMM) fitting algorithm is applied to recover facial geometries from multiview RGB-D data, which takes advantages of a powerful 3DMM basis constructed with extensive data generation and perturbation. Our 3DMM has much larger expressive capacities than conventional 3DMM, allowing us to recover more accurate facial geometry using merely linear basis. For reflectance synthesis, we present a hybrid approach that combines parametric fitting andConvolutional Neural Networks (CNNs)to synthesize high-resolution albedo/normal maps with realistic hair/pore/wrinkle details. Results show that our system can produce faithful 3D digital human faces with extremely realistic details. The main code and the newly constructed 3DMM basis is publicly available. Linchao Bao, Xiangkai Lin, Haoxian Zhang, Xuefei Zhe, Hao-Zhi Huang 0001, Xinwei Jiang, Jue Wang 0001, Dong Yu 0001, Zhengyou Zhang |
ACM Trans. Graph. | 12 |
| 2021 | Balance Control of a Novel Wheel-legged Robot: Design and ExperimentsabstractThis paper presents a balance control technique for a novel wheel-legged robot. We first derive a dynamic model of the robot and then apply a linear feedback controller based on output regulation and linear quadratic regulator (LQR) methods to maintain the standing of the robot on the ground without moving backward and forward mightily. To take into account nonlinearities of the model and obtain a large domain of stability, a nonlinear controller based on the interconnection and damping assignment - passivity-based control (IDA-PBC) method is exploited to control the robot in more general scenarios. Physical experiments are performed with various control tasks. Experimental results demonstrate that the proposed linear output regulator can maintain the standing of the robot, while the proposed nonlinear controller can balance the robot under an initial starting angle far away from the equilibrium point, or under a changing robot height. Shuai Wang 0007, Leilei Cui 0002, Jingfan Zhang, Jie Lai, Yu Zheng 0001, Zhengyou Zhang, Zhong-Ping Jiang |
ICRA | 8 |
| 2021 | Run Like a Dog: Learning Based Whole-Body Control Framework for Quadruped Gait Style TransferabstractIn this paper, a learning-based whole-body loco-motion controller is proposed, which enables quadruped robots to perform running in the style of real animals. We use a low-level controller based on multi-rigid body dynamics to calculate desired torques for each joint, while the high-level neural network policy planning the expected gait and foothold. The policy is trained with reinforcement learning, so that the robot can track a variety of trajectories according to the gait patterns recorded from real-world dogs. We transfer the walking and running gait style to quadrupeds in simulation, involving pace, trot, high-speed gallop and natural transitions. The performance is evaluated by the synchronization rate of contact state between the policy result and the recorded sequence. In the experiments, the robot runs steadily at a speed of 2 m/s and showcases a notable synchronization rate of about 80%. Without prior knowledge, the policy demonstrates a realistic foothold distribution that covers the central area of the torso, which is prevalent in running animals. Fulong Yin, Annan Tang, Liangwei Xu, Yu Zheng 0001, Zhengyou Zhang, Xiangyu Chen 0001 |
IROS | 6 |
| 2021 | Digital Human in an Integrated Physical-Digital World (IPhD)abstractWith the rapid development of digital technologies such as VR, AR, XR, and more importantly the almost ubiquitous mobile broadband coverage, we are entering an Integrated Physical-Digital World (IPhD), the tight integration of virtual world with the physical world. The IPhD is characterized with four key technologies: Virtualization of the physical world, Realization of the virtual world, Holographic internet, and Intelligent Agent. Internet will continue its development with faster speed and broader bandwidth, and will eventually be able to communicate holographic contents including 3D shape, appearance, spatial audio, touch sensing and smell. Intelligent agents, such as digital human, and digital/physical robots, travels between digital and physical worlds. In this talk, we will describe our work on digital human for this IPhD world. This includes: computer vision techniques for building digital humans, multimodal text-to-speech synthesis (voice and lip shapes), speech-driven face animation, neural-network-based body motion control, human-digital-human interaction, and an emotional video game anchor. Zhengyou Zhang |
ACM Multimedia | 1 |
| 2021 | Joint Hand-Object 3D Reconstruction From a Single Image With Cross-Branch Feature FusionabstractAccurate 3D reconstruction of the hand and object shape from a hand-object image is important for understanding human-object interaction as well as human daily activities. Different from bare hand pose estimation, hand-object interaction poses a strong constraint on both the hand and its manipulated object, which suggests that hand configuration may be crucial contextual information for the object, and vice versa. However, current approaches address this task by training a two-branch network to reconstruct the hand and object separately with little communication between the two branches. In this work, we propose to consider hand and object jointly in feature space and explore the reciprocity of the two branches. We extensively investigate cross-branch feature fusion architectures with MLP or LSTM units. Among the investigated architectures, a variant with LSTM units that enhances object feature with hand feature shows the best performance gain. Moreover, we employ an auxiliary depth estimation module to augment the input RGB image with the estimated depth map, which further improves the reconstruction accuracy. Experiments conducted on public datasets demonstrate that our approach significantly outperforms existing approaches in terms of the reconstruction accuracy of objects. Yujin Chen, Zhigang Tu 0001, Ruizhi Chen, Linchao Bao, Zhengyou Zhang, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Nonlinear Balance Control of an Unmanned Bicycle: Design and ExperimentsabstractIn this paper, nonlinear control techniques are exploited to balance an unmanned bicycle with enlarged stability domain. We consider two cases. For the first case when the autonomous bicycle is balanced by the flywheel, the steering angle is set to zero, and the torque of the flywheel is used as the control input. The controller is designed based on the Interconnection and Damping Assignment Passivity Based Control (IDA-PBC) method. For the second case when the bicycle is balanced by the handlebar, the bicycle's velocity is high, and the flywheel is turned off. The angular velocity of the handlebar is used as the control input and the balance controller is designed based on feedback linearization. In these cases, the global stability of the closed-loop unmanned bicycle is theoretically proved based on Lyapunov theory. The experiments are conducted to validate the efficacy of the proposed nonlinear balance controllers. Leilei Cui 0002, Shuai Wang 0007, Jie Lai, Xiangyu Chen 0001, Zhengyou Zhang, Zhong-Ping Jiang |
IROS | 6 |
| 2020 | Gain Scheduled Controller Design for Balancing an Autonomous BicycleabstractIn this paper, the gain scheduling technique is applied to design a balance controller for an autonomous bicycle with an inertia wheel. Previously, two different balance controllers are needed depending on whether the bicycle is stationary or dynamic. The switch between the two different controllers may cause the instability of the autonomous bicycle. Our proposed gain scheduled controller can balance the autonomous bicycle in both stationary and dynamic cases. A physical system is built and experiments are carried out to demonstrate the effectiveness of the gain scheduled controller. Shuai Wang 0007, Leilei Cui 0002, Jie Lai, Xiangyu Chen 0001, Yu Zheng 0001, Zhengyou Zhang, Zhong-Ping Jiang |
IROS | 7 |
| 2020 | A Flexible Dual-Core Optical Waveguide Sensor for Simultaneous and Continuous Measurement of Contact Force and PositionabstractHaving the merits of chemical inertness and immunity to electromagnetic interference, light weight, small size, and softness, optical waveguides have attracted much attention in making tactile sensors recently. This paper presents a new design of waveguide using two layers of cores, one of which has an uniform width and the other has an incremental width. It is deduced and verified that the contact force can be derived from the light power loss in the uniform-width core, while the contact position can be derived from the light power loss in the other core together with the estimated force. By this dual-core design, a single waveguide can simultaneously and continuously measure the contact force and position along it, which makes it very suited for integration on some thin long robotic parts, such as robotic fingers. A hardware experiment has been conducted to demonstrate its effectiveness on a two-finger gripper in an assembly task. The dual-core waveguide achieves 2 mm spatial resolution and 0.1 N sensitivity. Zhong Zhang 0015, Yu Zheng 0001, Jia Pan 0001, Xiong Li 0001, Zhengyou Zhang |
IROS | 6 |
| 2020 | Jointly Learning Visual Poses and Pose Lexicon for Semantic Action RecognitionabstractA novel method for semantic action recognition through learning a pose lexicon is presented in this paper. A pose lexicon comprises a set of semantic poses, a set of visual poses, and a probabilistic mapping between the visual and semantic poses. This paper assumes that both the visual poses and mapping are hidden and proposes a method to simultaneously learn a visual pose model that estimates the likelihood of an observed video frame being generated from hidden visual poses, and a pose lexicon model establishes the probabilistic mapping between the hidden visual poses and the semantic poses parsed from textual instructions. Specifically, the proposed method consists of two-level hidden Markov models. One level represents the alignment between the visual poses and semantic poses. The other level represents a visual pose sequence, and each visual pose is modeled as a Gaussian mixture. An expectation-maximization algorithm is developed to train a pose lexicon. With the learned lexicon, action classification is formulated as a problem of finding the maximum posterior probability of a given sequence of video frames that follows a given sequence of semantic poses, constrained by the most likely visual pose and the alignment sequences. The proposed method was evaluated on MSRC-12, WorkoutSU-10, WorkoutUOW-18, Combined-15, Combined-17, and Combined-50 action datasets using cross-subject, cross-dataset, zero-shot, and seen/unseen protocols. Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Challenges of Multimodal Interaction in the Era of Human-Robot CoexistenceabstractWith the rapid progress in computing and sensory technologies, we will enter the era of human-robot coexistence in the not-too-distant future, and it is time to address the challenges of multimodal interaction. Should a robot take the form of humanoid? Is it better for robots to behave as a second-class citizen or as an equal part of the society as human? Should the communication between human and robot be symmetric or is it okay to be asymmetric? And how about the communication between robots with human presence? What does it mean by emotional intelligence for robots? With the inevitable physical interaction between human and robot, how to guarantee safety? What is the ethical and moral model for robots and how do they follow? Zhengyou Zhang |
ICMI | 1 |
| 2019 | Curriculum-guided Hindsight Experience ReplayabstractIn off-policy deep reinforcement learning, it is usually hard to collect sufficient successful experiences with sparse rewards to learn from. Hindsight experience replay (HER) enables an agent to learn from failures by treating the achieved state of a failed experience as a pseudo goal. However, not all the failed experiences are equally useful to different learning stages, so it is not efficient to replay all of them or uniform samples of them. In this paper, we propose to 1) adaptively select the failed experiences for replay according to the proximity to the true goals and the curiosity of exploration over diverse pseudo goals, and 2) gradually change the proportion of the goal-proximity and the diversity-based curiosity in the selection criteria: we adopt a human-like learning strategy that enforces more curiosity in earlier stages and changes to larger goal-proximity later. This Goal-and-Curiosity-driven Curriculum Learning'' leads toCurriculum-guided HER (CHER)'', which adaptively and dynamically controls the exploration-exploitation trade-off during the learning process via hindsight experience selection. We show that CHER improves the state of the art in challenging robotics environments. Tianyi Zhou 0001, Yali Du 0001, Lei Han 0001, Zhengyou Zhang |
NeurIPS | 5 |
| 2018 | End-to-End Convolutional Semantic EmbeddingsabstractSemantic embeddings for images and sentences have been widely studied recently. The ability of deep neural networks on learning rich and robust visual and textual representations offers the opportunity to develop effective semantic embedding models. Currently, the state-of-the-art approaches in semantic learning first employ deep neural networks to encode images and sentences into a common semantic space. Then, the learning objective is to ensure a larger similarity between matching image and sentence pairs than randomly sampled pairs. Usually, Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) are employed for learning image and sentence representations, respectively. On one hand, CNNs are known to produce robust visual features at different levels and RNNs are known for capturing dependencies in sequential data. Therefore, this simple framework can be sufficiently effective in learning visual and textual semantics. On the other hand, different from CNNs, RNNs cannot produce middle-level (e.g. phrase-level in text) representations. As a result, only global representations are available for semantic learning. This could potentially limit the performance of the model due to the hierarchical structures in images and sentences. In this work, we apply Convolutional Neural Networks to process both images and sentences. Consequently, we can employ mid-level representations to assist global semantic learning by introducing a new learning objective on the convolutional layers. The experimental results show that our proposed textual CNN models with the new learning objective lead to better performance than the state-of-the-art approaches. Quanzeng You, Zhengyou Zhang, Jiebo Luo 0001 |
CVPR | 2 |
| 2018 | Depth Super-Resolution on RGB-D Video Sequences With Large Displacement 3D MotionabstractTo enhance the resolution and accuracy of depth data, some video-based depth super-resolution methods have been proposed which utilizes its neighboring depth images in the temporal domain. They often consist of two main stages: motion compensation of temporally neighboring depth images and fusion of compensated depth images. However, large displacement 3D motion often leads to compensation error, and the compensation error is further introduced into the fusion. A video-based depth super-resolution method with novel motion compensation and fusion approaches is proposed in this paper. We claim that, 3D Nearest Neighboring Field (NNF) is a better choice than using positions with true motion displacement for depth enhancements. To handle large displacement 3D motion, the compensation stage utilized 3D NNF instead of true motion used in previous methods. Next, the fusion approach is modeled as a regression problem to predict the super-resolution result efficiently for each depth image by using its compensated depth images. A new deep convolutional neural network architecture is designed for fusion, which is able to employ a large amount of video data for learning the complicated regression function. We comprehensively evaluate our method on various RGB-D video sequences to show its superior performance. Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Zhengyou Zhang, Yunde Jia |
IEEE Trans. Image Process. | 5 |
| 2017 | Adversarial Ranking for Language GenerationabstractGenerative adversarial networks (GANs) have great successes on synthesizing data. However, the existing GANs restrict the discriminator to be a binary classifier, and thus limit their learning capacity for tasks that need to synthesize output with rich structures such as natural language descriptions. In this paper, we propose a novel generative adversarial network, RankGAN, for generating high-quality language descriptions. Rather than training the discriminator to learn and assign absolute binary predicate for individual data sample, the proposed RankGAN is able to analyze and rank a collection of human-written and machine-written sentences by giving a reference group. By viewing a set of data samples collectively and evaluating their quality through relative ranking scores, the discriminator is able to make better assessment which in turn helps to learn a better generator. The proposed RankGAN is optimized through the policy gradient technique. Experimental results on multiple public datasets clearly demonstrate the effectiveness of the proposed approach. Dianqi Li, Xiaodong He 0001, Ming-Ting Sun, Zhengyou Zhang |
NIPS | 5 |
| 2017 | Semantic action recognition by learning a pose lexicon
Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona, Zhengyou Zhang |
Pattern Recognit. | 4 |
| 2017 | Guest Editorial Introduction to the Special Issue on Group and Crowd Behavior Analysis for Intelligent Multicamera Video SurveillanceabstractDespite significant progress in human behavior analysis over the past few years, most of today’s state-of-the-art algorithms focus on analyzing individual behavior in a simple environment monitored by a single camera. Recently, the widespread availability of cameras and a growing need for public safety have shifted the attention of researchers in video surveillance from individual behavior analysis to group and crowd behavior analysis in multicamera networks. Group behavior analysis provides a novel level for describing events, which are semantically more meaningful, highlighting barely visible relational connections among people. Crowd behavior analysis can also be used for anomaly detection such as panic scenarios, dangerous situations, and illegal behaviors in public spaces. Hongxun Yao, Andrea Cavallaro, Thierry Bouwmans, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Training deep networks for facial expression recognition with crowd-sourced label distributionabstractCrowd sourcing has become a widely adopted scheme to collect ground truth labels. However, it is a well-known problem that these labels can be very noisy. In this paper, we demonstrate how to learn a deep convolutional neural network (DCNN) from noisy labels, using facial expression recognition as an example. More specifically, we have 10 taggers to label each input image, and compare four different approaches to utilizing the multiple labels: majority voting, multi-label learning, probabilistic label drawing, and cross-entropy loss. We show that the traditional majority voting scheme does not perform as well as the last two approaches that fully leverage the label distribution. An enhanced FER+ data set with multiple labels for each face image will also be shared with the research community. Emad Barsoum, Cha Zhang, Cristian Canton, Zhengyou Zhang |
ICMI | 4 |
| 2016 | Guest Editorial: Human Activity Understanding from 2D and 3D Data
Junsong Yuan 0001, Wanqing Li 0001, Zhengyou Zhang, David J. Fleet, Jamie Shotton |
Int. J. Comput. Vis. | 3 |
| 2016 | Camera calibration: a personal retrospective
Zhengyou Zhang |
Mach. Vis. Appl. | 1 |
| 2016 | Handling Occlusion and Large Displacement Through Improved RGB-D Scene Flow EstimationabstractThe accuracy of scene flow is restricted by several challenges such as occlusion and large displacement motion. When occlusion happens, the positions inside the occluded regions lose their corresponding counterparts in preceding and succeeding frames. Large displacement motion will increase the complexity of motion modeling and computation. Moreover, occlusion and large displacement motion are highly related problems in scene flow estimation, e.g., large displacement motion often leads to considerably occluded regions in the scene. An improved dense scene flow method based on red-green-blue-depth (RGB-D) data is proposed in this paper. To handle occlusion, we model the occlusion status for each point in our problem formulation, and jointly estimate the scene flow and occluded regions. To deal with large displacement motion, we employ an over-parameterized scene flow representation to model both the rotation and translation components of the scene flow, since large displacement motion cannot be well approximated using translational motion only. Furthermore, we employ a two-stage optimization procedure for this overparameterized scene flow representation. In the first stage, we propose a new RGB-D PatchMatch method, which is mainly applied in the RGB-D image space to reduce the computational complexity introduced by the large displacement motion. According to the quantitative evaluation based on the Middlebury data set, our method outperforms other published methods. The improved performance is also comprehensively confirmed on the real data acquired by Kinect sensor. Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Philip A. Chou, Zhengyou Zhang, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | Dimension Reduction With Extreme Learning MachineabstractData may often contain noise or irrelevant information, which negatively affect the generalization capability of machine learning algorithms. The objective of dimension reduction algorithms, such as principal component analysis (PCA), non-negative matrix factorization (NMF), random projection (RP), and auto-encoder (AE), is to reduce the noise or irrelevant information of the data. The features of PCA (eigenvectors) and linear AE are not able to represent data as parts (e.g. nose in a face image). On the other hand, NMF and non-linear AE are maimed by slow learning speed and RP only represents a subspace of original data. This paper introduces a dimension reduction framework which to some extend represents data as parts, has fast learning speed, and learns the between-class scatter subspace. To this end, this paper investigates a linear and non-linear dimension reduction framework referred to as extreme learning machine AE (ELM-AE) and sparse ELM-AE (SELM-AE). In contrast to tied weight AE, the hidden neurons in ELM-AE and SELM-AE need not be tuned, and their parameters (e.g, input weights in additive neurons) are initialized using orthogonal and sparse random weights, respectively. Experimental results on USPS handwritten digit recognition data set, CIFAR-10 object recognition, and NORB object recognition data set show the efficacy of linear and non-linear ELM-AE and SELM-AE in terms of discriminative capability, sparsity, training time, and normalized mean square error. Liyanaarachchi Lekamalage Chamara Kasun, Guang-Bin Huang, Zhengyou Zhang |
IEEE Trans. Image Process. | 4 |
| 2015 | Deeply-Supervised NetsabstractWe propose deeply-supervised nets (DSN), a method that simultaneously minimizes classification error and improves the directness and transparency of the hidden layer learning process. We focus our attention on three aspects of traditional convolutional-neural-network-type (CNN-type) architectures: (1) transparency in the effect intermediate layers have on overall classification; (2) discriminativeness and robustness of learned features, especially in early layers; (3) training effectiveness in the face of “vanishing” gradients. To combat these issues, we introduce “companion” objective functions at each hidden layer, in addition to the overall objective function at the output layer (an integrated strategy distinct from layer-wise pre-training). We also analyze our algorithm using techniques extended from stochastic gradient methods. The advantages provided by our method are evident in our experimental results, showing state-of-the-art performance on MNIST, CIFAR-10, CIFAR-100, and SVHN. Chen-Yu Lee, Saining Xie, Patrick W. Gallagher, Zhengyou Zhang, Zhuowen Tu |
AISTATS | 4 |
| 2015 | ImmerseBoard: Immersive Telepresence Experience using a Digital WhiteboardabstractImmerseBoard is a system for remote collaboration through a digital whiteboard that gives participants a 3D immersive experience, enabled only by an RGBD camera (Microsoft Kinect) mounted on the side of a large touch display. Using 3D processing of the depth images, life-sized rendering, and novel visualizations, ImmerseBoard emulates writing side-by-side on a physical whiteboard, or alternatively on a mirror. User studies involving three tasks show that compared to standard video conferencing with a digital whiteboard, ImmerseBoard provides participants with a quantitatively better ability to estimate their remote partners' eye gaze direction, gesture direction, intention, and level of agreement. Moreover, these quantitative capabilities translate qualitatively into a heightened sense of being together and a more enjoyable experience. ImmerseBoard's form factor is suitable for practical and easy installation in homes and offices. Keita Higuchi, Yinpeng Chen, Philip A. Chou, Zhengyou Zhang, Zicheng Liu 0001 |
CHI | 4 |
| 2015 | Maximum a posteriori estimation of room impulse responsesabstractEstimating room impulse responses (RIRs) has a number of applications, including personalized audio, analyzing and improving acoustic behavior of concert halls, listening room compensation, sound source localization, and many others. RIRs have been estimated in essentially the same fashion for the last 50 years: Compute the cross correlation between a signal played at point A, and the signal received at point B. Best results are obtained when the signal played is white noise, or a maximum length sequence. No prior knowledge is exploited in computing the RIR, which is simply assumed to be the cross correlation between played and received signals. In contrast, research in adaptive RIR estimation (a.k.a. adaptive Acoustic Echo Cancellation) has made huge progress by (among other things) incorporating models for the RIR. In this paper we propose a new RIR estimation technique, based on a maximum a posteriori formulation. More specifically, we estimate the room reverberation time, as well as the room noise level, and use those as priors for the RIR estimation. Comparison with ground truth shows an average improvement of 12 dB compared to traditional methods. Dinei A. F. Florêncio, Zhengyou Zhang |
ICASSP | 2 |
| 2015 | VTouch: Vision-enhanced interaction for large touch displaysabstractWe propose a system that augments touch input with visual understanding of the user to improve interaction with a large touch-sensitive display. A commodity color plus depth sensor such as Microsoft Kinect adds the visual modality and enables new interactions beyond touch. Through visual analysis, the system understands where the user is, who the user is, and what the user is doing even before the user touches the display. Such information is used to enhance interaction in multiple ways. For example, a user can use simple gestures to bring up menu items such as color palette and soft keyboard; menu items can be shown where the user is and can follow the user; hovering can show information to the user before the user commits to touch; the user can perform different functions (for example writing and erasing) with different hands; and the user's preference profile can be maintained, distinct from other users. User studies are conducted and the users very much appreciate the value of these and other enhanced interactions. Yinpeng Chen, Zicheng Liu 0001, Philip A. Chou, Zhengyou Zhang |
ICME | 4 |
| 2015 | Vision-enhanced Immersive Interaction and Remote Collaboration with Large Touch DisplaysabstractLarge displays are becoming commodity, and more and more, they are touch-enabled. In this keynote, we describe a system called ViiBoard (Vision-enhanced Immersive Interaction with touch Board) that enables natural interaction and immersive remote collaboration with large touch displays by adding a commodity color plus depth sensor. It consists of two parts. The first part is called VTouch that augments touch input with visual understanding of the user to improve interaction with a large touch-sensitive display such as Microsoft Surface Hub. An RGBD sensor such as Microsoft Kinect adds the visual modality and enables new interactions beyond touch. Through visual analysis, the system understands where the user is, who the user is, and what the user is doing even before the user touches the display. Such information is used to enhance interaction in multiple ways. For example, a user can use simple gestures to bring up menu items such as color palette and soft keyboard; menu items can be shown where the user is and can follow the user; hovering can show information to the user before the user commits to touch; the user can perform different functions (for example writing and erasing) with different hands; and the user's preference profile can be maintained, distinct from other users. User studies are conducted and the users very much appreciate the value of these and other enhanced interactions. Zhengyou Zhang |
ACM Multimedia | 1 |
| 2015 | A survey on face detection in the wild: Past, present and future
Stefanos Zafeiriou, Cha Zhang, Zhengyou Zhang |
Comput. Vis. Image Underst. | 3 |
| 2015 | Visual Understanding with RGB-D Sensors: An Introduction to the Special Issueabstract10.1145/2732265 Richang Hong, Shuicheng Yan, Zhengyou Zhang |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2014 | Can Visual Recognition Benefit from Auxiliary Information in Training?
Qilin Zhang 0004, Gang Hua 0001, Wei Liu 0005, Zicheng Liu 0001, Zhengyou Zhang |
ACCV (1) | 5 |
| 2014 | Towards accurate and robust cross-ratio based gaze trackers through learning from simulationabstractCross-ratio (CR) based methods offer many attractive properties for remote gaze estimation using a single camera in an uncalibrated setup by exploiting invariance of a plane projectivity. Unfortunately, due to several simplification assumptions, the performance of CR-based eye gaze trackers decays significantly as the subject moves away from the calibration position. In this paper, we introduce an adaptive homography mapping for achieving gaze prediction with higher accuracy at the calibration position and more robustness under head movements. This is achieved with a learning-based method for compensating both spatially-varying gaze errors and head pose dependent errors simultaneously in a unified framework. The model of adaptive homography is trained offline using simulated data, saving a tremendous amount of time in data collection. We validate the effectiveness of the proposed approach using both simulated and real data from a physical setup. We show that our method compares favorably against other state-of-the-art CR based methods. Jia-Bin Huang 0001, Qin Cai, Zicheng Liu 0001, Narendra Ahuja, Zhengyou Zhang |
ETRA | 5 |
| 2014 | Improving cross-ratio-based eye tracking techniques by leveraging the binocular fixation constraintabstractThe cross-ratio approach has recently attracted increasing attention in eye-gaze tracking due to its simplicity in setting up a tracking system. Its accuracy, however, is lower than that of the model-based approach, and substantial efforts have been devoted to improving its accuracy. Binocular fixation is essential for humans to have good depth perception, and this paper presents a technique leveraging this constraint. It is used in two ways: First, in estimating jointly the homography matrices for both eyes, and second, in estimating the eye gaze itself. Experimental results with both synthetic and real data show that the proposed approach produces significantly better results than using a single eye and also better than averaging the independent results from the two eyes. Zhengyou Zhang, Qin Cai |
ETRA | 1 |
| 2014 | Facial expression tracking from head-mounted, partially observing camerasabstractHead-mounted displays (HMDs) have gained more and more interest recently. They can enable people to communicate with each other from anywhere, at anytime. However, since most HMDs today are only equipped with cameras pointing outwards, the remote party would not be able to see the user wearing the HMD. In this paper, we present a system for facial expression tracking based on head-mounted, inward looking cameras, such that the user can be represented with animated avatars at the remote party. The main challenge is that the cameras can only observe partial faces since they are very close to the face. We experiment with multiple machine learning algorithms to estimate facial expression parameters based on training data collected with the assistance of a Kinect depth sensor. Our results show that we can reliably track people's facial expression even from very limited view angles of the cameras. Bernardino Romera-Paredes, Cha Zhang, Zhengyou Zhang |
ICME | 3 |
| 2014 | Deblur a blurred RGB image with a sharp NIR image through local linear mappingabstractImage acquisition in a low light environment requires long exposure to achieve acceptable signal-to-noise ratio, which however causes blurry effect. This paper addresses this problem by using a sharp near-infrared (NIR) image when the environment has sufficient NIR light. We assume that an RGB and NIR image pair has a linear mapping in a local area and that the mapping function is valid for both the blur and sharp image pairs. Using this property, we solve the sharp RGB images from a blurred RGB image and the corres ponding s harp NIR image. The effectiveness of the proposed algorithm is verified with both synthetic and real captured datasets. Tao Yue 0003, Ming-Ting Sun, Zhengyou Zhang, Jin-Li Suo, Qionghai Dai |
ICME | 3 |
| 2014 | WWN: Integration with coarse-to-fine, supervised and reinforcement learningabstractThe cost of autonomous development is substantial. Although supervised learning is effective, the cost demand on teachers is often too high to be constantly applied. Reinforcement learning can take advantage of physical reality due to environmental feedback and inspections. Information required in reinforcement learning is not as specific as is required in supervised learning. Integration theories, methods, and analysis of these two learning strategies are still rare in the literature although such integration has been well known in the animal kingdom. Based on our prior work on a general purpose framework called Developmental Network and its embodiment Where-What-Network, we present our theory, method, and analysis for integration of supervised learning and reinforcement learning in this paper. Different from all other known work on reinforcement learning, this DN framework uses fully emergent representation to avoid the brittleness and task-specific representations. Central in the integration is not just to provide a freedom for the teacher to choose the mode of learning, which is necessary especially when the physical non-living world is an implicit teacher, but the mechanism of scaffolding. In our experiment the scaffolding is reflected by allowing the location motor(LM) neurons to gradually refine representation through splitting(mitosis) in a coarse to fine scheme. We report our experimental work in a very challenging learning setting: both object and backgrounds are unknown(cluttered settings) and concepts(e.g. location and type) emerge from agent-environment interactions, instead of rigidly handcrafted. Zejia Zheng, Juyang Weng, Zhengyou Zhang |
IJCNN | 3 |
| 2014 | Improving multiview face detection with multi-task deep convolutional neural networksabstractMultiview face detection is a challenging problem due to dramatic appearance changes under various pose, illumination and expression conditions. In this paper, we present a multi-task deep learning scheme to enhance the detection performance. More specifically, we build a deep convolutional neural network that can simultaneously learn the face/nonface decision, the face pose estimation problem, and the facial landmark localization problem. We show that such a multi-task learning scheme can further improve the classifier's accuracy. On the challenging FDDB data set, our detector achieves over 3% improvement in detection rate at the same false positive rate compared with other state-of-the-art methods. Cha Zhang, Zhengyou Zhang |
WACV | 2 |
| 2013 | Tensor-Based Human Body ModelingabstractIn this paper, we present a novel approach to model 3D human body with variations on both human shape and pose, by exploring a tensor decomposition technique. 3D human body modeling is important for 3D reconstruction and animation of realistic human body, which can be widely used in Tele-presence and video game applications. It is challenging due to a wide range of shape variations over different people and poses. The existing SCAPE model is popular in computer vision for modeling 3D human body. However, it considers shape and pose deformations separately, which is not accurate since pose deformation is person-dependent. Our tensor-based model addresses this issue by jointly modeling shape and pose deformations. Experimental results demonstrate that our tensor-based model outperforms the SCAPE model quite significantly. We also apply our model to capture human body using Microsoft Kinect sensors with excellent results. Yinpeng Chen, Zicheng Liu 0001, Zhengyou Zhang |
CVPR | 3 |
| 2013 | Wide-Baseline Hair Capture Using Strand-Based RefinementabstractWe propose a novel algorithm to reconstruct the 3D geometry of human hairs in wide-baseline setups using strand-based refinement. The hair strands are first extracted in each 2D view, and projected onto the 3D visual hull for initialization. The 3D positions of these strands are then refined by optimizing an objective function that takes into account cross-view hair orientation consistency, the visual hull constraint and smoothness constraints defined at the strand, wisp and global levels. Based on the refined strands, the algorithm can reconstruct an approximate hair surface: experiments with synthetic hair models achieve an accuracy of ~3mm. We also show real-world examples to demonstrate the capability to capture full-head hair styles as well as hair in motion with as few as 8 cameras. Linjie Luo, Cha Zhang, Zhengyou Zhang, Szymon Rusinkiewicz |
CVPR | 3 |
| 2013 | Modeling the effects of neuromodulation on internal brain areas: Serotonin and dopamineabstractThe effects of neuromodulator, such as serotonin and dopamine, on individual neurons in the brain have been known qualitatively. However, it is challenging to computationally model such effects in an emergent network, as the elements of internal representations do not have a static, task-specific meaning. Weng and coworkers modeled the effects of serotonin and dopamine on only motor neurons in emergent networks. In this work, we extend the effects of serotonin and dopamine to all neurons inside the emergent network. Our new theory is that although serotonin and dopamine indicate events of different natures (aversive and appetitive), they produce similar effects on internal non-motor neurons in that they increase their learning rates from the cases without serotonin and dopamine. This is because the presence of serotonin and dopamine indicates a higher importance of the event compared with baseline cases. Experimentally, we show that the enhanced developmental network learns faster under a limited resource. Zejia Zheng, Kui Qian, Juyang Weng, Zhengyou Zhang |
IJCNN | 4 |
| 2013 | "Pattern Recognition" special issue: Sparse representation for event recognition in video surveillance
Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang, Lisa M. Brown |
Pattern Recognit. | 4 |
| 2013 | Robust Part-Based Hand Gesture Recognition Using Kinect SensorabstractThe recently developed depth sensors, e.g., the Kinect sensor, have provided new opportunities for human-computer interaction (HCI). Although great progress has been made by leveraging the Kinect sensor, e.g., in human body tracking, face recognition and human action recognition, robust hand gesture recognition remains an open problem. Compared to the entire human body, the hand is a smaller object with more complex articulations and more easily affected by segmentation errors. It is thus a very challenging problem to recognize hand gestures. This paper focuses on building a robust part-based hand gesture recognition system using Kinect sensor. To handle the noisy hand shapes obtained from the Kinect sensor, we propose a novel distance metric, Finger-Earth Mover's Distance (FEMD), to measure the dissimilarity between hand shapes. As it only matches the finger parts while not the whole hand, it can better distinguish the hand gestures of slight differences. The extensive experiments demonstrate that our hand gesture recognition system is accurate (a 93.2% mean accuracy on a challenging 10-gesture dataset), efficient (average 0.0750 s per frame), robust to hand articulations, distortions and orientation or scale changes, and can work in uncontrolled environments (cluttered backgrounds and lighting conditions). The superiority of our system is further demonstrated in two real-life HCI applications. Zhou Ren, Junsong Yuan 0001, Jingjing Meng, Zhengyou Zhang |
IEEE Trans. Multim. | 4 |
| 2013 | Model-based hand pose estimation via spatial-temporal hand parsing and 3D fingertip localization
Hui Liang 0003, Junsong Yuan 0001, Daniel Thalmann, Zhengyou Zhang |
Vis. Comput. | 4 |
| 2012 | Mining noisy tagging from multi-label spaceabstractIn this paper we study the problem of mining noisy tagging. Most of the existing discriminative classification methods to this problem only consider one tag at a time as the classification target, and completely ignore the rest of the given tags at the same time. In this paper we argue that all the given multiple tags can be utilized simultaneously as an additional feature and the information contained in the multi-label space can be taken advantage of to improve the performance of the classification. We first propose a novel distance measure to compute the distance between instances in the multi-label space. Then we propose several novel methods to incorporate the information of the multi-label space into the discriminative classification methods in one view learning or in two views learning to solve a general multi-label classification problem and to mitigate the influence of the noise in the classification. We apply the proposed solutions to the problem with a more specific context - noisy image annotation, and evaluate the proposed methods on a standard dataset from the related literature. Experiments show that they are superior to the peer methods in the existing literature on solving the problem of mining noisy tagging. Zhongang Qi, Ming Yang 0012, Zhongfei Zhang, Zhengyou Zhang |
CIKM | 4 |
| 2012 | 3D scene reconstruction by multiple structured-light based commodity depth camerasabstractCommodity depth cameras have attracted a lot of research interest recently, in particular the structured-light based Kinect cameras available on the mass market. One important application of such cameras is 3D scene reconstruction and view synthesis. However, a single depth camera often has limited field of view and there is missing depth information when synthesizing a virtual view from a new viewpoint. In this paper, we study the problem of 3D scene reconstruction from multiple structured-light based depth cameras. Since multiple cameras may cause severe interference in the regions where the projected light overlaps, we present a novel planesweeping based algorithm to handle such interference. The proposed algorithm takes into account the correlation between multiple projectors and the infrared images as well as the correlation between the infrared images, thereby recovering the depth information for both overlapped and non-overlapped regions. Simulation results demonstrate that the proposed solution is very effective on various scenes. Cha Zhang, Wenwu Zhu 0001, Zhengyou Zhang, Zixiang Xiong, Philip A. Chou |
ICASSP | 4 |
| 2012 | Virtual View Reconstruction Using Temporal InformationabstractThe most significant problem in generating virtual views from a limited number of video camera views is handling areas that have become dis-occluded by shifting the virtual view away from the camera view. We propose using temporal information to address this problem, based on the notion that dis-occluded areas may have been seen by some camera in some previous frames. We formulate the problem as one of estimating the underlying state of the object in a stochastic dynamical system, given a sequence of observations. We apply the formulation to improving the visual quality of virtual views generated from a single “color plus depth” camera, and show that our algorithm achieves better results than depth image based rendering using standard inpainting. Shujie Liu 0001, Philip A. Chou, Cha Zhang, Zhengyou Zhang, Chang Wen Chen |
ICME | 4 |
| 2012 | Multi-view learning from imperfect taggingabstractIn many real-world applications, tagging is imperfect: incomplete, inconsistent, and error-prone. Solutions to this problem will generate societal and technical impacts. In this paper, we investigate this arguably new problem: learning from imperfect tagging. We propose a general and effective learning scheme called the Multi-view Imperfect Tagging Learning (MITL) to this problem. The main idea of MITL lies in extracting the information of the imperfectly tagged training dataset from multiple views to differentiate the data points in the role of classification. Further, a novel discriminative classification method is proposed under the framework of MITL, which explicitly makes use of the given multiple labels simultaneously as an additional feature to deliver a more effective classification performance than the existing literature where one label is considered at a time as the classification target while the rest of the given labels are completely ignored at the same time. The proposed methods can not only complete the incomplete tagging but also denoise the noisy tagging through an inductive learning. We apply the general solution to the problem with a more specific context - imperfect image annotation, and evaluate the proposed methods on a standard dataset from the related literature. Experiments show that they are superior to the peer methods on solving the problem of learning from imperfect tagging in cross-media. Zhongang Qi, Ming Yang 0012, Zhongfei Zhang, Zhengyou Zhang |
ACM Multimedia | 4 |
| 2012 | Auditory augmented reality: Object sonification for the visually impairedabstractAugmented reality applications have focused on visually integrating virtual objects into real environments. In this paper, we propose an auditory augmented reality, where we integrate acoustic virtual objects into the real world. We sonify objects that do not intrinsically produce sound, with the purpose of revealing additional information about them. Using spatialized (3D) audio synthesis, acoustic virtual objects are placed at specific real-world coordinates, obviating the need to explicitly tell the user where they are. Thus, by leveraging the innate human capacity for 3D sound source localization and source separation, we create an audio natural user interface. In contrast with previous work, we do not create acoustic scenes by transducing low-level (for instance, pixel-based) visual information. Instead, we use computer vision methods to identify high-level features of interest in an RGB-D stream, which are then sonified as virtual objects at their respective real-world coordinates. Since our visual and auditory senses are inherently spatial, this technique naturally maps between these two modalities, creating intuitive representations. We evaluate this concept with a head-mounted device, featuring modes that sonify flat surfaces, navigable paths and human faces. Flavio P. Ribeiro, Dinei A. F. Florêncio, Philip A. Chou, Zhengyou Zhang |
MMSP | 4 |
| 2012 | Introduction to the Special Issue on Mobile Vision
Gang Hua 0001, Yun Fu 0001, Matthew Turk 0001, Marc Pollefeys, Zhengyou Zhang |
Int. J. Comput. Vis. | 5 |
| 2012 | Societally connected multimedia across culturesabstractThe advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu |
J. Zhejiang Univ. Sci. C | 2 |
| 2012 | Hierarchical Filtered Motion for Action Recognition in Crowded VideosabstractAction recognition with cluttered and moving background is a challenging problem. One main difficulty lies in the fact that the motion field in an action region is contaminated by the background motions. We propose a hierarchical filtered motion (HFM) method to recognize actions in crowded videos by the use of motion history image (MHI) as basic representations of motion because of its robustness and efficiency. First, we detect interest points as the two-dimensional Harris corners with recent motion, e.g., locations with high intensities in the MHI. Then, a global spatial motion smoothing filter is applied to the gradients of the MHI to eliminate isolated unreliable or noisy motions. At each interest point, a local motion field filter is applied to the smoothed gradients of the MHI by computing structure proximity between any pixel in the local region and the interest point. Thus, the motion at a pixel is enhanced or weakened based on its structure proximity with the interest point. To validate its effectiveness, we characterize the spatial and temporal features by histograms of oriented gradient in the intensity image and the MHI, respectively, and use a Gaussian-mixture-model-based classifier for action recognition. The performance of the proposed approach achieves the state-of-the-art results on the KTH dataset that has clean background. More importantly, we perform cross-dataset action classification and detection experiments, where the KTH dataset is used for training, while the microsoft research (MSR) action dataset II that consists of crowded videos with people moving in the background is used for testing. Our experiments show that the proposed HFM method significantly outperforms existing techniques. Yingli Tian, Liangliang Cao, Zicheng Liu 0001, Zhengyou Zhang |
IEEE Trans. Syst. Man Cybern. Part C | 4 |
| 2011 | What did i miss?: in-meeting review using multimodal accelerated instant replay (air) conferencingabstractPeople sometimes miss small parts of meetings and need to quickly catch up without disrupting the rest of the meeting. We developed an Accelerated Instant Replay (AIR) Conferencing system for videoconferencing that enables users to catch up on missed content while the meeting is ongoing. AIR can replay parts of the conference using four different modalities: audio, video, conversation transcript, and shared workspace. We performed two studies to evaluate the system. The first study explored the benefit of AIR catch-up during a live meeting. The results showed that when the full videoconference was reviewed (i.e., all four modalities) at an accelerated rate, users were able to correctly recall a similar amount of information as when listening live. To better understand the benefit of full review, a follow-up study more closely examined the benefits of each of the individual modalities. The results show that users (a) preferred using audio along with any other modality to using audio alone, (b) were most confident and performed best when audio was reviewed with all other modalities, (c) compared to audio-only, had better recall of facts and explanations when reviewing audio together with the shared workspace and transcript modalities, respectively, and (d) performed similarly with audio-only and audio with video review. Sasa Junuzovic, Kori Inkpen, Rajesh Hegde, Zhengyou Zhang, John C. Tang, Christopher Brooks 0001 |
CHI | 4 |
| 2011 | Towards ideal window layouts for multi-party, gaze-aware desktop videoconferencing
Sasa Junuzovic, Kori Inkpen, Rajesh Hegde, Zhengyou Zhang |
Graphics Interface | 4 |
| 2011 | Realistic audio in immersive video conferencingabstractWith increasing computation power, network bandwidth, and improvements in display and capture technologies, fully immersive conferencing and tele-immersion is becoming ever closer to reality. Outside of video, one of the key components needed is high quality spatialized audio. This paper presents an implementation of a relatively low complexity, simple solution which allows realistic audio spatialization of arbitrary positions in a 3D video conference. When combined with pose tracking, it also allows the audio to change relative to which position on the screen the viewer is looking at. Sanjeev Mehrotra, Wei-Ge Chen, Zhengyou Zhang, Philip A. Chou |
ICME | 3 |
| 2011 | A novel see-through screen based on weave fabricsabstractSee-through screens (STS) have found important applications in remote collaboration systems to enhance non-verbal communication and gaze awareness. Existing STS designs often sacrifice the display quality significantly, rendering low-contrast images that discount the overall user experience. In this paper, we present a novel see-through screen solution based on weave fabrics. Such fabrics are known to be acoustically transparent and used to build professional projection screens for Hollywood studios. We place a cam-era immediately behind the screen and synchronize it with a 120Hz projector to perform time-multiplexing display and video capture. By focusing the camera at the user 4–5 feet away from the screen, the image of the weave fabric will be severely blurred. We present the imaging principle of the setup, and derive image processing techniques to enhance the quality of the captured video. The overall system is low cost, has much better display quality than existing systems, and can be used to build wall-size see-through screens for various applications. Cha Zhang, Ruigang Yang, Tim Large, Zhengyou Zhang |
ICME | 4 |
| 2011 | Calibration between depth and color sensors for commodity depth camerasabstractCommodity depth cameras have created many interesting new applications in the research community recently. These applications often require the calibration information between the color and the depth cameras. Traditional checkerboard based calibration schemes fail to work well for the depth camera, since its corner features cannot be reliably detected in the depth image. In this paper, we present a maximum likelihood solution for the joint depth and color calibration based on two principles. First, in the depth image, points on the checker board shall be co-planar, and the plane is known from color camera calibration. Second, additional point correspondences between the depth and color images may be manually specified or automatically established to help improve calibration accuracy. Uncertainty in depth values has been taken into account systematically. The proposed algorithm is reliable and accurate, as demonstrated by extensive experimental results on simulated and real-world examples. Cha Zhang, Zhengyou Zhang |
ICME | 2 |
| 2011 | Mining partially annotated imagesabstractIn this paper, we study the problem of mining partially annotated images. We first define what the problem of mining partially annotated images is, and argue that in many real-world applications annotated images are typically partially annotated and thus that the problem of mining partially annotated images exists in many situations. We then propose an effective solution to this problem based on a statistical model we have developed called the Semi-Supervised Correspondence Hierarchical Dirichlet Process (SSCHDP). The main idea of this model lies in exploiting the information pertaining to partially annotated images or even unannotated images to achieve semi-supervised learning under the HDP structure. We apply this model to completing the annotations appropriately for partially annotated images in the training data and then to predicting the annotations appropriately and completely for all the unannotated images either in the training data or in any unseen data beyond the training process. Experiments show that SSC-HDP is superior to the peer models from the recent literature when they are applied to solving the problem of mining partially annotated images. Zhongang Qi, Ming Yang 0012, Zhongfei Zhang, Zhengyou Zhang |
KDD | 4 |
| 2011 | Innovating the multimedia experienceabstractIn this panel, each panelist will present their view of the current state-of-the-art of research and product innovations in the three major areas of multimedia experience: visual, auditory and gaming. We will discuss examples of innovation that enhance the consumption and sharing of multimedia (video, audio, graphics etc.) and thus increase quality of user experience. Another major focus of this panel is to open the discussion on how to innovate new multimedia user experiences. Khaled El-Maleh, Haohong Wang, Susie J. Wee, Hong Heather Yu, James D. Johnston, Zhengyou Zhang |
ACM Multimedia | 6 |
| 2011 | Modeling and representing events in multimediaabstractThis paper presents an overview of the Joint Workshop on Modeling and Representing Events (JMRE), which is held as part of ACM Multimedia 2011. JMRE is concerned with the understanding of events from multimedia, and with using events in order to better organize and consume multimedia. Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang |
ACM Multimedia | 8 |
| 2011 | Robust hand gesture recognition with kinect sensorabstractHand gesture based Human-Computer-Interaction (HCI) is one of the most natural and intuitive ways to communicate between people and machines, since it closely mimics how human interact with each other. In this demo, we present a hand gesture recognition system with Kinect sensor, which operates robustly in uncontrolled environments and is insensitive to hand variations and distortions. Our system consists of two major modules, namely, hand detection and gesture recognition. Different from traditional vision-based hand gesture recognition methods that use color-markers for hand detection, our system uses both the depth and color information from Kinect sensor to detect the hand shape, which ensures the robustness in cluttered environments. Besides, to guarantee its robustness to input variations or the distortions caused by the low resolution of Kinect sensor, we apply a novel shape distance metric called Finger-Earth Mover's Distance (FEMD) for hand gesture recognition. Consequently, our system operates accurately and efficiently. In this demo, we demonstrate the performance of our system in two real-life applications, arithmetic computation and rock-paper-scissors game. Zhou Ren, Jingjing Meng, Junsong Yuan 0001, Zhengyou Zhang |
ACM Multimedia | 4 |
| 2011 | Robust hand gesture recognition based on finger-earth mover's distance with a commodity depth cameraabstractThe recently developed depth sensors, e.g., the Kinect sensor, have provided new opportunities for human-computer interaction (HCI). Although great progress has been made by leveraging the Kinect sensor, e.g. in human body tracking and body gesture recognition, robust hand gesture recognition remains an open problem. Compared to the entire human body, the hand is a smaller object with more complex articulations and more easily affected by segmentation errors. It is thus a very challenging problem to recognize hand gestures. This paper focuses on building a robust hand gesture recognition system using the Kinect sensor. To handle the noisy hand shape obtained from the Kinect sensor, we propose a novel distance metric for hand dissimilarity measure, called Finger-Earth Mover's Distance (FEMD). As it only matches fingers while not the whole hand shape, it can better distinguish hand gestures of slight differences. The extensive experiments demonstrate the accuracy, efficiency, and robustness of our hand gesture recognition system. Zhou Ren, Junsong Yuan 0001, Zhengyou Zhang |
ACM Multimedia | 3 |
| 2011 | Interpolation of combined head and room impulse response for audio spatializationabstractAudio spatialization is becoming an important part of creating realistic experiences needed for immersive video conferencing and gaming. Using a combined head and room impulse response (CHRIR) has been recently proposed as an alternative to using separate head related transfer functions (HRTF) and room impulse responses (RIR). Accurate measurements of the CHRIR at various source and listener locations and orientations are needed to perform good quality audio spatialization. However, it is infeasible to accurately measure or model the CHRIR for all possible locations and orientations. Therefore, low-complexity and accurate interpolation techniques are needed to perform audio spatialization in real-time. In this paper, we present a frequency domain interpolation technique which naturally interpolates the interaural level difference (ILD) and interaural time difference (ITD) for each frequency component in the spectrum. The proposed technique allows for an accurate and low-complexity interpolation of the CHRIR as well as allowing for a low-complexity audio spatialization technique which can be used for both headphones as well as loudspeakers. Sanjeev Mehrotra, Wei-Ge Chen, Zhengyou Zhang |
MMSP | 3 |
| 2011 | Low-complexity, near-lossless coding of depth maps from kinect-like depth camerasabstractDepth cameras are gaining interest rapidly in the market as depth plus RGB is being used for a variety of applications ranging from foreground/background segmentation, face tracking, activity detection, and free viewpoint video rendering. In this paper, we present a low-complexity, near-lossless codec for coding depth maps. This coding requires no buffering of video frames, is table-less, can encode or decode a frame in close to 5ms with little code optimization, and provides between 7:1 to 16:1 compression ratio for near-lossless coding of 16-bit depth maps generated by the Kinect camera. Sanjeev Mehrotra, Zhengyou Zhang, Qin Cai, Cha Zhang, Philip A. Chou |
MMSP | 2 |
| 2011 | ViewMark: An interactive videoconferencing system for mobile devicesabstractViewMark, a server-client based interactive mobile videoconferencing system is proposed in this paper to enhance the remote meeting experience for mobile users. Compared with the state-of-the-art mobile videoconferencing technology, ViewMark is novel in allowing a mobile user to interactively change the viewpoint of the remote video, create viewmarks, and hear with spatial audio. In addition, ViewMark also streams the screen of the presentation slides to mobile devices. In this paper, we introduce the system design of ViewMark in details, compare the devices that can be used to implement interactive videoconferencing, and demonstrate the prototype system we have built on Windows Mobile platform. Shu Shi, Zhengyou Zhang |
MMSP | 2 |
| 2011 | An effecive night video enhancement algorithmabstractNight video enhancement is important for video surveillance since many objects or activities of interest occur in a dark environment which cannot be seen easily without enhancement. In this paper, we discuss several problems of existing techniques for illumination-fusion based night video enhancement, which fuses video frames from day-time backgrounds and night-time video. We then present a simple enhancement algorithm without these problems. The algorithm uses an additive enhancement term with foreground object extraction and constrained low-passed object illumination to avoid light-inversion and sensitivity problems and to reduce ghost patterns. Experimental results show the effectiveness and robustness of the proposed algorithm. Yunbo Rao, Zhong-Ho Chen, Ming-Ting Sun, Yu-Feng Hsu, Zhengyou Zhang |
VCIP | 5 |
| 2011 | Introduction to the ICME2010 Special IssueabstractThe 15 papers in this special issue are extended versions of papers presented at the 2010 IEEE International Conference on Multimedia and Expo (ICME), held in Singapore on July 19-23, 2010. These papers cover a wide range of topics in multimedia including user interface, content understanding, mobility, 3-D processing, storage, and forensics. Zicheng Liu 0001, Ming-Ting Sun, Chia-Wen Lin, Zhengyou Zhang, Zhu Liu 0001, Homer H. Chen, Yap-Peng Tan, Oscar C. Au |
IEEE Trans. Multim. | 4 |
| 2010 | Exploring spatialized audio & video for distributed conversationsabstractPrevious work has demonstrated the benefits of spatial audio conferencing over monophonic when listening to a group conversation. In this paper we examined three-way distributed conversations while varying the presence of spatial video and audio. Our results demonstrate significant benefits to adding spatialized video to an audio conference. Specifically, users perceived that the conversations were of higher quality, they were more engaged, and they were better able to keep track of the conversation. In contrast, no significant benefits were found when mono audio was replaced by spatialized audio. The results of this work are important in that they provide strong evidence for continued exploration of spatialized video, and also suggest that the benefits of spatialized audio may have less of an impact when video is also spatialized. Kori Inkpen, Rajesh Hegde, Mary Czerwinski, Zhengyou Zhang |
CSCW | 4 |
| 2010 | 3D Deformable Face Tracking with a Commodity Depth Camera
Qin Cai, David Gallup, Cha Zhang, Zhengyou Zhang |
ECCV (3) | 4 |
| 2010 | Action detection using multiple spatial-temporal interest point featuresabstractThis paper considers the problem of detecting actions from cluttered videos. Compared with the classical action recognition problem, this paper aims to estimate not only the scene category of a given video sequence, but also the spatial-temporal locations of the action instances. In recent years, many feature extraction schemes have been designed to describe various aspects of actions. However, due to the difficulty of action detection, e.g., the cluttered background and potential occlusions, a single type of features cannot solve the action detection problems perfectly in cluttered videos. In this paper, we attack the detection problem by combining multiple Spatial-Temporal Interest Point (STIP) features, which detect salient patches in the video domain, and describe these patches by feature of local regions. The difficulty of combining multiple STIP features for action detection is two folds: First, the number of salient patches detected by different STIP methods varies across different salient patches. How to combine such features is not considered by existing fusion methods. Second, the detection in the videos should be efficient, which excludes many slow machine learning algorithms. To handle these two difficulties, we propose a new approach which combines Gaussian Mixture Model with Branch-and-Bound search to efficiently locate the action of interest. We build a new challenging dataset for our action detection task, and our algorithm obtains impressive results. On classical KTH dataset, our method outperforms the state-of-the-art methods. Liangliang Cao, Yingli Tian, Zicheng Liu 0001, Benjamin Z. Yao, Zhengyou Zhang, Thomas S. Huang |
ICME | 5 |
| 2010 | AIR conferencing: accelerated instant replay for in-meeting multimodal reviewabstractWhen people attend meetings they may miss parts of the discussion if they, for example, step out to take a phone call, go to the bathroom, or have a momentary lapse in concentration. As a result, they may need to catch up on what they missed upon returning to the meeting. Asking other attendees for a recap is often disruptive. To avoid such disruptions, we have developed an Accelerated Instant Replay (AIR) Conferencing system for videoconferencing that enables participants to privately catch up to an ongoing meeting. We explored several mechanisms where the meeting content is replayed at an accelerated rate so that the participants can catch up to the live discussion reasonably quickly. Kori Inkpen, Rajesh Hegde, Sasa Junuzovic, Christopher Brooks 0001, John C. Tang, Zhengyou Zhang |
ACM Multimedia | 6 |
| 2010 | Overview of ACM international workshop on connected multimediaabstractFollowing the very first international workshop on connected multimedia held in Hangzhou, China, in October of 2009 jointly sponsored by US National Science Foundation and Zhejiang University of China, this is the very first ACM International Workshop on Connected Multimedia in conjunction with ACM International Conference on Multimedia held in Florence, Italy, in October of 2010. In this workshop overview, we first define what we mean by connected multimedia, and then briefly overview the program of this workshop. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang |
ACM Multimedia | 2 |
| 2010 | Enhancing stereophonic teleconferencing with microphone arrays through sound field warpingabstractIt has been proven that spatial audio enhances the realism of sound for teleconferencing. Previously, solutions have been proposed for multiparty conferencing where each remote participant is assumed to have his/her own microphone, and for conferencing between two rooms where the microphones in one room are connected to the equal number of loudspeakers in the other room. Either approach has its limitations. Hence, we propose a new scheme to improve stereophonic conferencing experience through an innovative use of microphone arrays. Instead of operating in the default mode where a single channel is produced using spatial filtering, we propose to transmit all channels forming a collection of spatial samples of the sound field. Those samples are warped appropriately at the remote site, and are spatialized together with audio streams from other remote sites if any, to produce the perception of a virtual sound field. Real-world audio samples are provided to showcase the proposed technique. The informal listening test shows that majority of the users prefer the new experience. Wei-Ge Chen, Zhengyou Zhang |
MMSP | 2 |
| 2010 | Group Event Detection With a Varying Number of Group Members for Video SurveillanceabstractThis paper presents a novel approach for automatic recognition of group activities for video surveillance applications. We propose to use a group representative to handle the recognition with a varying number of group members, and use an asynchronous hidden Markov model (AHMM) to model the relationship between people. Furthermore, we propose a group activity detection algorithm which can handle both symmetric and asymmetric group activities, and demonstrate that this approach enables the detection of hierarchical interactions between people. Experimental results show the effectiveness of our approach. Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Efficient Scale-Space Spatiotemporal Saliency Tracking for Distortion-Free Video Retargeting
Gang Hua 0001, Cha Zhang, Zicheng Liu 0001, Zhengyou Zhang, Ying Shan |
ACCV (2) | 4 |
| 2009 | Boosted multi-task learning for face verification with applications to web image and video searchabstractFace verification has many potential applications including filtering and ranking image/video search results on celebrities. Since these images/videos are taken under uncontrolled environments, the problem is very challenging due to dramatic lighting and pose variations, low resolutions, compression artifacts, etc. In addition, the available number of training images for each celebrity may be limited, hence learning individual classifiers for each person may cause overfitting. In this paper, we propose two ideas to meet the above challenges. First, we propose to use individual bins, instead of whole histograms, of Local Binary Patterns (LBP) as features for learning, which yields significant performance improvements and computation reduction in our experiments. Second, we present a novel Multi-Task Learning (MTL) framework, called Boosted MTL, for face verification with limited training data. It jointly learns classifiers for multiple people by sharing a few boosting classifiers in order to avoid overfitting. The effectiveness of Boosted MTL and LBP bin features is verified with a large number of celebrity images/videos from the web. Cha Zhang, Zhengyou Zhang |
CVPR | 3 |
| 2009 | Multimodal collaboration and human-computer interactionabstractThe research effort at Microsoft research on multimodal collaboration and human-computer interaction aims at developing tools that allow people across geographically distributed sites to interact collaboratively with immersive experience. Our prototype systems consist of cameras, displays, speakers, microphones, computer controllable lights, and/or input devices such as touch sensitive surface, stylus, keyboard, and mouse. They require real-time processing a huge amount of data, such as foreground-background substraction, region-of-interest extraction, color estimation and correction, speaker detection, stereo matching, 3D reconstruction and rendering, without mentioning audio and video encoding and decoding possibly involving multiple microphones and cameras. Some of the processing can be easily parallelizable through general-purpose computation on graphics processing units (GPGPU) or on a multi-core processor machine, while others are not so trivial. In this extended summary, the author describe two projects: Visual echo cancellation in shared tele-collaborative space, and distributed meeting capture and broadcasting system. During the talk, the author will also present two recent projects: personal telepresence station and situated interaction. Zhengyou Zhang |
ICME | 1 |
| 2009 | Group Event Detection for Video SurveillanceabstractThis paper presents a novel approach for automatic recognition of group activities for video surveillance applications. We propose to use a group representative to handle the recognition with flexible or varying number of group members, and use an asynchronous hidden Markov model (AHMM) to model the relationship between two people. Furthermore, we propose a group activity detection algorithm which can handle symmetric and asymmetric group activities, and demonstrate that this approach enables the detection of hierarchical interactions between people. Experimental results show the effectiveness of our approach. Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang |
ISCAS | 4 |
| 2009 | Highly realistic audio spatialization for multiparty conferencing using headphonesabstractIt is known that during multi-party conferencing spatialized audio which maps remote participants' voices to distinct virtual locations improves the listening experience. In this paper, we consider the case when the audio is rendered through headphones due to e.g. privacy reasons. Although existing headphone spatial audio techniques abound, most lack the desired realism dictated by listeners' expectation of naturalness in audio conferencing. In light of the situation, we propose a novel approach of spatial audio processing for headphones, where we measure the combined head and room impulse responses (CHRIRs) in an actual physical setting which are then directly used to spatialize remote participants' voices. Through proper processing of the CHRIRs, our solution is able to offer a higher degree of realism, closely approximating binaural recordings. We note that the success, however, comes with certain limitations and presented solutions to mitigate the shortcomings of the proposed technique. As a result, a user can adapt the source localization and the degree of reverberation to her own subjective preferences. We also show that the computation load can be effectively reduced and becomes very reasonable for modern computer hardware. Wei-Ge Chen, Zhengyou Zhang |
MMSP | 2 |
| 2009 | Face Relighting from a Single Image under Arbitrary Unknown Lighting ConditionsabstractIn this paper, we present a new method to modify the appearance of a face image by manipulating the illumination condition, when the face geometry and albedo information is unknown. This problem is particularly difficult when there is only a single image of the subject available. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using a spherical harmonic representation. Moreover, morphable models are statistical ensembles of facial properties such as shape and texture. In this paper, we integrate spherical harmonics into the morphable model framework by proposing a 3D spherical harmonic basis morphable model (SHBMM). The proposed method can represent a face under arbitrary unknown lighting and pose simply by three low-dimensional vectors, i.e., shape parameters, spherical harmonic basis parameters, and illumination coefficients, which are called the SHBMM parameters. However, when the image was taken under an extreme lighting condition, the approximation error can be large, thus making it difficult to recover albedo information. In order to address this problem, we propose a subregion-based framework that uses a Markov random field to model the statistical distribution and spatial coherence of face texture, which makes our approach not only robust to extreme lighting conditions, but also insensitive to partial occlusions. The performance of our framework is demonstrated through various experimental results, including the improved rates for face recognition under extreme lighting conditions. Yang Wang 0001, Lei Zhang 0002, Zicheng Liu 0001, Gang Hua 0001, Zhengyou Zhang, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2009 | Active Lighting for Video ConferencingabstractIn consumer video conferencing, lighting conditions are usually not ideal thus the image qualities are poor. Lighting affects image quality on two aspects: brightness and skin tone. While there has been much research on improving the brightness of the captured images including contrast enhancement and noise removal (which can be thought of as components for brightness improvement), little attention has been paid to the skin tone aspect. In contrast, it is a common knowledge for professional stage lighting designers that lighting affects not only the brightness but also the color tone which plays a critical role in the perceived look of the host and the mood of the stage scene. Inspired by stage lighting design, we propose an active lighting system which automatically adjusts the lighting so that the image looks visually appealing. The system consists of computer controllable light emitting diode light sources of different colors so that it improves not only the brightness but also the skin tone of the face. Given that there is no quantitative formula on what makes a good skin tone, we use a data driven approach to learn a good skin tone model from a collection of photographs taken by professional photographers. We have developed a working system and conducted user studies to validate our approach. Mingxuan Sun 0001, Zicheng Liu 0001, Jingyu Qiu, Zhengyou Zhang, Mike Sinclair |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Taylor expansion based classifier adaptation: Application to person detectionabstractBecause of the large variation across different environments, a generic classifier trained on extensive data-sets may perform sub-optimally in a particular test environment. In this paper, we present a general framework for classifier adaptation, which improves an existing generic classifier in the new test environment. Viewing classifier learning as a cost minimization problem, we perform classifier adaptation by combining the cost function on the old data-sets with the cost function on the data-set collected from the new environment. The former term is further approximated with its second order Taylor expansion to reduce the amount of information that needs to be saved for adaptation. Unlike traditional approaches that are often designed for a specific application or classifier, our scheme is applicable to various types of classifiers and user labels. We demonstrate this property on two popular classifiers (logistic regression and boosting), while using two types of user labels (direct labels and similarity labels). Extensive experiments conducted for the task of person detection in conference-room environments show that significant performance improvement can be achieved with our proposed method. Cha Zhang, Raffay Hamid, Zhengyou Zhang |
CVPR | 3 |
| 2008 | Why does PHAT work well in lownoise, reverberative environments?abstractAmong many existing time difference of arrival (TDOA) based sound source localization (SSL) algorithms, the Phase Transform (PHAT) is extremely popular for its excellent performance in low noise environments, even under relatively heavy reverberation. However, PHAT was developed as a heuristic approach and its working principle has not been completely understood. In this paper, we present the relationship between PHAT and a maximum likelihood (ML) framework for multi-microphone sound source localization. We show that when the environment noise approaches zero, PHAT is indeed a special case of the ML algorithm, which explains its good performance under low noise environments. In addition, we show that as long as the noise stays low, PHAT remains optimal in ML sense even when the room reverberation is heavy, which explains its robustness over reverberation. Cha Zhang, Dinei A. F. Florêncio, Zhengyou Zhang |
ICASSP | 3 |
| 2008 | Human activity recognition for video surveillanceabstractThis paper presents a novel approach for automatic recognition of human activities from video sequences. We first group features with high correlations into Category Feature Vectors (CFVs). Each activity is then described by a combination of GMMs (Gaussian Mixture Models) with each GMM representing the distribution of a CFV. We show that this approach offers flexibility to add new events and to deal with the problem of lacking training data for building models for unusual events. For improving the recognition accuracy, a Confident-Frame-based Recognizing algorithm (CFR) is proposed to recognize the human activity, where the video frames which have high confidence for recognition an activity (Confident-Frames) are used as a specialized model for classifying the rest of the video frames. Experimental results show the effectiveness of the proposed approach. Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang |
ISCAS | 4 |
| 2008 | Requirements and recommendations for an enhanced meeting viewing experienceabstractWe have found that viewing recorded meetings using traditional meeting viewers whose interfaces consist of an automatic speaker and a fixed context view does not provide sufficient information and control to the users. In particular, a survey of users who watch meeting recordings on a regular basis revealed that it is also useful to provide (1) speaker-related information, including who the speaker is talking to, looking at, and being interrupted by, and (2) more control of the interface, including changing the relative sizes of the speaker and context views and navigating within the context view. We present a 3D interface prototype designed specifically to meet these requirements when viewing recorded meetings. We describe in detail the results of a user study comparing the effectiveness of the new and traditional style interfaces with respect to these requirements. Based on this study, we present a set of guidelines for future interfaces. Sasa Junuzovic, Rajesh Hegde, Zhengyou Zhang, Philip A. Chou, Zicheng Liu 0001, Cha Zhang |
ACM Multimedia | 3 |
| 2008 | Graphical modeling and decoding of human actionsabstractThis paper presents a graphical model for learning and recognizing human actions. Specifically, we propose to encode actions in a weighted directed graph, referred to as action graph, where nodes of the graph represent salient postures that are used to characterize the actions and shared by all actions. The weight between two nodes measures the transitional probability between the two postures. An action is encoded as one or multiple paths in the action graph. The salient postures are modeled using Gaussian Mixture Models (GMM). Both the salient postures and action graph are automatically learned from training samples through unsupervised clustering and expectation and maximization (EM) algorithm. Experimental results have verified the performance of the proposed model, its tolerance to noise and viewpoints and its robustness across different subjects and datasets. Wanqing Li 0001, Zhengyou Zhang, Zicheng Liu 0001 |
MMSP | 2 |
| 2008 | Semantic saliency driven camera control for personal remote collaborationabstractThis paper presents a camera combo system for personal remote collaboration applications. The system consists of two different cameras. One camera has a wide field of view, and the other can pan/tilt/zoom (PTZ) based on analysis of the images captured by the wide angle camera. Unlike traditional approaches which usually drive the PTZ camera to follow the person or his/her head, our system is capable of capturing general objects of interest in remote collaboration. For instance, when the user raises something trying to show it to the remote person, our system will automatically position the PTZ camera to zoom in at the object. At the core of our system is a semantic saliency map that overcomes many limitations of low-level saliency maps computed from preliminary image features. We demonstrate how such a semantic saliency map can be computed through contextual analysis, sign analysis and transitional analysis, and how it can be used for PTZ camera control with a novel information loss optimization based virtual director. The effectiveness of the proposed method is demonstrated with real-world sequences. Cha Zhang, Zicheng Liu 0001, Zhengyou Zhang |
MMSP | 3 |
| 2008 | Robust and Accurate Visual Echo Cancelation in a Full-duplex Projector-Camera SystemabstractIn this paper we study the problem of "visual echo" in a full-duplex projector-camera system for telecollaboration applications. Visual echo is defined as the appearance of projected contents observed by the camera. It can potentially saturate the projected contents, similar to audio echo in telephone conversation. Our approach to visual echo cancellation includes an offline calibration procedure that records the geometric and photometric transfer between the projector and the camera in a look-up table. During run-time, projected contents in the captured video are identified using the calibration information and suppressed, therefore achieving the goal of cancelling visual echo. Our approach can accurately handle full-color images under arbitrary reflectance of display surfaces and photometric response of the projector or camera. It is robust to geometric registration errors and quantization effects and is therefore particularly effective for high-frequency contents such as texts and hand drawings. We demonstrate the effectiveness of our approach with a variety of real images in a full-duplex projector-camera system. Miao Liao, Ruigang Yang, Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Multisensory processing for speech enhancement and magnitude-normalized spectra for speech modeling
Amarnag Subramanya, Zhengyou Zhang, Zicheng Liu 0001, Alex Acero |
Speech Commun. | 2 |
| 2008 | Expandable Data-Driven Graphical Modeling of Human Actions Based on Salient PosturesabstractThis paper presents a graphical model for learning and recognizing human actions. Specifically, we propose to encode actions in a weighted directed graph, referred to as action graph, where nodes of the graph represent salient postures that are used to characterize the actions and are shared by all actions. The weight between two nodes measures the transitional probability between the two postures represented by the two nodes. An action is encoded asoneor multiple paths in the action graph. The salient postures are modeled using Gaussian mixture models (GMMs). Both the salient postures and action graph are automatically learned from training samples through unsupervised clustering and expectation and maximization (EM) algorithm. The proposed action graph not only performs effective and robust recognition of actions, but it can also be expanded efficiently with new actions. An algorithm is also proposed for adding a new action to a trained action graph without compromising the existing action graph. Extensive experiments on widely used and challenging data sets have verified the performance of the proposed methods, its tolerance to noise and viewpoints, its robustness across different subjects and data sets, as well as the effectiveness of the algorithm for learning new actions. Wanqing Li 0001, Zhengyou Zhang, Zicheng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Activity Recognition Using a Combination of Category Components and Local Models for Video SurveillanceabstractThis paper presents a novel approach for automatic recognition of human activities for video surveillance applications. We propose to represent an activity by a combination of category components and demonstrate that this approach offers flexibility to add new activities to the system and an ability to deal with the problem of building models for activities lacking training data. For improving the recognition accuracy, a confident-frame-based recognition algorithm is also proposed, where the video frames with high confidence for recognizing an activity are used as a specialized local model to help classify the remainder of the video frames. Experimental results show the effectiveness of the proposed approach. Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Camera Calibration With Three Noncollinear Points Under Special MotionsabstractPlane-based (2-D) camera calibration is becoming a hot research topic in recent years because of its flexibility. However, at least four image points are needed in every view to denote the coplanar feature in the 2-D camera calibration. Can we do the camera calibration by using the calibration object that only has three points? Some 1-D camera calibration techniques use the setup of three collinear points with known distances, but it is a kind of special conditions of calibration object setup. How about the general setup-three noncollinear points? We propose a new camera calibration algorithm based on the calibration objects with three noncollinear points. Experiments with simulated data and real images are carried out to verify the theoretical correctness and numerical robustness of our results. Because the objects with three noncollinear points have special properties in camera calibration, they are midway between 1-D and 2-D calibration objects. Our method is actually a new kind of camera calibration algorithm. Yuncai Liu, Zhengyou Zhang |
IEEE Trans. Image Process. | 3 |
| 2008 | Maximum Likelihood Sound Source Localization and Beamforming for Directional Microphone Arrays in Distributed MeetingsabstractIn distributed meeting applications, microphone arrays have been widely used to capture superior speech sound and perform speaker localization through sound source localization (SSL) and beamforming. This paper presents a unified maximum likelihood framework of these two techniques, and demonstrates how such a framework can be adapted to create efficient SSL and beamforming algorithms for reverberant rooms and unknown directional patterns of microphones. The proposed method is closely related to steered response power-based algorithms, which are known to work extremely well in real-world environments. We demonstrate the effectiveness of the proposed method on challenging synthetic and real-world datasets, including over six hours of recorded meetings. Cha Zhang, Dinei A. F. Florêncio, Demba Ba 0001, Zhengyou Zhang |
IEEE Trans. Multim. | 4 |
| 2008 | Boosting-Based Multimodal Speaker Detection for Distributed Meeting VideosabstractIdentifying the active speaker in a video of a distributed meeting can be very helpful for remote participants to understand the dynamics of the meeting. A straightforward application of such analysis is to stream a high resolution video of the speaker to the remote participants. In this paper, we present the challenges we met while designing a speaker detector for the Microsoft RoundTable distributed meeting device, and propose a novel boosting-based multimodal speaker detection (BMSD) algorithm. Instead of separately performing sound source localization (SSL) and multiperson detection (MPD) and subsequently fusing their individual results, the proposed algorithm fuses audio and visual information at feature level by using boosting to select features from a combined pool of both audio and visual features simultaneously. The result is a very accurate speaker detector with extremely high efficiency. In experiments that includes hundreds of real-world meetings, the proposed BMSD algorithm reduces the error rate of SSL-only approach by 24.6%, and the SSL and MPD fusion approach by 20.9%. To the best of our knowledge, this is the first real-time multimodal speaker detection algorithm that is deployed in commercial products. Cha Zhang, Pei Yin, Yong Rui, Ross Cutler, Paul A. Viola, Xinding Sun, Nelson Pinto, Zhengyou Zhang |
IEEE Trans. Multim. | 8 |
| 2007 | Face Re-Lighting from a Single Image under Harsh Lighting ConditionsabstractIn this paper, we present a new method to change the illumination condition of a face image, with unknown face geometry and albedo information. This problem is particularly difficult when there is only one single image of the subject available and it was taken under a harsh lighting condition. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using spherical harmonic representation. However, the approximation error can be large under harsh lighting conditions thus making it difficult to recover albedo information. In order to address this problem, we propose a subregion based framework that uses a Markov Random Field to model the statistical distribution and spatial coherence of face texture, which makes our approach not only robust to harsh lighting conditions, but insensitive to partial occlusions as well. The performance of our framework is demonstrated through various experimental results, including the improvement to the face recognition rate under harsh lighting conditions. Yang Wang 0001, Zicheng Liu 0001, Gang Hua 0001, Zhengyou Zhang, Dimitris Samaras |
CVPR | 5 |
| 2007 | Energy-Based Sound Source Localization and Gain Normalization for Ad Hoc Microphone ArraysabstractWe present an energy-based technique to estimate both microphone and speaker/talker locations from an ad hoc network of microphones. An example of such ad hoc microphone network is a set of microphones built in the laptops that some meeting participants bring in a meeting room. Compared with traditional sound source localization approaches based on time of flight, our technique does not require accurate synchronization, and it does not require each laptop to emit special signals. We estimate the meeting participants' positions based on average energies of their speech signals. In addition, we present a technique, which is independent of the volumes of the speakers, to estimate the relative gains of the microphones. This is crucial to aggregate various audio channels from the ad hoc microphone network into a single stream for audio conferencing. Zicheng Liu 0001, Zhengyou Zhang, Li-wei He, Philip A. Chou |
ICASSP (2) | 2 |
| 2007 | A Generative-Discriminative Framework using Ensemble Methods for Text-Dependent Speaker VerificationabstractSpeaker verification can be treated as a statistical hypothesis testing problem. The most commonly used approach is the likelihood ratio test (LRT), which can be shown to be optimal using the Neymann-Pearson lemma. However, in most practical situations the Neymann-Pearson lemma does not apply. In this paper, we present a more robust approach that makes use of a hybrid generative-discriminative framework for text-dependent speaker verification. Our algorithm makes use of a generative models to learn the characteristics of a speaker and then discriminative models to discriminate between a speaker and an impostor. One of the advantages of the proposed algorithm is that it does not require us to retrain the generative model. The proposed model, on an average, yields 36.41% relative improvement in EER over a LRT. Amarnag Subramanya, Zhengyou Zhang, Arun C. Surendran, Patrick Nguyen, Mukund Narasimhan, Alex Acero |
ICASSP (4) | 2 |
| 2007 | Maximum Likelihood Sound Source Localization for Multiple Directional MicrophonesabstractThis paper presents a maximum likelihood (ML) framework for multi-microphone sound source localization (SSL). Besides deriving the framework, we focus on making the connection and contrast between the ML-based algorithm and popular steered response power (SRP) SSL algorithms such as phase transform (SRP-PHAT). We also show under our ML framework how challenging conditions such as directional microphone arrays and reverberations can be handled. The computational cost of our method is low-similar to SRP-PHAT. The effectiveness of the proposed method is shown on a large dataset with 99 real-world audio sequences recorded by directional circular microphone arrays in over 50 different meeting rooms. Cha Zhang, Zhengyou Zhang, Dinei A. F. Florêncio |
ICASSP (1) | 2 |
| 2007 | Exploring Discriminative Learning for Text-Independent Speaker RecognitionabstractSpeaker verification is a technology of verifying the claimed identity of a speaker based on the speech signal from the speaker (voice print). To learn the score of similarity between each pair of target and trial utterances, we investigated two different discriminative learning frameworks: Fisher mapping followed by SVM learning and utterance transform followed by iterative cohort modeling (ICM). In both methods, a mapping is applied to map speech utterance from a variable-length acoustic feature sequence into a fixed dimensional vector. SVM learning constructs a classifier in the mapped vector space for speaker verification. ICM learns a metric in this vector space by incorporating discriminative learning methods. The obtained metric is then used by a nearest neighbor classifier for speaker verification. The experiments conducted on NIST02 corpus show that both discriminative learning methods outperform the baseline GMM-UBM system. Furthermore, we observe that the ICM-based method is more effective than the SVM-based method, indicating that the metric learning scheme is more powerful in constructing a better metric in the mapped vector space. Ming Liu 0009, Zhengyou Zhang, Mark Hasegawa-Johnson, Thomas S. Huang |
ICME | 2 |
| 2007 | Learning-Based Perceptual Image Quality Improvement for Video ConferencingabstractIt is well known that in professional TV show filming, stage lighting has to be carefully designed in order to make the host and the scene look visually appealing. The lighting affects not only the brightness but also the color tone which plays a critical role in the perceived look of the host and the mood of the stage. In contrast, during video conferencing, the lighting is usually far from ideal thus the perceived image quality is low. There has been a lot of research on improving the brightness of the captured images. In this paper, we propose a learning-based technique to improve the perceptual image quality by enhancing both brightness and color tone. The basic idea is to learn the color statistics from a training set of images which look visually appealing, and adjust the color of an input image so that its color statistics matches those in the training set. To validate our approach, we have conducted user study and the results show that our technique significantly improves the perceived image quality. Zicheng Liu 0001, Cha Zhang, Zhengyou Zhang |
ICME | 3 |
| 2007 | Frequency domain correspondence for speaker normalizationabstractDue to physiology and linguistic difference between speakers, the spectrum pattern for the same phoneme of two speakers can be quite dissimilar. Without appropriate alignment on the frequency axis, the misalignment will reduce the modeling efficiency resutling in performance degradation. In this paper, a novel data-driven framework is proposed to build the alignment of the frequency axes of two speakers. This alignment between two frequency axes is essentially a frequency domain correspondence of the two speakers. To establish the correspondence, we formulate the task as a global optimal matching problem. The local matching of frequency bins is achieved by comparing the local feature of the spectrogram along the frequency bins. The local feature is actually capturing the local pattern in the spectrogram. Given the local matching score, a dynamic programming is then applied to find the optimal correspondence. Experiments on TIMIT corpus and TIDIGITS corpus clearly show the effectiveness of this method. 1. Ming Liu 0009, Mark Hasegawa-Johnson, Thomas S. Huang, Zhengyou Zhang |
INTERSPEECH | 5 |
| 2007 | Real-Time Whiteboard Capture and Processing Using a Video Camera for Remote CollaborationabstractThis paper describes our recently developed system which captures pen strokes on physical whiteboards in real time using an off-the-shelf video camera. Unlike many existing tools, our system does not instrument the pens or the whiteboard. It analyzes the sequence of captured video images in real time, classifies the pixels into whiteboard background, pen strokes and foreground objects (e.g., people in front of the whiteboard), extracts newly written pen strokes, and corrects the color to make the whiteboard completely white. This allows us to transmit whiteboard contents using very low bandwidth to remote meeting participants. Combined with other teleconferencing tools such as voice conference and application sharing, our system becomes a powerful tool to share ideas during online meetings Li-wei He, Zhengyou Zhang |
IEEE Trans. Multim. | 2 |
| 2007 | Head-Size Equalization for Improved Visual Perception in Video ConferencingabstractIn a video conferencing setting, people often use an elongated meeting table with the major axis along the camera direction. A standard wide-angle perspective image of this setting creates significant foreshortening, thus the people sitting at the far end of the table appear very small relative to those nearer the camera. This has two consequences. First, it is difficult for the remote participants to see the faces of those at the far end, thus affecting the experience of the video conferencing. Second, it is a waste of the screen space and network bandwidth because most of the pixels are used on the background instead of on the faces of the meeting participants. In this paper, we present a novel technique, called Spatially-Varying-Uniform scaling functions, to warp the images to equalize the head sizes of the meeting participants without causing undue distortion. This technique works for both the 180-degree views where the camera is placed at one end of the table and the 360-degree views where the camera is placed at the center of the table. We have implemented this algorithm on two types of camera arrays: one with 180-degree view, and the other with 360-degree view. On both hardware devices, image capturing, stitching, and head-size equalization are run in real time. In addition, we have conducted user study showing that people clearly prefer head-size equalized images. Zicheng Liu 0001, Michael F. Cohen, Deepti Bhatnagar, Ross Cutler, Zhengyou Zhang |
IEEE Trans. Multim. | 5 |
| 2006 | Automatic Business Card Scanning with a CameraabstractIn this paper, we present a system to automatically extract, rectify and enhance business card images. First the business card image patch is automatically segmented by minimizing a novel local-global variational energy. Second a quadrangle is fitted to the segmented image patch. With the four corner points of the quadrangle, we then estimate the physical aspect ratio of the business card and obtain a homography to rectify the quadrangle back to rectangular shape. We finally enhance the contrast of the rectified business card image using a S-shaped curve. Extensive experiments demonstrated the efficacy and robustness of our system. Gang Hua 0001, Zicheng Liu 0001, Zhengyou Zhang, Ying Wu 0001 |
ICIP | 3 |
| 2006 | Automatic Real-Time Barcode Localization in Complex ScenesabstractMany applications exist for automatically finding and reading barcodes in complex scenes with a camera. The key problem is to search barcodes in a complex scene and supply them for a reading subsystem. However, illumination, rotation, perspective distortion and multiple barcodes circumstances make barcode localization difficult. By jointly analyzing texture and shape, we propose a realtime barcode localization method that is robust to the above problems. A user only needs to put a barcode in front of a camera at around 15 cm for common Web cameras, and the system then automatically locates and decodes the barcode. This paper only focuses at automatic localization. We will demonstrate the complete system live at the conference. Shi Han, Mo Yi, Zhengyou Zhang |
ICIP | 5 |
| 2006 | Speech Modelingwith Magnitude-Normalized Complex Spectra and Its Application to Multisensory Speech EnhancementabstractA good speech model is essential for speech enhancement, but it is very difficult to build because of huge intra- and extra-speaker variation. We present a new speech model for speech enhancement, which is based on statistical models of magnitude-normalized complex spectra of speech signals. Most popular speech enhancement techniques work in the spectrum space, but the large variation of speech strength, even from the same speaker, makes accurate speech modeling very difficult because the magnitude is correlated across all frequency bins. By performing magnitude normalization for each speech frame, we are able to get rid of the magnitude variation and to build a much better speech model with only a small number of Gaussian components. This new speech model is applied to speech enhancement for our previously developed microphone headsets that combine a conventional air microphone with a bone sensor. Much improved results have been obtained Amarnag Subramanya, Zhengyou Zhang, Zicheng Liu 0001, Alex Acero |
ICME | 2 |
| 2006 | A novel framework of text-independent speaker verification based on utterance transform and iterative cohort modelingabstractA novel framework for text-independent speaker verification is proposed. The framework is based on a new interpretation of Universal Background Model. The UBM in our framework actually defines a transform which maps the variable length observation into a fixed dimensional supervector(supervector space). Each speech utterance is then mapped into a point in this supervector space. The similarity measure in this vector space is progressively refined via an iterative cohort modeling scheme. The experiments on NIST 2002 corpus show the effectiveness of this new framework. Overall the EER drops from the baseline system(with T-Norm) 9.21 % to final improved system(without T-Norm) 8.07%. The new framework can effectively reduce the data dependence in the final output score which is clearly indicated in the second sets of experiments. The EER after T-Norm of final system marginally increases by relatively 1.73 % compared to the EER of baseline system drops 16.12 % relatively after T-Norm. Also, the relative improvement of DCF after T-Norm is marginal for the final improved system (2.47%) compared to 33.68 % in baseline system. It clear shows that the iterative cohort modeling effectively reduce the data dependence of the final scores, so that T-Norm will not further improve the system performance. Also, the performance of novel frame clearly increases as the iteration grows which suggest that the framework progressively refine the similarity measure on the supervector space with the iterative cohort modeling. Index Terms: speaker verification, utterance transform, iterative cohort modeling. Ming Liu 0009, Huazhong Ning, Thomas S. Huang, Zhengyou Zhang |
INTERSPEECH | 4 |
| 2006 | Editorial
Zhengyou Zhang |
Int. J. Comput. Vis. | 1 |
| 2006 | Iterative Local-Global Energy Minimization for Automatic Extraction of Objects of InterestabstractWe propose a novel global-local variational energy to automatically extract objects of interest from images. Previous formulations only incorporate local region potentials, which are sensitive to incorrectly classified pixels during iteration. We introduce a global likelihood potential to achieve better estimation of the foreground and background models and, thus, better extraction results. Extensive experiments demonstrate its efficacy. Gang Hua 0001, Zicheng Liu 0001, Zhengyou Zhang, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2005 | Linear Combination Representation for Outlier Detection in Motion TrackingabstractIn this paper we show that Ullman and Basri's linear combination (LC) representation, which was originally proposed for alignment-based object recognition, can be used for outlier detection in motion tracking with an affine camera. For this task LC can be realized either on image frames or feature trajectories, and therefore two methods are developed which we call linear combination of frames and linear combination of trajectories. For robust estimation of the linear combination coefficients, the support vector regression (SVR) algorithm is used and compared with the RANSAC method. SVR based on quadratic programming optimization can efficiently deal with more than 50 percent outliers and delivers more consistent results than RANSAC in our experiments. The linear combination representation can use SVR in a straightforward manner while previous factorization-based or subspace separation methods cannot. Experimental results are presented using real video sequences to demonstrate the effectiveness of our LC+SVR approaches, including a quantitative comparison of SVR and RANSAC. Guodong Guo, Charles R. Dyer, Zhengyou Zhang |
CVPR (2) | 3 |
| 2005 | Real-time whiteboard capture and processing using a video camera for teleconferencingabstractThis paper describes our recently developed system which captures pen strokes on whiteboards in real time using an off-the-shelf video camera. Unlike many existing tools, our system does not instrument the pens or the whiteboard. It analyzes the sequence of captured video images in real time, classifies the pixels into whiteboard background, pen strokes and foreground objects (e.g., people in front of the whiteboard), and extracts newly written pen strokes. This allows us to transmit whiteboard contents using very low bandwidth to remote meeting participants. Combined with other teleconferencing tools such as voice conference and application sharing, our system becomes a powerful tool to share ideas during online meetings. Li-wei He, Zhengyou Zhang |
ICASSP (2) | 2 |
| 2005 | Leakage Model and Teeth Clack Removal for Air- and Bone-Conductive Integrated MicrophonesabstractContinuing our previous work (Zhang et al. (2004), Liu et al. (2004)) on using air- and bone-conductive integrated microphones, and in particular on using the direct filtering approach (Liu et al. (2004)) for speech enhancement in noisy environments, we present in this paper a refined version of the direct filtering algorithm. The new algorithm takes into account explicitly the leakage of background noise into the bone channel. We also present a new algorithm that detects and removes an artifact known as teeth clacks. Experiments show that the addition of the above algorithms improves system performance to a large extent even in highly nonstationary noisy environments. Zicheng Liu 0001, Amarnag Subramanya, Zhengyou Zhang, Jasha Droppo, Alex Acero |
ICASSP (1) | 3 |
| 2005 | Automatic Head-size Equalization in Panorama Images for Video ConferencingabstractIn panorama images captured by omni-directional cameras during video conferencing, the image sizes of the people around the conference table are not uniform due to the varying distances to the camera. Spatially varying-uniform (SVU) scaling functions have been proposed to warp a panorama image smoothly such that the participants have similar sizes on the image. To generate the SVU function, one needs to segment the table boundaries, which was generated manually in the previous work. In this paper, we propose a robust algorithm to automatically segment the table boundaries. To ensure the robustness, we apply a symmetry voting scheme to filter out noisy points on the edge map. Trigonometry and quadratic fitting methods are developed to fit a continuous curve to the remaining edge points. We report experimental results on both synthetic and real images. Ya Chang, Ross Cutler, Zicheng Liu 0001, Zhengyou Zhang, Alex Acero, Matthew Turk 0001 |
ICME | 4 |
| 2005 | Multi-sensory speech processing: incorporating automatically extracted hidden dynamic informationabstractWe describe a novel technique for multi-sensory speech processing for enhancing noisy speech and for improved noise-robust speech recognition. Both air- and bone-conductive microphones are used to capture speech data where the bone sensor contains virtually noise-free hidden dynamic information of clean speech in the form of formant trajectories. The distortion in the bone-sensor signal such as teeth-clacking and noise leakage can be effectively removed by making use of the automatically extracted formant information from the bone-sensor signal. This paper reports an improved technique for synthesizing speech waveforms based on the LPC cepstra computed analytically from the formant trajectories. When this new signal stream is fused with the other available speech data streams, we achieved improved performance for noisy speech recognition. Amarnag Subramanya, Li Deng 0001, Zicheng Liu 0001, Zhengyou Zhang |
ICME | 4 |
| 2005 | A graphical model for multi-sensory speech processing in air-and-bone conductive microphonesabstractIn continuation of our previous work on using an air-and-boneconductive microphone for speech enhancement, in this paper we propose a graphical model based approach to estimating the clean speech signal given the noisy observations in the air sensor. We also show how the same model can be used as a speech/non-speech classifier. With the aid of MOS (mean opinion score) tests we show, that the performance of the proposed model is better in comparison to our previously proposed direct filtering algorithm. Amarnag Subramanya, Zhengyou Zhang, Zicheng Liu 0001, Jasha Droppo, Alex Acero |
INTERSPEECH | 2 |
| 2005 | Editorial
Zhengyou Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2004 | Office presence detection using multimodal context informationabstractAn office presence detection system is presented. Context information from multi-sensory inputs is integrated to infer a user's activities in an office. We design a layered architecture to model human activities with different granularities. An IHDR (incremental hierarchical discriminant regression) tree is used to generate models automatically for acoustic signals from unsegmented auditory streams, with a high adaptive capability to new settings. Hidden Markov models (HMM) are implemented to detect human motion patterns. The outputs of the above two components are fed into high-level HMMs to analyze human activities. Experimental results of the real-time prototype system are reported. Juyang Weng, Zhengyou Zhang |
ICASSP (3) | 3 |
| 2004 | Note-taking with a camera: whiteboard scanning and image enhancementabstractThe paper describes a system for scanning the content on a whiteboard into a computer by a digital camera and also for enhancing the visual quality of whiteboard images. Because digital cameras are becoming accessible to average users, more and more people use digital cameras to take pictures of whiteboards instead of copying manually, thus increasing productivity significantly. However, the images are usually taken from an angle, resulting in undesired perspective distortion. They also contain other distracting regions such as walls. We have developed a system that automatically locates the boundary of a whiteboard, crops out the whiteboard region, rectifies it into a rectangle, and corrects the color to make the whiteboard completely white. In the case where a single image is not enough (e.g., large whiteboard and low-res camera), we have developed a robust feature-based technique to stitch multiple overlapping images automatically. We therefore reproduce the whiteboard content as a faithful electronic document which can be archived or shared with others. The system has been tested extensively, and very good results have been obtained. Zhengyou Zhang, Li-wei He |
ICASSP (3) | 1 |
| 2004 | Multi-sensory microphones for robust speech detection, enhancement and recognitionabstractIn this paper, we present new hardware prototypes that integrate several heterogeneous sensors into a single headset and describe the underlying DSP techniques for robust speech detection, enhancement and recognition in highly non-stationary noisy environments. We also speculate other business uses with this type of device. Zhengyou Zhang, Zicheng Liu 0001, Mike Sinclair, Alex Acero, Li Deng 0001, Jasha Droppo, Xuedong Huang 0001, Yanli Zheng |
ICASSP (3) | 1 |
| 2004 | Visual echo cancellation in a projector-camera-whiteboard systemabstractWe propose to incorporate a whiteboard into a projector-camera system. The whiteboard serves as the writing surface (input) as well as the projecting surface (output). The ability to write and draw on top of computer-projected content opens up many new opportunities for real-time collaborations between people located on-site and remotely. Such applications inevitably require extracting handwritings from video images that contain both handwritings and the projected content. By analogy with echo cancellation in audio conferencing, we call this problem visual echo cancellation. This paper presents one approach to accomplish the task. Our visual echo cancellation algorithm estimates the incident light and derives the surface albedo based on both incident light and reflection. By estimating the albedo, we can extract the writings and recover their colors. Our approach includes two basic components of projector-camera systems: geometric calibration and color calibration. The first one solves the mapping between the position in the camera view and the position in the projector screen, while the second one solves the mapping between the actual color of the projected content and that seen by the camera. Hanning Zhou, Zhengyou Zhang, Thomas S. Huang |
ICIP | 2 |
| 2004 | Nonlinear information fusion in multi-sensor processing - extracting and exploiting hidden dynamics of speech captured by a bone-conductive microphoneabstractOne well-known difficulty in creating an effective human-machine interface via the speech input is the adverse effects of concurrent acoustic noise. To overcome this challenge, we have developed a joint hardware and software solution. A novel bone-conductive microphone is integrated with a regular air-conductive one in a single headset. These two simultaneous sensors capture the distinct signal properties in the speech embedded in acoustic noise. The focus of this paper is the exploration of the type of dynamic properties that are relatively invariant between the bone-conductive sensor's signal and the clean speech signal; the latter would not be available to the recognizer. Our approach is based on a nonlinear processing technique that estimates the unobserved (hidden) vocal tract resonances, as a representation of such invariant hidden dynamics, from the available bone-sensor signal. The information about these dynamic aspects of the clean speech is then fused with the other noisy measurements that aims to improve the recognition system's robustness to acoustic distortion. The fusion technique is based on a combination of three sets of signals including the synthesized speech signal using the vocal tract resonance dynamics extracted nonlinearly from the bone-sensor signal. Li Deng 0001, Zicheng Liu 0001, Zhengyou Zhang, Alex Acero |
MMSP | 3 |
| 2004 | Direct filtering for air- and bone-conductive microphonesabstractAir- and bone-conductive integrated microphones have been introduced by the authors [Y. Zheng, et al., 2003, Z. Zhang et al., 2004] for speech enhancement in noisy environments. In this paper, we present a novel technique, called direct filtering, to combine the two channels from the air- and bone-conductive microphone for speech enhancement. Compared to the previous technique, the advantage of the direct filtering is that it does not require any training, and it is speaker independent. Experiments show that this technique effectively removes noises and significantly improves speech recognition accuracy even in highly non-stationary noisy environments. Zicheng Liu 0001, Zhengyou Zhang, Alex Acero, Jasha Droppo, Xuedong Huang 0001 |
MMSP | 2 |
| 2004 | Robust and Rapid Generation of Animated Faces from Video Images: A Model-Based Modeling Approach
Zhengyou Zhang, Zicheng Liu 0001, Dennis Adler, Michael F. Cohen, Erik Hanson, Ying Shan |
Int. J. Comput. Vis. | 1 |
| 2004 | Automatic Eyeglasses Removal from Face ImagesabstractIn this paper, we present an intelligent image editing and face synthesis system that automatically removes eyeglasses from an input frontal face image. Although conventional image editing tools can be used to remove eyeglasses by pixel-level editing, filling in the deleted eyeglasses region with the right content is a difficult problem. Our approach works at the object level where the eyeglasses are automatically located, removed as one piece, and the void region filled. Our system consists of three parts: eyeglasses detection, eyeglasses localization, and eyeglasses removal. First, an eye region detector, trained offline, is used to approximately locate the region of eyes, thus the region of eyeglasses. A Markov-chain Monte Carlo method is then used to accurately locate key points on the eyeglasses frame by searching for the global optimum of the posterior. Subsequently, a novel sample-based approach is used to synthesize the face image without the eyeglasses. Specifically, we adopt a statistical analysis and synthesis approach to learn the mapping between pairs of face images with and without eyeglasses from a database. Extensive experiments demonstrate that our system effectively removes eyeglasses. Ce Liu 0001, Harry Shum, Ying-Qing Xu, Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2004 | Eye Gaze Correction with Stereovision for Video-TeleconferencingabstractThe lack of eye contact in desktop video teleconferencing substantially reduces the effectiveness of video contents. While expensive and bulky hardware is available on the market to correct eye gaze, researchers have been trying to provide a practical software-based solution to bring video-teleconferencing one step closer to the mass market. This paper presents a novel approach: Based on stereo analysis combined with rich domain knowledge (a personalized face model), we synthesize, using graphics hardware, a virtual video that maintains eye contact. A 3D stereo head tracker with a personalized face model is used to compute initial correspondences across two views. More correspondences are then added through template and feature matching. Finally, all the correspondence information is fused together for view synthesis using view morphing techniques. The combined methods greatly enhance the accuracy and robustness of the synthesized views. Our current system is able to generate an eye-gaze corrected video stream at five frames per second on a commodity 1 GHz PC. Ruigang Yang, Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Camera Calibration with One-Dimensional ObjectsabstractCamera calibration has been studied extensively in computer vision and photogrammetry and the proposed techniques in the literature include those using 3D apparatus (two or three planes orthogonal to each other or a plane undergoing a pure translation, etc.), 2D objects (planar patterns undergoing unknown motions), and 0D features (self-calibration using unknown scene points). Yet, this paper proposes a new calibration technique using 1D objects (points aligned on a line), thus filling the missing dimension in calibration. In particular, we show that camera calibration is not possible with free-moving 1D objects, but can be solved if one point is fixed. A closed-form solution is developed if six or more observations of such a 1D object are made. For higher accuracy, a nonlinear technique based on the maximum likelihood criterion is then used to refine the estimate. Singularities have also been studied. Besides the theoretical aspect, the proposed technique is also important in practice especially when calibrating multiple cameras mounted apart from each other, where the calibration objects are required to be visible simultaneously. Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Why take notes? Use the whiteboard capture systemabstractThe paper describes our recently developed system which captures both whiteboard content and audio signals of a meeting using a digital still camera and a microphone. Our system can be retrofitted to any existing whiteboard. It computes the time stamps of pen strokes on the whiteboard by analyzing the sequence of captured snapshots. It also automatically produces a set of key frames representing all the written content on the whiteboard before each erasure. Therefore, the whiteboard content serves as a visual index to browse the audio meeting efficiently. It is a complete system which not only captures the whiteboard content, but also helps users to view and manage the captured meeting content efficiently and securely. Li-wei He, Zicheng Liu 0001, Zhengyou Zhang |
ICASSP (5) | 3 |
| 2003 | Incremental motion estimation through modified bundle adjustmentabstractAn incremental motion estimation scheme for long image sequence analysis is introduced. It applies to a sliding window of triplet images and maintains local motion consistency without resort to post-concatenation. This is possible thanks to our newly developed local process called three-view partial bundle adjustment. Unlike previous approaches that rely only on point matches across three or more views, our local process also takes into account all available two-view matches, leading to more accurate motion estimation. For sparse image sequences, two-view matches are very reliable and it becomes even more important to use them since the number of matches across more views decreases quickly. In this case, our incremental method can produce results very close to those obtained with a global bundle adjustment but in a fraction of time. Experiments with both synthetic and real data have been conducted to compare the proposed technique with other techniques, and have shown our technique to be clearly superior. Zhengyou Zhang, Ying Shan |
ICIP (2) | 1 |
| 2002 | Eye Gaze Correction with Stereovision for Video-Teleconferencing
Ruigang Yang, Zhengyou Zhang |
ECCV (2) | 2 |
| 2002 | Camera Calibration with One-Dimensional Objects
Zhengyou Zhang |
ECCV (4) | 1 |
| 2002 | Distributed meetings: a meeting capture and broadcasting systemabstractThe common meeting is an integral part of everyday life for most workgroups. However, due to travel, time, or other constraints, people are often not able to attend all the meetings they need to. Teleconferencing and recording of meetings can address this problem. In this paper we describe a system that provides these features, as well as a user study evaluation of the system. The system uses a variety of capture devices (a novel 360° camera, a whiteboard camera, an overview camera, and a microphone array) to provide a rich experience for people who want to participate in a meeting from a distance. The system is also combined with speaker clustering, spatial indexing, and time compression to provide a rich experience for people who miss a meeting and want to watch it afterward. Ross Cutler, Yong Rui, Anoop Gupta, Jonathan J. Cadiz, Ivan Tashev, Li-wei He, Alex Colburn, Zhengyou Zhang, Zicheng Liu 0001, Steve Silverberg |
ACM Multimedia | 8 |
| 2002 | New Measurements and Corner-Guidance for Curve Matching with Probabilistic Relaxation
Ying Shan, Zhengyou Zhang |
Int. J. Comput. Vis. | 2 |
| 2001 | Image-Based Surface Detail TransferabstractWe present a novel technique, called Image-Based Surface Detail Transfer, to transfer geometric details from one surface to another with simple 2D image operations. The basic observation is that, without knowing its 3D geometry, geometric details (local deformations) can be extracted from a single image of an object in a way independent of its surface reflectance, and furthermore, these geometric details can be transferred to modify the appearance of other objects directly in images. We show examples including surface detail transfer between real objects, as well as between real and synthesized objects. Ying Shan, Zicheng Liu 0001, Zhengyou Zhang |
CVPR (2) | 3 |
| 2001 | Multibody Grouping via Orthogonal Subspace DecompositionabstractMultibody structure from motion could be solved by the factorization approach. However, the noise measurements would make the segmentation difficult when analyzing the shape interaction matrix. This paper presents an orthogonal subspace decomposition and grouping technique to approach such a problem. We decompose the object shape spaces into signal subspaces and noise subspaces. We show that the signal subspaces of the object shape spaces are orthogonal to each other. Instead of using the shape interaction matrix contaminated by noise, we introduce the shape signal subspace distance matrix for shape space grouping. Outliers could be easily identified by this approach. The robustness of the proposed approach lies in the fact that the shape space decomposition alleviates the influence of noise, and has been verified with extensive experiments. Ying Wu 0001, Zhengyou Zhang, Thomas S. Huang, John Y. Lin |
CVPR (2) | 2 |
| 2001 | Determining Reflectance Parameters and Illumination Distribution from a Sparse Set of Images for View-dependent Image Synthesis
Ko Nishino, Zhengyou Zhang, Katsushi Ikeuchi |
ICCV | 2 |
| 2001 | Model-Based Bundle Adjustment with Application to Face ModelingabstractWe present a new model-based bundle adjustment algorithm to recover the 3D model of a scene/object from a sequence of images with unknown motions. Instead of representing scene/object by a collection of isolated 3D features (usually points), our algorithm uses a surface controlled by a small set of parameters. Compared with previous model-based approaches, our approach has the following advantages. First instead of using the model space as a regularizer we directly use it as our search space, thus resulting in a more elegant formulation with fewer unknowns and fewer equations. Second, our algorithm automatically associates tracked points with their correct locations on the surfaces, thereby eliminating the need for a prior 2D-to-3D association. Third, regarding face modeling, we use a very small set of face metrics (meaningful deformations) to parameterize the face geometry, resulting in a smaller search space and a better posed system. Experiments with both synthetic and real data show that this new algorithm is faster, more accurate and more stable than existing ones. Ying Shan, Zicheng Liu 0001, Zhengyou Zhang |
ICCV | 3 |
| 2001 | Cloning Your Own Face with a Desktop CameraabstractWe have developed an easy and cost-effective system that constructs textured 3D animated face models from videos with minimal user interaction. Our system first takes, with an ordinary video camera, images of a face of a person sitting in front of the camera turning the head from one side to the other. After five manual clicks on two images to tell the system where the eye corners, nose top and mouth corners are, the system automatically generates a realistic looking 3D human head model and the constructed model can be animated immediately (different poses, facial expressions and talking). A user, with a PC and a video camera, can use our system to generate hisher face model in a few minutes. The face model can then be imported in hisher favorite game, and the user sees themselves and their friends take part in the game they are playing. We will demonstrate the system on a laptop computer live at the conference, and participants can try it to model their own faces. Zhengyou Zhang, Zicheng Liu 0001, Dennis Adler, Michael F. Cohen, Erik Hanson, Ying Shan |
ICCV | 1 |
| 2001 | Expressive expression mapping with ratio imagesabstractFacial expressions exhibit not only facial feature motions, but also subtle changes in illumination and appearance (e.g., facial creases and wrinkles). These details are important visual cues, but they are difficult to synthesize. Traditional expression mapping techniques consider feature motions while the details in illumination changes are ignored. In this paper, we present a novel technique for facial expression mapping. We capture the illumination change of one person's expression in what we call an expression ratio image (ERI). Together with geometric warping, we map an ERI to any other person's face image to generate more expressive facial expressions. Zicheng Liu 0001, Ying Shan, Zhengyou Zhang |
SIGGRAPH | 3 |
| 2001 | Estimating the Fundamental Matrix by Transforming Image Points in Projective Space
Zhengyou Zhang, Charles T. Loop |
Comput. Vis. Image Underst. | 1 |
| 2001 | Rapid modeling of animated faces from videoabstractAbstract Generating realistic 3D human face models and facial animations has been a persistent challenge in computer graphics. We have developed a system that constructs textured 3Dface models from videos with minimal user interaction. Our system takes images andvideo sequences of a face with an ordinary video camera. After five manual clicks ontwo images to tell the system where the eye corners, nose top and mouth corners are, thesystem automatically generates a realistic looking 3D human head model and the constructed model can be animated immediately. A user with a PC and an ordinary camera can use our system to generate his/her face model in a few minutes. Copyright © 2001 John Wiley & Sons, Ltd. Zicheng Liu 0001, Zhengyou Zhang, Chuck Jacobs, Michael F. Cohen |
Comput. Animat. Virtual Worlds | 2 |
| 2000 | Corner Guided Curve Matching and its Application to Scene ReconstructionabstractCorners and curves are important image features in many vision-based applications. Corners are usually more stable and easier to match than curves, while curves contain richer information of scene structure. In previously work, corners are often used to recover the epipolar geometry between two views, which is then used in curve matching to reduce the search space. However, information of the scene structure contained in this set of matched corners is ignored. In this paper we present a curve matching algorithm that is guided by a set of matched corners. Within a probabilistic framework, the role of the corner guidance is explicitly defined by a set of similarity-invariant unary measurements and by a similarity function. The similarity function provides stronger capability of resolving matching ambiguity than the epipolar constraint, and is integrated into a relaxation scheme to reduce computational complexity and improve accuracy of curve matching. Experimental results clearly demonstrate the benefit of integrating corner matches into the curve matching procedure. Ying Shan, Zhengyou Zhang |
CVPR | 2 |
| 2000 | Rapid modeling of animated faces from video images
Zicheng Liu 0001, Zhengyou Zhang, Chuck Jacobs, Michael F. Cohen |
ACM Multimedia | 2 |
| 2000 | A Flexible New Technique for Camera CalibrationabstractWe propose a flexible technique to easily calibrate a camera. It only requires the camera to observe a planar pattern shown at a few (at least two) different orientations. Either the camera or the planar pattern can be freely moved. The motion need not be known. Radial lens distortion is modeled. The proposed procedure consists of a closed-form solution, followed by a nonlinear refinement based on the maximum likelihood criterion. Both computer simulation and real data have been used to test the proposed technique and very good results have been obtained. Compared with classical techniques which use expensive equipment such as two or three orthogonal planes, the proposed technique is easy to use and flexible. It advances 3D computer vision one more step from laboratory environments to real world use. Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Computing Rectifying Homographies for Stereo VisionabstractImage rectification is the process of applying a pair of 2D projective transforms, or homographies, to a pair of images whose epipolar geometry is known so that epipolar lines in the original images map to horizontally aligned lines in the transformed images. We propose a novel technique for image rectification based on geometrically well defined criteria such that image distortion due to rectification is minimized. This is achieved by decomposing each homography into a specialized projective transform, a similarity transform, followed by a shearing transform. The effect of image distortion at each stage is carefully considered. Charles T. Loop, Zhengyou Zhang |
CVPR | 2 |
| 1999 | Efficient Bundle Adjustment with Virtual Key Frames: A Hierarchical Approach to Multi-Frame Structure from MotionabstractIn this paper we present an efficient hierarchical approach to structure from motion for long image sequences. There are two key elements to our approach: accurate 3D reconstruction for each segment and efficient bundle adjustment for the whole sequence. The image sequence is first divided into a number of segments so that feature points can be reliably tracked across each segment. Each segment has a long baseline to ensure accurate 3D reconstruction. To efficiently bundle adjust 3D structures from ail segments, we reduce the number of frames in each segment by introducing "virtual keyframes". The virtual frames encode the 3D structure of each segment along with its uncertainty but they form a small subset of the original frames. Our method achieves significant speedup over conventional bundle adjustment methods. Harry Shum, Zhengyou Zhang, Qifa Ke |
CVPR | 2 |
| 1999 | Flexible Camera Calibration by Viewing a Plane from Unknown OrientationsabstractProposes a flexible new technique to easily calibrate a camera. It only requires the camera to observe a planar pattern shown at a few (at least two) different orientations. Either the camera or the planar pattern can be freely moved. The motion need not be known. Radial lens distortion is modeled. The proposed procedure consists of a closed-form solution followed by a nonlinear refinement based on the maximum likelihood criterion. Both computer simulation and real data have been used to test the proposed technique, and very good results have been obtained. Compared with classical techniques which use expensive equipment, such as two or three orthogonal planes, the proposed technique is easy to use and flexible. It advances 3D computer vision one step from laboratory environments to real-world use. The corresponding software is available from the author's Web page (). Zhengyou Zhang |
ICCV | 1 |
| 1999 | What Can Be Determined from a Full and a Weak Perspective Image?abstractThis paper presents a first investigation on the structure from motion problem from the combination of full and weak perspective images. This problem arises in multiresolution object modeling where multiple zoomed-in or close-up views are combined with wider or distant reference views. The narrow field-of-view (FOV) images from the zoomed-in or closeup views can be approximated as weak perspective projection. Using a full perspective projection model for the narrow FOV images, although more accurate, actually leads to instabilities during the estimation process due to the non-linearities in the imaging model. The weak perspective approximation leads to more stable estimation algorithms, although at the cost of a small amount of modeling inaccuracy. Previous work in structure from motion focused either on two (or more) perspective images or on a set of weak perspective (more generally, affine) images. The main contribution of this paper is the study of the SFM problem for the much neglected case of one perspective and one (or more) weak perspective image. We show that in contrast to the case of a pair of weak perspective images, there is adequate information to recover Euclidean structure from a single perspective and a single weak perspective image. The epipolar geometry is simpler than with two perspective images leading to simpler and more stable estimation algorithms. Computer simulation shows that more stable results can be obtained with the technique presented in this paper than if two images are both considered to be full perspective. Zhengyou Zhang, P. Anandan 0001, Harry Shum |
ICCV | 1 |
| 1999 | Feature-Based Facial Expression Recognition: Sensitivity Analysis and Experiments with A Multilayer PerceptronabstractIn this paper, we report our experiments on feature-based facial expression recognition within an architecture based on a two-layer perceptron. We investigate the use of two types of features extracted from face images: the geometric positions of a set of fiducial points on a face, and a set of multiscale and multiorientation Gabor wavelet coefficients at these points. They can be used either independently or jointly. The recognition performance with different types of features has been compared, which shows that Gabor wavelet coefficients are much more powerful than geometric positions. Furthermore, since the first layer of the perceptron actually performs a nonlinear reduction of the dimensionality of the feature space, we have also studied the desired number of hidden units, i.e. the appropriate dimension to represent a facial expression in order to achieve a good recognition rate. It turns out that five to seven hidden units are probably enough to represent the space of facial expressions. Then, we have investigated the importance of each individual fiducial point to facial expression recognition. Sensitivity analysis reveals that points on cheeks and on forehead carry little useful information. After discarding them, not only the computational efficiency increases, but also the generalization performance slightly improves. Finally, we have studied the significance of image scales. Experiments show that facial expression recognition is mainly a low frequency process, and a spatial resolution of 64 pixels × 64 pixels is probably enough. Zhengyou Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1998 | Image-Based Geometrically-Correct Photorealistic Scene/Object Modeling (IBPhM): A Review
Zhengyou Zhang |
ACCV (2) | 1 |
| 1998 | A New Multistage Approach to Motion and Structure Estimation by Gradually Enforcing Geometric Constraints
Zhengyou Zhang |
ACCV (2) | 1 |
| 1998 | Head Pose Determination from One Image Using a Generic Model
Zhengyou Zhang, Shigeru Akamatsu, Koichiro Deguchi |
FG | 2 |
| 1998 | Comparison Between Geometry-Based and Gabor-Wavelets-Based Facial Expression Recognition Using Multi-Layer Perceptron
Zhengyou Zhang, Michael J. Lyons 0001, Michael Schuster, Shigeru Akamatsu |
FG | 1 |
| 1998 | Understanding the Relationship Between the Optimization Criteria in Two-View Motion AnalysisabstractThe three best known criteria in two-view motion analysis are based respectively, on the distances between points and their corresponding epipolar lines, on the gradient-weighted epipolar errors, and on the distances between points and the reprojections of their reconstructed paints. The last one has a better statistical interpretation, but is, however much slower than the first two. In this paper we show that the last two criteria are equivalent when the epipoles are at infinity, and differ from each other only a little even when the epipoles are in the image. The first two criteria are equivalent only when the epipoles are at infinity and when the observed object has the same scale in the two images. This suggests that the second criterion is sufficient in practice because of its computational efficiency. The result is valid for both calibrated and uncalibrated images. Zhengyou Zhang |
ICCV | 1 |
| 1998 | Modeling Geometric Structure and Illumination Variation of a Scene from Real ImagesabstractWe present in this paper a system which automatically builds, from real images, a scene model containing both 3D geometric information of the scene structure and its photometric information under various illumination conditions. The geometric structure is recovered from images taken from distinct viewpoints. Structure-from-motion and correlation-based stereo techniques are used to match pixels between images of different viewpoints and to reconstruct the scene in 3D space. The photometric property is extracted from images taken under different illumination conditions (orientation, position and intensity of the light sources). This is achieved by computing a low-dimensional linear space of the spatio-illumination volume, and is represented by a set of basis images. The model that has been built can be used to create realistic renderings from different viewpoints and illumination conditions. Applications include object recognition, virtual reality and product advertisement. Zhengyou Zhang |
ICCV | 1 |
| 1998 | Euclidean Structure from Uncalibrated Images Using Fuzzy Domain Knowledge: Application to Facial Images SynthesisabstractUse of uncalibrated images has found many applications such as image synthesis. However, it is not easy to specify the desired position of the new image in projective or affine space. This paper proposes to recover Euclidean structure from uncalibrated images using domain knowledge such as distances and angles. The knowledge we have is usually about an object category, but not very precise for the particular object being considered. The variation (fuzziness) is modeled as a Gaussian variable. Six types of common knowledge are formulated. Once we have an Euclidean description, the task to specify the desired position in Euclidean space becomes trivial. The proposed technique is then applied to synthesis of new facial images. A number of difficulties existing in image synthesis are identified and solved. For example, we propose to use edge points to deal with occlusion. Zhengyou Zhang, Katsunori Isono, Shigeru Akamatsu |
ICCV | 1 |
| 1998 | Determining the Epipolar Geometry and its Uncertainty: A Review
Zhengyou Zhang |
Int. J. Comput. Vis. | 1 |
| 1998 | On the Optimization Criteria Used in Two-View Motion AnalysisabstractThe three best-known criteria in two-view motion analysis are based, respectively, on the distances between points and their corresponding epipolar lines, on the gradient-weighted epipolar errors, and on the distances between points and the re-projections of their reconstructed points. The last one has a better statistical interpretation, but is significantly slower than the first two. The author shows that, given a reasonable initial guess of the epipolar geometry, the last two criteria are equivalent when the epipoles are at infinity, and differ from each other only a little even when the epipoles are in the image, as shown experimentally. The first two criteria are equivalent only when the epipoles are at infinity and when the observed object/scene has the same scale in the two images. This suggests that the second criterion is sufficient in practice because of its computational efficiency. Experiments with several thousand computer simulations and four sets of real data confirm the analysis. Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Self-Maintaining Camera Calibration Over TimeabstractThe success of an intelligent robotic system depends on the performance of its vision-system which in turn depends to a great extend upon the quality of its calibration. During the execution of a task the vision-system is subject to external influences such as vibrations, thermal expansion etc. which affect and possibly render invalid the initial calibration. Moreover it is possible that the parameters of the vision-system like, for example, the zoom or the focus are altered intentionally in order to perform specific vision-tasks. This paper describes a technique for automatically maintaining calibration of stereovision systems over time without using again any particular calibration apparatus. It uses all available information, i.e. both spatial and temporal data. Uncertainty is systematically manipulated and maintained. Synthetical and real data are used to validate the proposed technique, and the results compare very favourably with those given by classical calibration methods. Zhengyou Zhang, Veit Schenk |
CVPR | 1 |
| 1997 | A General Expression of the Fundamental Matrix for Both Perspective and Affine Cameras
Zhengyou Zhang |
IJCAI | 1 |
| 1997 | Characterizing the Uncertainty of the Fundamental Matrix
Gabriela Csurka, Cyril Zeller, Zhengyou Zhang, Olivier D. Faugeras |
Comput. Vis. Image Underst. | 3 |
| 1997 | A Tighter Lower Bound on the Spetsakis-Aloimonos Trilinear Constraints
Zhengyou Zhang |
Comput. Vis. Image Underst. | 1 |
| 1997 | Parameter estimation techniques: a tutorial with application to conic fitting
Zhengyou Zhang |
Image Vis. Comput. | 1 |
| 1997 | A stereovision system for a planetary rover: calibration, correlation, registration, and fusion
Zhengyou Zhang |
Mach. Vis. Appl. | 1 |
| 1996 | On the epipolar geometry between two images with lens distortionabstractIn order to achieve a 3D, either Euclidean or projective, reconstruction with high precision, one has to consider lens distortion. In almost all work on multiple-views problems in computer vision, a camera is modeled as a pinhole. Lens distortion has usually been corrected off-line. This paper intends to consider lens distortion as an integral part of a camera. We first describe the epipolar geometry between two images with lens distortion. For a point in one image, its corresponding point in the other image should lie on a so-called epipolar curve. We then investigate the possibility of estimating the distortion parameters and the fundamental matrix based on the generalized epipolar constraint. Experimental results with computer simulation show that the distortion parameters can be estimated correctly if the noise in image points is low and the lens distortion is severe. Otherwise, it is better to treat the cameras as being distortion-free. Zhengyou Zhang |
ICPR | 1 |
| 1996 | Motion of an uncalibrated stereo rig: self-calibration and metric reconstructionabstractWe address in this paper the problem of self-calibration and metric reconstruction (up to a scale factor) from one unknown motion of an uncalibrated stereo rig. The epipolar constraint is first formulated for two uncalibrated images. The problem then becomes one of estimating unknowns such that the discrepancy from the epipolar constraint, in terms of sum of squared distances between points and their corresponding epipolar lines, is minimized. Although the full self-calibration is theoretically possible, we assume in this paper that the coordinates of the principal point of each camera are known. Then, the initialization of the unknowns can be done based on our previous work on self-calibration of a single moving camera, which requires to solve a set of so-called Kruppa equations. Redundancy of the information contained in a sequence of stereo images makes this method more robust than using a sequence of monocular images. Real data has been used to test the proposed method, and the results obtained are quite good. We also show experimentally that it is very difficult to estimate precisely the coordinates of the principal points of cameras. A variation of as high as several dozen pixels in the principal point coordinates does not affect significantly the 3-D reconstruction. Zhengyou Zhang, Quang-Tuan Luong, Olivier D. Faugeras |
IEEE Trans. Robotics Autom. | 1 |
| 1995 | Motion of a Stereo Rig: Strong Weak and Self Calibration
Zhengyou Zhang |
ACCV | 1 |
| 1995 | Multi-Sensor Multi-Target Tracking - Strategies for Events that become InvisibleabstractWithin the context of multi-sensor, multi-target tracking, this paper addresses the problems of dealing with objects in the world that, for various reasons, were being observed but are no longer. A modification to the usual Mahalanobis distance metric used for matching observations to Extended Kalman Filters is proposed, and also two strategies for terminating tracks and tracking objects through blind zones are described. Experimental results obtained from a sensor-equipped vehicle navigating in road traffic are presented. 1 David Hutber, Zhengyou Zhang |
BMVC | 2 |
| 1995 | An Automatic and Robust Algorithm for Determining Motion and Structure from Two Perspective Images
Zhengyou Zhang |
CAIP | 1 |
| 1995 | Estimating Motion and Structure from Correspondences of Line Segments between Two Perspective ImagesabstractPresents an algorithm for determining 3D motion and structure from correspondences of line segments between two perspective images. To our knowledge, the paper is the first investigation of use of line segments in motion and structure from motion. Classical methods use their geometric abstraction, namely straight lines, but then three images are necessary for the motion and structure determination process. We show that two views are in general sufficient when we use line segments. The assumption we use is that two matched line segments contain the projection of a common part of the corresponding line segment in space. Indeed this is what we use to match line segments between different views. Both synthetic and real data have been used to test the proposed algorithm, and excellent results have been obtained with real data containing a relatively large set of line segments. The results are comparable with those obtained using stereo calibration.> Zhengyou Zhang |
ICCV | 1 |
| 1995 | A Robust Technique for Matching two Uncalibrated Images Through the Recovery of the Unknown Epipolar Geometry
Zhengyou Zhang, Rachid Deriche, Olivier D. Faugeras, Quang-Tuan Luong |
Artif. Intell. | 1 |
| 1995 | Estimating Motion and Structure from Correspondences of Line Segments between Two Perspective ImagesabstractPresents an algorithm for determining 3D motion and structure from correspondences of line segments between two perspective images. To the author's knowledge, this paper is the first investigation of use of line segments in motion and structure from motion. Classical methods use their geometric abstraction, namely straight lines, but then three images are necessary for the motion and structure determination process. In this paper the author shows that it is possible to recover motion from two views when using line segments. The assumption used is that two matched line segments contain the projection of a common part of the corresponding line segment in space, i.e., they overlap. Indeed, this is what the author uses to match line segments between different views. This assumption constrains the possible motion between two views to an open set in motion parameter space. A heuristic, consisting of maximizing the overlap, leads to a unique solution. Both synthetic and real data have been used to test the proposed algorithm, and excellent results have been obtained with real data containing a relatively large set of line segments. Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1995 | Motion and Structure of Four Points from One Motion of a Stereo Rig with Unknown Extrinsic ParametersabstractWe describe an analytical method for recovering 3D motion and structure of four or more points from one motion of a stereo rig. The extrinsic parameters are unknown. The motion of the stereo rig is also unknown. Because of the exploitation of information redundancy, the approach gains over the traditional "motion and structure from motion" approach in that less features and less motions are required, and thus more robust estimation of motion and structure can be obtained. Since the constraint on the rotation matrix is not fully exploited in the analytical method, nonlinear minimization can be used to improve the result. We propose to estimate directly the motion and structure by minimizing the difference between the measured positions and the predicted ones in the image plane. Both computer simulated data and real data are used to validate the proposed algorithm, and very promising results are obtained. Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1994 | A Two-Stage Approach to Multi-Sensor Temporal Data FusionabstractThis paper proposes a two-stage architecture for multi-sensor temporal data fusion. The first stage uses extended Kalman filters to track tokens seen by each sensor, and the second stage links the tokens corresponding to the same real-world event. Two pairs of strategies are presented relating to the initial data association between tokens and filters, together with decision rules for switching between them. One pair is for the 'bootstrap' phase and the 'continuous' phase for an event, and the other distinguishes between a complex tracking task and a simpler one. The application of the techniques to a driver's assistant system is described. David Hutber, Zhengyou Zhang |
BMVC | 2 |
| 1994 | Self-calibration of an Uncalibrated Stereo Rig from One Unknown Motion
Zhengyou Zhang, Quang-Tuan Luong, Olivier D. Faugeras |
BMVC | 1 |
| 1994 | Robust Recovery of the Epipolar Geometry for an Uncalibrated Stereo Rig
Rachid Deriche, Zhengyou Zhang, Quang-Tuan Luong, Olivier D. Faugeras |
ECCV (1) | 2 |
| 1994 | A new and efficient iterative approach to image matchingabstractThis paper addresses the problem of matching two images with unknown epipolar geometry. A new and efficient iterative algorithm is proposed, which is a part of the author's project on developing a robust image matching technique. The author defines a new measure of matching support, which allows less contribution and higher tolerance of deformation with respect to affine transformations from distant matches than from nearby ones. A new strategy for updating matches is developed, which only selects those matches having both high matching support and low matching ambiguity. The update strategy is different from the classical "winner-take-all", which evolves too soon and is easily stuck at a local minimum, and also from "loser-take-nothing", which is usually very slow. The proposed algorithm has been tested with two dozen image pairs of very different types of scenes, and very good results have been obtained. It works remarkably well in a scene with many repetitive patterns. Zhengyou Zhang |
ICPR (1) | 1 |
| 1994 | Motion of an uncalibrated stereo rig: self-calibration and metric reconstructionabstractWe address in this paper the problem of self-calibration and metric reconstruction (up to a scale) from one unknown motion of an uncalibrated stereo rig. The epipolar constraint is first formulated for two uncalibrated images. The problem then becomes one of estimating unknowns such that the discrepancy from the epipolar constraint, in terms of sum of squared distances between points and their corresponding epipolar lines, is minimized. Redundancy of the information contained in a sequence of stereo images makes this method more robust than using a sequence of monocular images. Real data have been used to test the proposed method, and the results obtained are quite good. Zhengyou Zhang, Quang-Tuan Luong, Olivier D. Faugeras |
ICPR (1) | 1 |
| 1994 | Iterative point matching for registration of free-form curves and surfaces
Zhengyou Zhang |
Int. J. Comput. Vis. | 1 |
| 1994 | Token tracking in a cluttered scene
Zhengyou Zhang |
Image Vis. Comput. | 1 |
| 1993 | Strategies for Tracking Tokens in a Cluttered SceneabstractTracking is an important approach to analyze long sequences of images in Computer Vision. Although it has extensively been studied in other domains such as in radar imagery, it was introduced only recently in Computer Vision, and is already recognized as an efficient approach to solving correspondence and motion problems. We describe in this paper some strategies for tracking with emphasis on practical importance. They include beam search for resolving multiple matches, support of existence for discarding false matches, and locking on reliable tokens and maximizing local rigidity for handling combinatorial explosion. We have implemented those strategies in a 3D line segment tracking algorithm and found them very useful. Zhengyou Zhang |
BMVC | 1 |
| 1993 | Point Matching for Registration of Free-Form Surfaces
Zhengyou Zhang |
CAIP | 1 |
| 1993 | Motion and structure of four points from one motion of a stereo rig with unknown extrinsic parametersabstractAn analytical method for recovering 3-D motion and structure of four or more points from one motion of a stereo rig is described. The extrinsic parameters are unknown. Because of the exploitation of information redundancy, the approach gains over the traditional motion and structure from motion approach in that less features and less motions are required. Thus, more robust estimation of motion and structure can be obtained. Since the constraint on the rotation matrix is not fully exploited in the analytical method, nonlinear minimization can be used to improve the result. Both computer simulated data and real data are used to validate the proposed algorithm. Very promising results are obtained.> Zhengyou Zhang |
CVPR | 1 |
| 1992 | On Local Matching of Free-form Curves
Zhengyou Zhang |
BMVC | 1 |
| 1992 | Finding Clusters and Planes from 3D Line Segments with Application to 3D Motion Determination
Zhengyou Zhang, Olivier D. Faugeras |
ECCV | 1 |
| 1992 | Three-dimensional motion computation and object segmentation in a long sequence of stereo frames
Zhengyou Zhang, Olivier D. Faugeras |
Int. J. Comput. Vis. | 1 |
| 1992 | Estimation of Displacements from Two 3-D Frames Obtained From StereoabstractA method for estimating 3D displacements from two stereo frames is presented. It is based on the hypothesize-and-verify paradigm used to match 3D line segments between the two frames. In order to reduce the complexity of the method, an assumption is made that objects are rigid. The formulate a set of complete rigidity constraints for 3D line segments and integrate the uncertainty of measurements in this formation. The hypothesize-and-verify stages of the method use an extended Kalman filter to produce estimates of the displacements and of their uncertainty. The algorithm is shown to work on indoor and natural scenes. It is also shown to be easily extended to the case in which several mobile objects are present. The method is quite robust, fast, and has been thoroughly tested on hundreds of real stereo frames.> Zhengyou Zhang, Olivier D. Faugeras |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1991 | Determining motion from 3D line segment matches: a comparative study
Zhengyou Zhang, Olivier D. Faugeras |
Image Vis. Comput. | 1 |
| 1990 | Determining motion from 3D line segment matches: a comparative studyabstractAbstract Motion estimation is a very important problem in dynamic scene analysis. Although it is easier to estimate motion parameters from 3D data than from 2D images, it is not trivial, since the 3D data we have are almost always corrupted by noise. A comparative study on motion estimation from 3D line segments is presented. Two representations of line segments and two representations of rotation are described. With different representations of line segments and rotation a number of methods for motion estimation are presented, including the extended Kalman filter a general minimization process and the singular value decomposition. These methods are compared using both synthetic and real data obtained by a trinocular stereo. It is observed that the extended Kalman filter with the rotation axis representation of rotation is preferable. Note that all methods discussed can be directly applied to 3D point data. Zhengyou Zhang, Olivier D. Faugeras |
BMVC | 1 |
| 1990 | Tracking and Motion Estimation in a Sequence of Stereo Frames
Zhengyou Zhang, Olivier D. Faugeras |
ECAI | 1 |
| 1990 | Tracking and grouping 3D line segmentsabstractThe authors address the motion tracking problem that arises in the context of a mobile vehicle navigating in an unknown environment where other mobile bodies such as human beings or robots may also be moving. A stereo rig mounted on the mobile vehicle provides a sequence of 3-D maps of the environment. The current stereo system is trinocular and the 3-D tokens used at the moment are line segments corresponding to significant intensity discontinuities in the images. Although the framework to solve the motion tracking problem developed in this work arises in this specific context, the authors believe it should be applicable in other contexts; in particular they could use other 3-D primitives, for example, points, combinations of points and lines, curves.> Zhengyou Zhang, Olivier D. Faugeras |
ICCV | 1 |
| 1990 | Building a 3D world model with a mobile robot: 3D line segment representation and integrationabstractA system whereby a mobile robot in an unknown environment can incrementally build a world model is described. The model discussed is segment-based. A trinocular system is used to build a local map of the environment. A global map is obtained by integrating a sequence of stereo frames taken while the robot navigates in the environment. Emphasis is on the representation of the uncertainty of 3D segments from stereo and on the integration of segments from multiple views. The representation is simple and very convenient for characterizing the uncertainty of segments. A Kalman filter is used to merge matched line segments. An important characteristic of the integration strategy is that a segment observed by the stereo system corresponds only to one part of the segment in space, so that the union of different observations gives a better estimate on the segment in space. The integration of 35 stereo frames taken in a robot room is described.> Zhengyou Zhang, Olivier D. Faugeras |
ICPR (1) | 1 |
| 1988 | Analysis Of A Sequence Of Stereo Scenes Containing Multiple Moving Objects Using Rigidity ConstraintsabstractIri this paper, we describe a method for comput.ing the rriovc~rrient. of objects as well as that of a mobile robot from ii scqiicrice of stereo frames. Stereo frames are obtained at ~iifl’(~reiit, instants by a stereo rig, when the mobile robot rIilvigat(>s in an unknown environment possibly containing ~)iiit: rrioving rigid objects. An approach based on rigidity ( oiihlraiiit,s is presented for registering two stereo frames. Wcs dernoristrate how the uncertainty of measurements can IN^ integrated with the formalism of the rigidity constraints. A iiew technique is described to match very noisy segments. ‘1‘11~ iiifluence of egomotion on observed movement:; of ob,j(~ 1,s is discussed in detail. Egomotion is fir:jt determined itiid then eliminated before determination of the motion of o1)jt:cl.s. The proposed algorithm is completely automatic. I~:xperirnerital results are provided. Some remarks conclude ths paper. Zhengyou Zhang, Olivier D. Faugeras, Nicholas Ayache |
ICCV | 1 |