Hao Jiang 0007

dblp:38/6049-7 · DBLP profile ↗
← Back
48ranked-venue papers
31as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 26 first-author · 7 since 2021Artificial intelligence and machine learning · 33 · 23 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Ego4D: Around the World in 3,600 Hours of Egocentric Video
abstract
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception.
Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Devansh Kukreja, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik
IEEE Trans. Pattern Anal. Mach. Intell.9
2024 The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
abstract
In recent years, the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions, where both visual and audio signals play a crucial role. While most prior work focus on learning about behaviors that directly involve the camera wearer, we introduce the Ego-Exocentric Conversational Graph Prediction problem, marking the first attempt to infer exocentric conversational interactions from egocentric videos. We propose a unified multi-modal framework-Audio- Visual Conversational Attention (AV-CONV), for the joint prediction of conversation behaviors-speaking and listening-for both the camera wearer as well as all other social partners present in the egocentric video. Specifically, we adopt the self-attention mechanism to model the representations across-time, across-subjects, and across-modalities. To validate our method, we conduct experiments on a challenging egocentric video dataset that includes multi-speaker and multi-conversation scenarios. Our results demonstrate the superior performance of our method compared to a series of baselines. We also present detailed ablation studies to assess the contribution of each component in our model. Check our Project Page.
Wenqi Jia 0001, Miao Liu 0007, Hao Jiang 0007, Ishwarya Ananthabhotla, James M. Rehg, Vamsi K. Ithapu, Ruohan Gao
CVPR3
2023 Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations
abstract
Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D environment by exploiting shared information in the egocentric audio-visual observations of participants in a natural conversation. Our hypothesis is that as multiple people (“egos”) move in a scene and talk among themselves, they receive rich audio-visual cues that can help uncover the unseen areas of the scene. Given the high cost of continuously processing egocentric visual streams, we further explore how to actively coordinate the sampling of visual information, so as to minimize redundancy and reduce power use. To that end, we present an audio-visual deep reinforcement learning approach that works with our shared scene mapper to selectively turn on the camera to efficiently chart out the space. We evaluate the approach using a state-of-the-art audio-visual simulator for 3D scenes as well as real-world video. Our model outperforms previous state-of-the-art mapping methods, and achieves an excellent cost-accuracy tradeoff. Project: http://vision.cs.utexas.edu/projects/chat2map.
Sagnik Majumder, Hao Jiang 0007, Pierre Moulon, Ethan Henderson, Paul Calamia, Kristen Grauman, Vamsi K. Ithapu
CVPR2
2023 Egocentric Auditory Attention Localization in Conversations
abstract
In a noisy conversation environment such as a dinner party, people often exhibit selective auditory attention, or the ability to focus on a particular speaker while tuning out others. Recognizing who somebody is listening to in a conversation is essential for developing technologies that can understand social behavior and devices that can augment human hearing by amplifying particular sound sources. The computer vision and audio research communities have made great strides towards recognizing sound sources and speakers in scenes. In this work, we take a step further by focusing on the problem of localizing auditory attention targets in egocentric video, or detecting who in a camera wearer's field of view they are listening to. To tackle the new and challenging Selective Auditory Attention Localization problem, we propose an end-to-end deep learning approach that uses egocentric video and multichannel audio to predict the heatmap of the camera wearer's auditory attention. Our approach leverages spatiotemporal audiovisual features and holistic reasoning about the scene to make predictions, and outperforms a set of baselines on a challenging multi-speaker conversation dataset. Project page: https://fkryan.github.io/saal
Fiona Ryan, Hao Jiang 0007, Abhinav Shukla, James M. Rehg, Vamsi K. Ithapu
CVPR2
2022 Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization
abstract
Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these social interactions first requires detecting and localizing the voice activities of the device wearer and the surrounding people. These tasks are challenging due to their egocentric nature: the wearer's head motion may cause motion blur, surrounding people may appear in difficult viewing angles, and there may be occlusions, visual clutter, audio noise, and bad lighting. Under these conditions, previous state-of-the-art active speaker detection methods do not give satisfactory results. Instead, we tackle the problem from a new setting using both video and multi-channel microphone array audio. We propose a novel end-to-end deep learning approach that is able to give robust voice activity detection and localization results. In contrast to previous methods, our method localizes active speakers from all possible directions on the sphere, even outside the camera's field of view, while simultaneously detecting the device wearer's own voice activity. Our experiments show that the proposed method gives superior results, can run in real time, and is robust against noise and clutter.
Hao Jiang 0007, Calvin Murdock, Vamsi K. Ithapu
CVPR1
2022 Ego4D: Around the World in 3, 000 Hours of Egocentric Video
abstract
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik
CVPR8
2021 On the Predictability of Hrtfs from Ear Shapes Using Deep Networks
abstract
Head-Related Transfer Function (HRTF) individualization is critical for immersive and realistic spatial audio rendering in augmented/virtual reality. Neither measurements nor simulations using 3D scans of head/ear are scalable for practical applications. More efficient machine learning approaches are being explored recently, to predict HRTFs from ear images or anthropometric features. However, it is not yet clear whether such models can provide an alternative for direct measurements or high-fidelity simulations. Here, we aim to address this question. Using 3D ear shapes as inputs, we explore the bounds of HRTF predictability using deep neural networks. To that end, we propose and evaluate two models, and identify the lowest achievable spectral distance error when predicting the true HRTF magnitude spectra.
Yaxuan Zhou, Hao Jiang 0007, Vamsi K. Ithapu
ICASSP2
2021 Egocentric Pose Estimation from Human Vision Span
abstract
Estimating camera wearer's body pose from an egocentric view (egopose) is a vital task in augmented and virtual reality. Existing approaches either use a narrow field of view front facing camera that barely captures the wearer, or an extended head-mounted top-down camera for maximal wearer visibility. In this paper, we tackle the egopose estimation from a more natural human vision span, where camera wearer can be seen in the peripheral view and depending on the head pose the wearer may become invisible or has a limited partial view. This is a realistic visual field for user-centric wearable devices like glasses which have front facing wide angle cameras. Existing solutions are not appropriate for this setting, and so, we propose a novel deep learning system taking advantage of both the dynamic features from camera SLAM and the body shape imagery. We compute 3D head pose, 3D body pose, the figure/ground separation, all at the same time while explicitly enforcing a certain geometric consistency across pose attributes. We further show that this system can be trained robustly with lots of existing mocap data so we do not have to collect and annotate large new datasets. Lastly, our system estimates egopose in real time and on the fly while maintaining high accuracy.
Hao Jiang 0007, Vamsi K. Ithapu
ICCV1
2019 Action4D: Online Action Recognition in the Crowd and Clutter
abstract
Recognizing every person's action in a crowded and cluttered environment is a challenging task in computer vision. We propose to tackle this challenging problem using a holistic 4D ``scan'' of a cluttered scene to include every detail about the people and environment. This leads to a new problem, i.e., recognizing multiple people's actions in the cluttered 4D representation. At the first step, we propose a new method to track people in 4D, which can reliably detect and follow each person in real time. Then, we build a new deep neural network, the Action4DNet, to recognize the action of each tracked person. Such a model gives reliable and accurate results in the real-world settings. We also design an adaptive 3D convolution layer and a novel discriminative temporal feature learning objective to further improve the performance of our model. Our method is invariant to camera view angles, resistant to clutter and able to handle crowd. The experimental results show that the proposed method is fast, reliable and accurate. Our method paves the way to action recognition in the real-world applications and is ready to be deployed to enable smart homes, smart factories and smart stores.
Quanzeng You, Hao Jiang 0007
CVPR2
2017 Detangling People: Individuating Multiple Close People and Their Body Parts via Region Assembly
abstract
Todays person detection methods work best when people are in common upright poses and appear reasonably well spaced out in the image. However, in many real images, thats not what people do. People often appear quite close to each other, e.g., with limbs linked or heads touching, and their poses are often not pedestrian-like. We propose an approach to detangle people in multi-person images. We formulate the task as a region assembly problem. Starting from a large set of overlapping regions from body part semantic segmentation and generic object proposals, our optimization approach reassembles those pieces together into multiple person instances. Since optimal region assembly is a challenging combinatorial problem, we present a Lagrangian relaxation method to accelerate the lower bound estimation, thereby enabling a fast branch and bound solution for the global optimum. As output, our method produces a pixel-level map indicating both 1) the body part labels (arm, leg, torso, and head), and 2) which parts belong to which individual person. Our results on challenging datasets show our method is robust to clutter, occlusion, and complex poses. It outperforms a variety of competing methods, including existing detector CRF methods and region CNN approaches. In addition, we demonstrate its impact on a proxemics recognition task, which demands a precise representation of whose body part is where in crowded images.
Hao Jiang 0007, Kristen Grauman
CVPR1
2017 Seeing Invisible Poses: Estimating 3D Body Pose from Egocentric Video
abstract
Understanding the camera wearers activity is central to egocentric vision, yet one key facet of that activity is inherently invisible to the camera-the wearers body pose. Prior work focuses on estimating the pose of hands and arms when they come into view, but this 1) gives an incomplete view of the full body posture, and 2) prevents any pose estimate at all in many frames, since the hands are only visible in a fraction of daily life activities. We propose to infer the invisible pose of a person behind the egocentric camera. Given a single video, our efficient learning-based approach returns the full body 3D joint positions for each frame. Our method exploits cues from the dynamic motion signatures of the surrounding scene-which change predictably as a function of body pose-as well as static scene structures that reveal the viewpoint (e.g., sitting vs. standing). We further introduce a novel energy minimization scheme to infer the pose sequence. It uses soft predictions of the poses per time instant together with a non-parametric model of human pose dynamics over longer windows. Our method outperforms an array of possible alternatives, including typical deep learning approaches for direct pose regression from images.
Hao Jiang 0007, Kristen Grauman
CVPR1
2016 3D Human Pose Estimation via Deep Learning from 2D Annotations
abstract
We propose a deep convolutional neural network for 3D human pose and camera estimation from monocular images that learns from 2D joint annotations. The proposed network follows the typical architecture, but contains an additional output layer which projects predicted 3D joints onto 2D, and enforces constraints on body part lengths in 3D. We further enforce pose constraints using an independently trained network that learns a prior distribution over 3D poses. We evaluate our approach on several benchmark datasets and compare against state-of-the-art approaches for 3D human pose estimation, achieving comparable performance. Additionally, we show that our approach significantly outperforms other methods in cases where 3D ground truth data is unavailable, and that our network exhibits good generalization properties.
Ernesto Brau, Hao Jiang 0007
3DV2
2016 A Bayesian part-based approach to 3D human pose and camera estimation
abstract
We present a Bayesian framework for estimating 3D human pose and camera from a single RGB image. We develop a generative model where a 3D pose is rendered onto an image (via the camera), which then generates a detection probability map for each body part. We represent a human pose with a set of 3D cylinders in space, one for each body part, and we place kinematic and self-intersection priors on the model. Importantly, we use a graphics engine (e.g., OpenGL) to render the pose, and use its built-in capabilities for color blending to efficiently compute the likelihood of the model given the observed probability maps, which are obtained by running a convolutional neural network classifier on a test image. We explore the space of 3D poses and camera configurations via the Hybrid Monte Carlo algorithm, with sampling moves designed specifically for this problem. We train the parameters of our prior and likelihood distributions using annotated poses from the CMU mocap database, and test our algorithm on two benchmark datasets, where we compare performance against state-of-the-art methods. Additionally, we demonstrate the flexibility of our framework by incorporating a likelihood function for depth images and showing the associated performance gains.
Ernesto Brau, Hao Jiang 0007
ICPR2
2016 Human's Scene Sketch Understanding
abstract
Human's sketch understanding is important. It has many applications in human computer interaction, multimedia, and computer vision. Recognizing human sketches is also challenging. Previous methods focus on single-object sketch recognition. Understanding human's scene sketch that involves multiple objects and their complex interactions has not been explored. In this paper, we tackle this new problem. We create the first scene sketch dataset "Scene250" and propose a deep learning method to understand human scene sketches. We propose "Scene-Net", a new deep convolutional neural network (CNN) structure, based on which we build a novel scene sketch recognition system. Our system has been tested on the collected scene sketch dataset and compared with other state-of-the-art CNNs and sketch recognition approaches. Our experimental results demonstrate that our method achieves the state of art.
Yuxiang Ye, Yijuan Lu, Hao Jiang 0007
ICMR3
2015 Matching bags of regions in RGBD images
abstract
We study the new problem of matching regions between a pair of RGBD images given a large set of overlapping region proposals. These region proposals do not have a tree hierarchy and are treated as bags of regions. Matching RGBD images using bags of region candidates with unstructured relations is a challenging combinatorial problem. We propose a linear formulation, which optimizes the region selection and matching simultaneously so that the matched regions have similar color histogram, shape, and small overlaps, the selected regions have a small number and overall low concavity, and they tend to cover both of the images. We efficiently compute the lower bound by solving a sequence of min-cost bipartite matching problems via Lagrangian relaxation and we obtain the global optimum using branch and bound. Our experiments show that the proposed method is fast, accurate, and robust against cluttered scenes.
Hao Jiang 0007
CVPR1
2015 Scale and Rotation Invariant Matching Using Linearly Augmented Trees
abstract
We propose a novel linearly augmented tree method for efficient scale and rotation invariant object matching. The proposed method enforces pairwise matching consistency defined on trees, and high-order constraints on all the sites of a template. The pairwise constraints admit arbitrary metrics while the high-order constraints use L1 norms and therefore can be linearized. Such a linearly augmented tree formulation introduces hyperedges and loops into the basic tree structure. But, different from a general loopy graph, its special structure allows us to relax and decompose the optimization into a sequence of tree matching problems that are efficiently solvable by dynamic programming. The proposed method also works on continuous scale and rotation parameters; we can match with a scale up to any large value with the same efficiency. Our experiments on ground truth data and a variety of real images and videos show that the proposed method is efficient, accurate and reliable.
Hao Jiang 0007, Tai-Peng Tian, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Finding Approximate Convex Shapes in RGBD Images
Hao Jiang 0007
ECCV (3)1
2013 A Linear Approach to Matching Cuboids in RGBD Images
abstract
We propose a novel linear method to match cuboids in indoor scenes using RGBD images from Kinect. Beyond depth maps, these cuboids reveal important structures of a scene. Instead of directly fitting cuboids to 3D data, we first construct cuboid candidates using super pixel pairs on a RGBD image, and then we optimize the configuration of the cuboids to satisfy the global structure constraints. The optimal configuration has low local matching costs, small object intersection and occlusion, and the cuboids tend to project to a large region in the image, the number of cuboids is optimized simultaneously. We formulate the multiple cuboid matching problem as a mixed integer linear program and solve the optimization efficiently with a branch and bound method. The optimization guarantees the global optimal solution. Our experiments on the Kinect RGBD images of a variety of indoor scenes show that our proposed method is efficient, accurate and robust against object appearance variations, occlusions and strong clutter.
Hao Jiang 0007, Jianxiong Xiao
CVPR1
2013 Human movement summarization and depiction from videos
abstract
Human movement summarization and depiction from videos is to automatically turn an input video into high level action illustrations, in which the movements of the body parts are visualized using arrows and motion particles. Motion depiction compactly illustrates how specific movements are performed. Previous action summarization methods reply on 3D motion capture or manually labeled data, without which depicting actions is a challenging task. In this paper, we propose a novel scheme to automatically summarize and depict human movements from 2D videos without 3D motion capture or manually labeled data. The proposed method first segments videos into sub-actions with an effective streamline matching scheme. Then, to estimate human movement, we propose a novel trajectory following method to track points by using both body part detection and optical flow. With the estimated movement, we depict the human articulated motion with arrows and motion particles. Our experiments on a variety of videos show that the proposed method is effective in summarizing complex human movements and generating compact depictions.
Yijuan Lu, Hao Jiang 0007
ICME2
2012 Linear solution to scale invariant global figure ground separation
abstract
We propose a novel linear method for scale invariant figure ground separation in images and videos. Figure ground separation is treated as a superpixel labeling problem. We optimize superpixel foreground and background labeling so that the object foreground estimation matches model color histogram, its area and perimeter are consistent with object shape prior, and the foreground superpixels form a connected region. This optimization problem is challenging due to high-order soft and hard global constraints among large number of superpixels. We devise a scale invariant linear method that gives an integer solution with a guaranteed error bound via a branch and cut procedure. The proposed method does not rely on motion continuity and works on static images and videos with abrupt motion. Our experimental results on both synthetic ground truth data and real images show that the proposed method is efficient and robust over object appearance changes, large deformation and strong background clutter.
Hao Jiang 0007
CVPR1
2012 Scale resilient, rotation invariant articulated object matching
abstract
A novel method is proposed for matching articulated objects in cluttered videos. The method needs only a single exemplar image of the target object. Instead of using a small set of large parts to represent an articulated object, the proposed model uses hundreds of small units to represent walks along paths of pixels between key points on an articulated object. Matching directly on dense pixels is key to achieving reliable matching when motion blur occurs. The proposed method fits the model to local image properties, conforms to structure constraints, and remembers the steps taken along a pixel path. The model formulation handles variations in object scaling, rotation and articulation. Recovery of the optimal pixel walks is posed as a special shortest path problem, which can be solved efficiently via dynamic programming. Further speedup is achieved via factorization of the path costs. An efficient method is proposed to find multiple walks and simultaneously match multiple key points. Experiments show that the proposed method is efficient and reliable and can be used to match articulated objects in fast motion videos with strong clutter and blurry imagery.
Hao Jiang 0007, Tai-Peng Tian, Kun He 0003, Stan Sclaroff
CVPR1
2012 Finding People Using Scale, Rotation and Articulation Invariant Matching
Hao Jiang 0007
ECCV (4)1
2011 Scale and rotation invariant matching using linearly augmented trees
abstract
We propose a novel linearly augmented tree method for efficient scale and rotation invariant object matching. The proposed method enforces pairwise matching consistency defined on trees, and high-order constraints on all the sites of a template. The pairwise constraints admit arbitrary metrics while the high-order constraints use L1 norms and therefore can be linearized. Such a linearly augmented tree formulation introduces hyperedges and loops into the basic tree structure, but different from a general loopy graph, its special structure allows us to relax and decompose the optimization into a sequence of tree matching problems efficiently solvable by dynamic programming. The proposed method also works on continuous scale and rotation parameters; we can match with a scale up to any large number with the same efficiency. Our experiments on ground truth data and a variety of real images and videos show that the proposed method is efficient, accurate and reliable.
Hao Jiang 0007, Tai-Peng Tian, Stan Sclaroff
CVPR1
2011 Human Pose Estimation Using Consistent Max Covering
abstract
A novel consistent max-covering method is proposed for human pose estimation. We focus on problems in which a rough foreground estimation is available. Pose estimation is formulated as a jigsaw puzzle problem in which the body part tiles maximally cover the foreground region, match local image features, and satisfy body plan and color constraints. This method explicitly imposes a global shape constraint on the body part assembly. It anchors multiple body parts simultaneously and introduces hyperedges in the part relation graph, which is essential for detecting complex poses. Using multiple cues in pose estimation, our method is resistant to cluttered foregrounds. We propose an efficient linear method to solve the consistent max-covering problem. A two-stage relaxation finds the solution in polynomial time. Our experiments on a variety of images and videos show that the proposed method is more robust than previous locally constrained methods.
Hao Jiang 0007
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Linear Scale and Rotation Invariant Matching
abstract
Matching visual patterns that appear scaled, rotated, and deformed with respect to each other is a challenging problem. We propose a linear formulation that simultaneously matches feature points and estimates global geometrical transformation in a constrained linear space. The linear scheme enables search space reduction based on the lower convex hull property so that the problem size is largely decoupled from the original hard combinatorial problem. Our method therefore can be used to solve large scale problems that involve a very large number of candidate feature points. Without using prepruning in the search, this method is more robust in dealing with weak features and clutter. We apply the proposed method to action detection and image matching. Our results on a variety of images and videos demonstrate that our method is accurate, efficient, and robust.
Hao Jiang 0007, Stella X. Yu, David R. Martin 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 Finding Human Poses in Videos Using Concurrent Matching and Segmentation
Hao Jiang 0007
ACCV (1)1
2010 Building a Videorama with Shallow Depth of Field
abstract
This paper presents a new automatic approach to building a videorama with shallow depth of field. We stitch the static background of video frames and render the dynamic foreground onto the enlarged background after foreground/background segmentation. To this end, we extract the depth information from a two-view video stream. We show that the depth cues combined with color cues improve segmentation. Finally, we use the depth cues to synthesize the shallow depth of field effects in the final videorama. Our approach stabilizes the camera motion as if the video was captured from a static camera and improves the visual quality with the increased field of view and shallow depth of field effects.
Soonmin Bae, Hao Jiang 0007
ICPR2
2010 Action Detection in Cluttered Video With Successive Convex Matching
abstract
We propose a novel successive convex matching method for human action detection in cluttered video. Human actions are represented as sequences of poses, and specific actions are detected by matching pose sequences. Since we represent actions as the evolution of poses and shapes, the proposed method can detect actions in videos that involve fast camera motions. Template sequence to video registration is nonlinear and highly nonconvex. Instead of directly solving the hard problem, our method convexifies it into a sequence of linear programs and refines the matching by successive trust region shrinkage. The proposed scheme further simplifies the linear programs by representing the target point space with a small set of basis points. The low complexity of the proposed method enables it to search efficiently in a large range. Experiments show that successive convex matching can robustly match a sequence of coupled shape templates simultaneously to target sequences and effectively detect specific actions in cluttered videos.
Hao Jiang 0007, Mark S. Drew, Ze-Nian Li
IEEE Trans. Circuits Syst. Video Technol.1
2009 Linear solution to scale and rotation invariant object matching
abstract
Images of an object undergoing ego- or camera-motion often appear to be scaled, rotated, and deformed versions of each other. To detect and match such distorted patterns to a single sample view of the object requires solving a hard computational problem that has eluded most object matching methods. We propose a linear formulation that simultaneously finds feature point correspondences and global geometrical transformations in a constrained solution space. Further reducing the search space based on the lower convex hull property of the formulation, our method scales well with the number of candidate features. Our results on a variety of images and videos demonstrate that our method is accurate, efficient, and robust over local deformation, occlusion, clutter, and large geometrical transformations.
Hao Jiang 0007, Stella X. Yu
CVPR1
2008 Global pose estimation using non-tree models
abstract
We propose a novel global pose estimation method to detect body parts of articulated objects in images based on non-tree graph models. There are two kinds of edges defined in the body part relation graph: Strong (tree) edges corresponding to the body plan that can enforce any type of constraint, and weak (non-tree) edges that express exclusion constraints arising from inter-part occlusion and symmetry conditions. We express optimal part localization as a multiple shortest path problem in a set of correlated trellises constructed from the graph model. Strong model edges generate the trellises, while weak model edges prohibit implausible poses by generating exclusion constraints among trellis nodes and edges. The optimization may be expressed as an integer linear program and solved using a novel two-stage relaxation scheme. Experiments show that the proposed method has a high chance of obtaining the globally optimal pose at low computational cost.
Hao Jiang 0007, David R. Martin 0001
CVPR1
2008 Finding Actions Using Shape Flows
Hao Jiang 0007, David R. Martin 0001
ECCV (2)1
2008 Optimizing Multiple Object Tracking and Best View Video Synthesis
abstract
We study schemes to tackle problems of optimizing multiple object tracking and best-view video synthesis. A novel linear relaxation method is proposed for the class of multiple object tracking problems where the inter-object interaction metric is convex and the intra-object term quantifying object state continuity may use any metric. This scheme models object tracking as multi-path searching. It explicitly models track interaction, such as object spatial layout consistency or mutual occlusion, and optimizes multiple object tracks simultaneously. The proposed scheme does not rely on track initialization and complex heuristics. It has much less average complexity than previous efficient exhaustive search methods such as extended dynamic programming and can find the global optimum with high probability. Given the tracking data from our method, optimizing best-view video synthesis using multiple-view videos is further studied, which is formulated as a recursive decision problem and optimized by a dynamic programming approach. The proposed object tracking and best-view synthesis methods have found successful applications in MyView - a system to enhance media content presentation of multiple-view video.
Hao Jiang 0007, Sidney S. Fels, James J. Little
IEEE Trans. Multim.1
2007 A Linear Programming Approach for Multiple Object Tracking
abstract
We propose a linear programming relaxation scheme for the class of multiple object tracking problems where the inter-object interaction metric is convex and the intra-object term quantifying object state continuity may use any metric. The proposed scheme models object tracking as a multi-path searching problem. It explicitly models track interaction, such as object spatial layout consistency or mutual occlusion, and optimizes multiple object tracks simultaneously. The proposed scheme does not rely on track initialization and complex heuristics. It has much less average complexity than previous efficient exhaustive search methods such as extended dynamic programming and is found to be able to find the global optimum with high probability. We have successfully applied the proposed method to multiple object tracking in video streams.
Hao Jiang 0007, Sidney S. Fels, James J. Little
CVPR1
2007 Matching by Linear Programming and Successive Convexification
abstract
We present a novel convex programming scheme to solve matching problems, focusing on the challenging problem of matching in a large search range and with cluttered background. Matching is formulated as metric labeling with L1 regularization terms, for which we propose a novel linear programming relaxation method and an efficient successive convexification implementation. The unique feature of the proposed relaxation scheme is that a much smaller set of basis labels is used to represent the original label space. This greatly reduces the size of the searching space. A successive convexification scheme solves the labeling problem in a coarse to fine manner. Importantly, the original cost function is reconvexified at each stage, in the new focus region only, and the focus region is updated so as to refine the searching result. This makes the method well-suited for large label set matching. Experiments demonstrate successful applications of the proposed matching scheme in object detection, motion estimation, and tracking.
Hao Jiang 0007, Mark S. Drew, Ze-Nian Li
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Shadow resistant tracking using inertia constraints
Hao Jiang 0007, Mark S. Drew
Pattern Recognit.1
2006 Successive Convex Matching for Action Detection
abstract
We propose human action detection based on a successive convex matching scheme. Human actions are represented as sequences of postures and specific actions are detected in video by matching the time-coupled posture sequences to video frames. The template sequence to video registration is formulated as an optimal matching problem. Instead of directly solving the highly non-convex problem, our method convexifies the matching problem into linear programs and refines the matching result by successively shrinking the trust region. The proposed scheme represents the target point space with small sets of basis points and therefore allows efficient searching. This matching scheme is applied to robustly matching a sequence of coupled binary templates simultaneously in a video sequence with cluttered backgrounds.
Hao Jiang 0007, Mark S. Drew, Ze-Nian Li
CVPR (2)1
2006 Unsupervised Discovery of Action Classes
abstract
In this paper we consider the problem of describing the action being performed by human figures in still images. We will attack this problem using an unsupervised learning approach, attempting to discover the set of action classes present in a large collection of training images. These action classes will then be used to label test images. Our approach uses the coarse shape of the human figures to match pairs of images. The distance between a pair of images is computed using a linear programming relaxation technique. This is a computationally expensive process, and we employ a fast pruning method to enable its use on a large collection of images. Spectral clustering is then performed using the resulting distances. We present clustering and image labeling results on a variety of datasets.
Yang Wang 0003, Hao Jiang 0007, Mark S. Drew, Ze-Nian Li, Greg Mori
CVPR (2)2
2006 Detecting Human Action in Active Video
abstract
We propose a novel scheme to detect human actions in active video. Active videos such as movies or sports broadcasting are taken purposively by "clever" photographers. They are object and action oriented and usually involve complex camera motions. Detecting actions in active videos is both important and challenging. We study a three-step scheme to detect complex human actions in such videos. The proposed method first locates potential objects and removes clutter with a composite filter scheme. The detected object candidates in successive frames are then associated to form object trajectories based on a consistent labeling formulation, and solved with belief propagation. Finally, specific human actions are detected in video with a linear programming matching approach that can efficiently deal with matching problems having a large target point set. The proposed method has been successfully applied in action detection for general videos and TV hockey games
Hao Jiang 0007, Ze-Nian Li, Mark S. Drew
ICME1
2005 Human Posture Recognition with Convex Programming
abstract
We present a novel human posture recognition method us ing convex programming based matching schemes. Instead of trying to segment the object from the background, we develop a novel multistage linear programming scheme to locate the target by searching for the best matching region based on an automatically acquired graph template. The linear programming based visual matching scheme gener ates relatively dense matching patterns and thus presents a key for robust object matching and human posture recogni tion. By matching distance transformations of edge maps, the proposed scheme is able to match figures with large ap pearance changes. We further present object recognition methods based on the similarity of the exemplar with the matching target. The proposed scheme can also be used for recognizing multiple targets in an image. Experiments show promising results for recognizing human postures in clut tered environments.
Hao Jiang 0007, Ze-Nian Li, Mark S. Drew
ICME1
2004 Optimizing Motion Estimation with Linear Programming and Detail-Preserving Variational Method
Hao Jiang 0007, Ze-Nian Li, Mark S. Drew
CVPR (1)1
2003 Nondiagonal color correction
abstract
A new color correction method is introduced which predicts how changing the color of the scene illuminant will affect a camera's RGB response. Like diagonal transformation color correction methods, the new method requires only 3-parameters. It therefore requires only the RGB color of the two illuminants be known. The method models the 9-parameters of a 3-by-3 linear transformation using a 3-dimensional linear model composed of 3 basis transformations. Experiments show that the method works better than the standard diagonal model unless the camera sensors are very sharply peaked, in which case the performance is essentially unchanged.
Brian V. Funt, Hao Jiang 0007
ICIP (1)2
2003 Shadow-resistant tracking in video
abstract
In this paper, we present a new method for tracking objects with shadows. Traditional motion-based tracking schemes cannot usually distinguish the shadow from the object itself, and this results in a falsely captured object shape, posing a severe difficulty for a pattern recognition task. In this paper we present a color processing scheme to project the image into an illumination invariant space such that the shadow's effect is greatly attenuated. The optical flow in this projected image together with the original image is used as a reference for object tracking so that we can extract the real object shape in the tracking process. We present a modified snake model for general video object tracking. Two new external forces are introduced into the snake equation based on the predictive contour such that (1) the active contour is attracted to a shape similar to the one in the previous video frame, and (2) chordal string constraints across the shape are applied so that the snake is correctly maintained when only partial features are obtained in some frames. The proposed method can deal with the problem of an object's ceasing movements temporarily, and can also avoid the problem of the snake tracking into the object interior. Global affine motion compensation makes the method applicable in a general video environment. Experimental results show that the proposed method can track the real object even if there is strong shadow influence.
Hao Jiang 0007, Mark S. Drew
ICME1
2002 A predictive contour inertia snake model for general video tracking
abstract
We present a modified snake model for the problem of general video object tracking. We introduce a new external force into the snake equation based on the predictive contour such that the active contour is attracted to a shape similar to the one in the previous video frame. New methods of contour prediction and contour smoothing are presented. The proposed methods can deal with the problem of an object's stopping movement temporarily and can also avoid the problem of the snake tracking into the object interior. Global affine motion estimation is applied to eliminate the effect of camera motion and hence the method can be applied in a general video environment. Experimental results show that the proposed method exhibits increased robustness over a traditional snake algorithm and works well for general video object tracking.
Hao Jiang 0007, Mark S. Drew
ICIP (3)1
2002 Error concealment using a diffusion based method
abstract
In this paper, we present a novel PDE based error concealment algorithm. We formulate the error concealment problem as a sequential optimization problem with both smoothing and orientation constraints. By introducing the orientation constraint we convert a nonlinear variational problem into a problem that is well posed and which can be solved without iterative operations. A modified orientation diffusion scheme is presented which is able to reconstruct complex orientation patterns within blocks which have been lost in an image. In the intensity reconstruction stage which follows orientation diffusion, optimization is performed based on the orientation estimates from the first stage together with the constraint of smoothness on block boundaries. We present an efficient numerical scheme which implements the method without iterations.
Hao Jiang 0007, Cecilia Moloney
ICIP (1)1
2002 A new direction adaptive scheme for image interpolation
abstract
We present a novel image interpolation method based on variational models with both smoothing and orientation constraints. By introducing the orientation constraint, we simplify the nonlinear PDE problem into a series of problems with explicit solutions. In our model, the gradient directions for the interpolated pixels are first estimated using a modified orientation diffusion method. Using these estimated gradient directions adaptive directional interpolation is carried out. An effective numerical implementation of the adaptive directional interpolation is presented for the case of upsampling by factors of two. This implementation had very low complexity and is well suited for real-time applications.
Hao Jiang 0007, Cecilia Moloney
ICIP (3)1
2002 Content analysis for audio classification and segmentation
abstract
We present our study of audio content analysis for classification and segmentation, in which an audio stream is segmented according to audio type or speaker identity. We propose a robust approach that is capable of classifying and segmenting an audio stream into speech, music, environment sound, and silence. Audio classification is processed in two steps, which makes it suitable for different applications. The first step of the classification is speech and nonspeech discrimination. In this step, a novel algorithm based on K-nearest-neighbor (KNN) and linear spectral pairs-vector quantization (LSP-VQ) is developed. The second step further divides nonspeech class into music, environment sounds, and silence with a rule-based classification scheme. A set of new features such as the noise frame ratio and band periodicity are introduced and discussed in detail. We also develop an unsupervised speaker segmentation algorithm using a novel scheme based on quasi-GMM and LSP correlation analysis. Without a priori knowledge, this algorithm can support the open-set speaker, online speaker modeling and real time segmentation. Experimental results indicate that the proposed algorithms can produce very satisfactory results.
Lie Lu, HongJiang Zhang, Hao Jiang 0007
IEEE Trans. Speech Audio Process.3
2001 A robust audio classification and segmentation method
abstract
In this paper, we present a robust algorithm for audio classification that is capable of segmenting and classifying an audio stream into speech, music, environment sound and silence. Audio classification is processed in two steps, which makes it suitable for different applications. The first step of the classification is speech and non-speech discrimination. In this step, a novel algorithm based on KNN and LSP VQ is presented. The second step further divides non-speech class into music, environment sounds and silence with a rule based classification scheme. Some new features such as the noise frame ratio and band periodicity are introduced and discussed in detail. Our experiments in the context of video structure parsing have shown the algorithms produce very satisfactory results.
Lie Lu, Hao Jiang 0007, HongJiang Zhang
ACM Multimedia2
2000 Integrating Visual, Audio and Text Analysis for News Video
abstract
We present a system developed for content-based broadcast news video browsing for home users. There are three main factors that distinguish our work from other similar ones. First, we have integrated the image and audio analysis results in identifying news segments. Second, we use the video OCR technology to detect text from frames, which provides a good source of textual information for story classification when transcripts and close captions are not available. Finally, natural language processing (NLP) technologies are used to perform automated categorization of news stories based on the texts obtained from close caption or video OCR process. Based on these video structure and content analysis technologies, we have developed two advanced video browsers for home users: intelligent highlight player and HTML-based video browser.
Lie Gu, Hao Jiang 0007, HongJiang Zhang
ICIP3