VLDB 2026 Research / reviewers in the wild / expert
Jia Xu 0011
dblp:95/3616-11
· DBLP profile ↗
22ranked-venue papers
6as first author
4since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
19 papers |
3D vision · 30% Segmentation and scene understanding · 11% Generative modeling · 10% | |
| Computer graphics and multimedia
7 papers |
Image and video processing · 45% Computer animation and physical simulation · 17% Visual content generation and editing · 13% |
Topics — the 30 heaviest of 50, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › motion estimation
optical flow |
1.5 | 4 | 2020 | Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo Matching · CVPR 2020 SelFlow: Self-Supervised Learning of Optical Flow · CVPR 2019 DDFlow: Learning Optical Flow with Unlabeled Data Distillation · AAAI 2019 |
Computer vision › 3D vision › motion estimation › optical flow
unsupervised optical flow |
0.8 | 2 | 2019 | SelFlow: Self-Supervised Learning of Optical Flow · CVPR 2019 DDFlow: Learning Optical Flow with Unlabeled Data Distillation · AAAI 2019 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.7 | 2 | 2020 | Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo Matching · CVPR 2020 Accurate Optical Flow via Direct Cost Volume Processing · CVPR 2017 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.7 | 2 | 2022 | Learning by Distillation: A Self-Supervised Learning Framework for Optical Flow Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 DDFlow: Learning Optical Flow with Unlabeled Data Distillation · AAAI 2019 |
Machine learning › Generative modeling › diffusion model
human motion generation |
0.7 | 1 | 2023 | Human MotionFormer: Transferring Human Motions with Vision Transformers · ICLR 2023 |
Machine learning › Generative modeling › video generation
motion transfer |
0.7 | 1 | 2023 | Human MotionFormer: Transferring Human Motions with Vision Transformers · ICLR 2023 |
Computer animation and physical simulation › motion synthesis
human motion synthesis |
0.7 | 1 | 2023 | Human MotionFormer: Transferring Human Motions with Vision Transformers · ICLR 2023 |
Image and video processing › motion estimation
optical flow |
0.6 | 1 | 2022 | Learning by Distillation: A Self-Supervised Learning Framework for Optical Flow Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › Video understanding and tracking
human motion prediction |
0.5 | 1 | 2021 | Action-guided 3D Human Motion Prediction · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › memory mechanism
memory bank |
0.5 | 1 | 2021 | Action-guided 3D Human Motion Prediction · NeurIPS 2021 |
Visual content generation and editing › image animation
human motion transfer |
0.5 | 1 | 2021 | Few-Shot Human Motion Transfer by Personalized Geometry and Texture Modeling · CVPR 2021 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.4 | 2 | 2015 | Learning to segment under various forms of weak supervision · CVPR 2015 Tell Me What You See and I Will Show You Where It Is · CVPR 2014 |
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation |
0.4 | 2 | 2015 | Learning to segment under various forms of weak supervision · CVPR 2015 Tell Me What You See and I Will Show You Where It Is · CVPR 2014 |
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning |
0.4 | 1 | 2019 | DHER: Hindsight Experience Replay for Dynamic Goals · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning › off-policy reinforcement learning › experience replay
hindsight experience replay |
0.4 | 1 | 2019 | DHER: Hindsight Experience Replay for Dynamic Goals · ICLR (Poster) 2019 |
Computer vision › 3D vision
motion estimation |
0.4 | 1 | 2019 | SelFlow: Self-Supervised Learning of Optical Flow · CVPR 2019 |
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency |
0.4 | 1 | 2019 | Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses · CVPR 2019 |
Computer vision › Vision and language
video grounding |
0.4 | 1 | 2019 | Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses · CVPR 2019 |
Computer vision › Vision and language › video grounding
weakly-supervised video grounding |
0.4 | 1 | 2019 | Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses · CVPR 2019 |
Image and video processing › image restoration
image denoising |
0.3 | 1 | 2018 | Learning to See in the Dark · CVPR 2018 |
Image and video processing › image enhancement
low-light image processing |
0.3 | 1 | 2018 | Learning to See in the Dark · CVPR 2018 |
Computer vision › 3D vision › stereo vision › stereo matching
semi-global matching |
0.3 | 1 | 2017 | Accurate Optical Flow via Direct Cost Volume Processing · CVPR 2017 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
0.2 | 1 | 2015 | Manifold-valued Dirichlet Processes · ICML 2015 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
bayesian regression |
0.2 | 1 | 2015 | Manifold-valued Dirichlet Processes · ICML 2015 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model |
0.2 | 1 | 2015 | Manifold-valued Dirichlet Processes · ICML 2015 |
Computer vision › 3D vision › geometric deep learning
manifold-valued data |
0.2 | 1 | 2015 | Manifold-valued Dirichlet Processes · ICML 2015 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
riemannian manifold |
0.2 | 1 | 2015 | Manifold-valued Dirichlet Processes · ICML 2015 |
Multimedia analysis and retrieval › video summarization
egocentric video summarization |
0.2 | 1 | 2015 | Gaze-enabled egocentric video summarization via constrained submodular maximization · CVPR 2015 |
Multimedia analysis and retrieval
video summarization |
0.2 | 1 | 2015 | Gaze-enabled egocentric video summarization via constrained submodular maximization · CVPR 2015 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
subspace learning |
0.2 | 2 | 2013 | GOSUS: Grassmannian Online Subspace Updates with Structured-Sparsity · ICCV 2013 Analyzing the Subspace Structure of Related Images: Concurrent Segmentation of Image Sets · ECCV (4) 2012 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 2.0teacher-student training · 1.5knowledge distillation · 1.5vision transformer · 1.3motion transfer · 1.3fully convolutional network · 1.2texture generator · 0.5query-read process · 0.5keypoint conditioning · 0.5action constraint loss · 0.5UV map · 0.5proxy task · 0.4data distillation · 0.4end-to-end training · 0.3spectral clustering · 0.2convex regularization · 0.2random walker · 0.1GPU computing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Human MotionFormer: Transferring Human Motions with Vision Transformers
Xintong Han, Chenbin Jin, Lihui Qian 0003, Huawei Wei, Haoye Dong, Yibing Song, Jia Xu 0011, Qifeng Chen 0001 |
ICLR | 10 |
| 2022 | Learning by Distillation: A Self-Supervised Learning Framework for Optical Flow EstimationabstractWe present DistillFlow, a knowledge distillation approach to learning optical flow. DistillFlow trains multiple teacher models and a student model, where challenging transformations are applied to the input of the student model to generate hallucinated occlusions as well as less confident predictions. Then, a self-supervised learning framework is constructed: confident predictions from teacher models are served as annotations to guide the student model to learn optical flow for those less confident predictions. The self-supervised learning framework enables us to effectively learn optical flow from unlabeled data, not only for non-occluded pixels, but also for occluded pixels. DistillFlow achieves state-of-the-art unsupervised learning performance on both KITTI and Sintel datasets. Our self-supervised pre-trained model also provides an excellent initialization for supervised fine-tuning, suggesting an alternate training paradigm in contrast to current supervised learning methods that highly rely on pre-training on synthetic data. At the time of writing, our fine-tuned models ranked 1st among all monocular methods on the KITTI 2015 benchmark, and outperform all published methods on the Sintel Final benchmark. More importantly, we demonstrate the generalization capability of DistillFlow in three aspects: framework generalization, correspondence generalization and cross-dataset generalization. Our code and models will be available on https://github.com/ppliuboy/DistillFlow. Michael R. Lyu, Irwin King, Jia Xu 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Few-Shot Human Motion Transfer by Personalized Geometry and Texture ModelingabstractWe present a new method for few-shot human motion transfer that achieves realistic human image generation with only a small number of appearance inputs. Despite recent advances in single person motion transfer, prior methods often require a large number of training images and take long training time. One promising direction is to perform few-shot human motion transfer, which only needs a few of source images for appearance transfer. However, it is particularly challenging to obtain satisfactory transfer results. In this paper, we address this issue by rendering a human texture map to a surface geometry (represented as a UV map), which is personalized to the source person. Our geometry generator combines the shape information from source images, and the pose information from 2D keypoints to synthesize the personalized UV map. A texture generator then generates the texture map conditioned on the texture of source images to fill out invisible parts. Furthermore, we may fine-tune the texture map on the manifold of the texture generator from a few source images at the test time, which improves the quality of the texture map without over-fitting or artifacts. Extensive experiments show the proposed method outperforms state-of-the-art methods both qualitatively and quantitatively. Our code is available at https://github.com/HuangZhiChao95/FewShotMotionTransfer. Zhichao Huang 0002, Xintong Han, Jia Xu 0011, Tong Zhang 0001 |
CVPR | 3 |
| 2021 | Action-guided 3D Human Motion PredictionabstractThe ability of forecasting future human motion is important for human-machine interaction systems to understand human behaviors and make interaction. In this work, we focus on developing models to predict future human motion from past observed video frames. Motivated by the observation that human motion is closely related to the action being performed, we propose to explore action context to guide motion prediction. Specifically, we construct an action-specific memory bank to store representative motion dynamics for each action category, and design a query-read process to retrieve some motion dynamics from the memory bank. The retrieved dynamics are consistent with the action depicted in the observed video frames and serve as a strong prior knowledge to guide motion prediction. We further formulate an action constraint loss to ensure the global semantic consistency of the predicted motion. Extensive experiments demonstrate the effectiveness of the proposed approach, and we achieve state-of-the-art performance on 3D human motion prediction. Jiangxin Sun, Zihang Lin, Xintong Han, Jianfang Hu, Jia Xu 0011, Wei-Shi Zheng 0001 |
NeurIPS | 5 |
| 2020 | Learning 3D Face Reconstruction with a Pose Guidance Network
Xintong Han, Michael R. Lyu, Irwin King, Jia Xu 0011 |
ACCV (5) | 5 |
| 2020 | Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo MatchingabstractIn this paper, we propose a unified method to jointly learn optical flow and stereo matching. Our first intuition is stereo matching can be modeled as a special case of optical flow, and we can leverage 3D geometry behind stereoscopic videos to guide the learning of these two forms of correspondences. We then enroll this knowledge into the state-of-the-art self-supervised learning framework, and train one single network to estimate both flow and stereo. Second, we unveil the bottlenecks in prior self-supervised learning approaches, and propose to create a new set of challenging proxy tasks to boost performance. These two insights yield a single model that achieves the highest accuracy among all existing unsupervised flow and stereo methods on KITTI 2012 and 2015 benchmarks. More remarkably, our self-supervised method even outperforms several state-of-the-art fully supervised methods, including PWC-Net and FlowNet2 on KITTI 2012. Irwin King, Michael R. Lyu, Jia Xu 0011 |
CVPR | 4 |
| 2019 | DDFlow: Learning Optical Flow with Unlabeled Data DistillationabstractWe present DDFlow, a data distillation approach to learning optical flow estimation from unlabeled data. The approach distills reliable predictions from a teacher network, and uses these predictions as annotations to guide a student network to learn optical flow. Unlike existing work relying on handcrafted energy terms to handle occlusion, our approach is data-driven, and learns optical flow for occluded pixels. This enables us to train our model with a much simpler loss function, and achieve a much higher accuracy. We conduct a rigorous evaluation on the challenging Flying Chairs, MPI Sintel, KITTI 2012 and 2015 benchmarks, and show that our approach significantly outperforms all existing unsupervised learning methods, while running at real time. Irwin King, Michael R. Lyu, Jia Xu 0011 |
AAAI | 4 |
| 2019 | Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering LossesabstractWe invest the problem of weakly-supervised video grounding, where only video-level sentences are provided. This is a challenging task, and previous Multi-Instance Learning (MIL) based image grounding methods turn to fail in the video domain. Recent work attempts to decompose the video-level MIL into frame-level MIL by applying weighted sentence-frame ranking loss over frames, but it is not robust and does not exploit the rich temporal information in videos. In this work, we address these issues by extending frame-level MIL with a false positive frame-bag constraint and modeling the visual feature consistency in the video. In specific, we design a contextual similarity between semantic and visual features to deal with sparse objects association across frames. Furthermore, we leverage temporal coherence by strengthening the clustering effect of similar features in the visual space. We conduct an extensive evaluation on YouCookII and RoboWatch datasets, and demonstrate our method significantly outperforms prior state-of-the-art methods. Jing Shi 0005, Jia Xu 0011, Boqing Gong, Chenliang Xu |
CVPR | 2 |
| 2019 | SelFlow: Self-Supervised Learning of Optical FlowabstractWe present a self-supervised learning approach for optical flow. Our method distills reliable flow estimations from non-occluded pixels, and uses these predictions as ground truth to learn optical flow for hallucinated occlusions. We further design a simple CNN to utilize temporal information from multiple frames for better flow estimation. These two principles lead to an approach that yields the best performance for unsupervised optical flow learning on the challenging benchmarks including MPI Sintel, KITTI 2012 and 2015. More notably, our self-supervised pre-trained model provides an excellent initialization for supervised fine-tuning. Our fine-tuned models achieve state-of-the-art results on all three datasets. At the time of writing, we achieve EPE=4.26 on the Sintel benchmark, outperforming all submitted methods. Michael R. Lyu, Irwin King, Jia Xu 0011 |
CVPR | 4 |
| 2019 | DHER: Hindsight Experience Replay for Dynamic Goals
Bei Shi, Boqing Gong, Jia Xu 0011, Tong Zhang 0001 |
ICLR (Poster) | 5 |
| 2018 | Learning to See in the DarkabstractImaging in low light is challenging due to low photon count and low SNR. Short-exposure images suffer from noise, while long exposure can induce blur and is often impractical. A variety of denoising, deblurring, and enhancement techniques have been proposed, but their effectiveness is limited in extreme conditions, such as video-rate imaging at night. To support the development of learning-based pipelines for low-light image processing, we introduce a dataset of raw short-exposure low-light images, with corresponding long-exposure reference images. Using the presented dataset, we develop a pipeline for processing low-light images, based on end-to-end training of a fully-convolutional network. The network operates directly on raw sensor data and replaces much of the traditional image processing pipeline, which tends to perform poorly on such data. We report promising results on the new dataset, analyze factors that affect performance, and highlight opportunities for future work. Chen Chen 0003, Qifeng Chen 0001, Jia Xu 0011, Vladlen Koltun |
CVPR | 3 |
| 2017 | Accurate Optical Flow via Direct Cost Volume ProcessingabstractWe present an optical flow estimation approach that operates on the full four-dimensional cost volume. This direct approach shares the structural benefits of leading stereo matching pipelines, which are known to yield high accuracy. To this day, such approaches have been considered impractical due to the size of the cost volume. We show that the full four-dimensional cost volume can be constructed in a fraction of a second due to its regularity. We then exploit this regularity further by adapting semi-global matching to the four-dimensional setting. This yields a pipeline that achieves significantly higher accuracy than state-of-the-art optical flow methods while being faster than most. Our approach outperforms all published general-purpose optical flow methods on both Sintel and KITTI 2015 benchmarks. Jia Xu 0011, René Ranftl, Vladlen Koltun |
CVPR | 1 |
| 2017 | Fast Image Processing with Fully-Convolutional NetworksabstractWe present an approach to accelerating a wide variety of image processing operators. Our approach uses a fully-convolutional network that is trained on input-output pairs that demonstrate the operator's action. After training, the original operator need not be run at all. The trained network operates at full resolution and runs in constant time. We investigate the effect of network architecture on approximation accuracy, runtime, and memory footprint, and identify a specific architecture that balances these considerations. We evaluate the presented approach on ten advanced image processing operators, including multiple variational models, multiscale tone and detail manipulation, photographic style transfer, nonlocal dehazing, and nonphoto-realistic stylization. All operators are approximated by the same model. Experiments demonstrate that the presented approach is significantly more accurate than prior approximation schemes. It increases approximation accuracy as measured by PSNR across the evaluated operators by 8.5 dB on the MIT-Adobe dataset (from 27.5 to 36 dB) and reduces DSSIM by a multiplicative factor of 3 compared to the most accurate prior approximation scheme, while being the fastest. We show that our models generalize across datasets and across resolutions, and investigate a number of extensions of the presented approach. Qifeng Chen 0001, Jia Xu 0011, Vladlen Koltun |
ICCV | 2 |
| 2015 | Gaze-enabled egocentric video summarization via constrained submodular maximizationabstractWith the proliferation of wearable cameras, the number of videos of users documenting their personal lives using such devices is rapidly increasing. Since such videos may span hours, there is an important need for mechanisms that represent the information content in a compact form (i.e., shorter videos which are more easily browsable/sharable). Motivated by these applications, this paper focuses on the problem of egocentric video summarization. Such videos are usually continuous with significant camera shake and other quality issues. Because of these reasons, there is growing consensus that direct application of standard video summarization tools to such data yields unsatisfactory performance. In this paper, we demonstrate that using gaze tracking information (such as fixation and saccade) significantly helps the summarization task. It allows meaningful comparison of different image frames and enables deriving personalized summaries (gaze provides a sense of the camera wearer's intent). We formulate a summarization model which captures common-sense properties of a good summary, and show that it can be solved as a submodular function maximization with partition matroid constraints, opening the door to a rich body of work from combinatorial optimization. We evaluate our approach on a new gaze-enabled egocentric video dataset (over 15 hours), which will be a valuable standalone resource. Jia Xu 0011, Lopamudra Mukherjee, Yin Li 0003, Jamieson Warner, James M. Rehg |
CVPR | 1 |
| 2015 | Learning to segment under various forms of weak supervisionabstractDespite the promising performance of conventional fully supervised algorithms, semantic segmentation has remained an important, yet challenging task. Due to the limited availability of complete annotations, it is of great interest to design solutions for semantic segmentation that take into account weakly labeled data, which is readily available at a much larger scale. Contrasting the common theme to develop a different algorithm for each type of weak annotation, in this work, we propose a unified approach that incorporates various forms of weak supervision - image level tags, bounding boxes, and partial labels - to produce a pixel-wise labeling. We conduct a rigorous evaluation on the challenging Siftflow dataset for various weakly labeled settings, and show that our approach outperforms the state-of-the-art by 12% on per-class accuracy, while maintaining comparable per-pixel accuracy. Jia Xu 0011, Alexander G. Schwing, Raquel Urtasun |
CVPR | 1 |
| 2015 | Manifold-valued Dirichlet ProcessesabstractStatistical models for manifold-valued data permit capturing the intrinsic nature of the curved spaces in which the data lie and have been a topic of research for several decades. Typically, these formulations use geodesic curves and distances defined locally for most cases - this makes it hard to design parametric models globally on smooth manifolds. Thus, most (manifold specific) parametric models available today assume that the data lie in a small neighborhood on the manifold. To address this ’locality’ problem, we propose a novel nonparametric model which unifies multivariate general linear models (MGLMs) using multiple tangent spaces. Our framework generalizes existing work on (both Euclidean and non-Euclidean) general linear models providing a recipe to globally extend the locally-defined parametric models (using a mixture of local models). By grouping observations into sub-populations at multiple tangent spaces, our method provides insights into the hidden structure (geodesic relationships) in the data. This yields a framework to group observations and discover geodesic relationships between covariates X and manifold-valued responses Y, which we call Dirichlet process mixtures of multivariate general linear models (DP-MGLM) on Riemannian manifolds. Finally, we present proof of concept experiments to validate our model. Hyunwoo J. Kim, Jia Xu 0011, Baba C. Vemuri |
ICML | 2 |
| 2014 | Tell Me What You See and I Will Show You Where It IsabstractWe tackle the problem of weakly labeled semantic segmentation, where the only source of annotation are image tags encoding which classes are present in the scene. This is an extremely difficult problem as no pixel-wise labelings are available, not even at training time. In this paper, we show that this problem can be formalized as an instance of learning in a latent structured prediction framework, where the graphical model encodes the presence and absence of a class as well as the assignments of semantic labels to superpixels. As a consequence, we are able to leverage standard algorithms with good theoretical properties. We demonstrate the effectiveness of our approach using the challenging SIFT-flow dataset and show average per-class accuracy improvements of 7% over the state-of-the-art. Jia Xu 0011, Alexander G. Schwing, Raquel Urtasun |
CVPR | 1 |
| 2014 | Spectral Clustering with a Convex Regularizer on Millions of Images
Maxwell D. Collins, Ji Liu 0002, Jia Xu 0011, Lopamudra Mukherjee |
ECCV (3) | 3 |
| 2013 | Incorporating User Interaction and Topological Constraints within Contour Completion via Discrete CalculusabstractWe study the problem of interactive segmentation and contour completion for multiple objects. The form of constraints our model incorporates are those coming from user scribbles (interior or exterior constraints) as well as information regarding the topology of the 2-D space after partitioning (number of closed contours desired). We discuss how concepts from discrete calculus and a simple identity using the Euler characteristic of a planar graph can be utilized to derive a practical algorithm for this problem. We also present specialized branch and bound methods for the case of single contour completion under such constraints. On an extensive dataset of ~ 1000 images, our experiments suggest that a small amount of side knowledge can give strong improvements over fully unsupervised contour completion methods. We show that by interpreting user indications topologically, user effort is substantially reduced. Jia Xu 0011, Maxwell D. Collins |
CVPR | 1 |
| 2013 | GOSUS: Grassmannian Online Subspace Updates with Structured-SparsityabstractWe study the problem of online subspace learning in the context of sequential observations involving structured perturbations. In online subspace learning, the observations are an unknown mixture of two components presented to the model sequentially - the main effect which pertains to the subspace and a residual/error term. If no additional requirement is imposed on the residual, it often corresponds to noise terms in the signal which were unaccounted for by the main effect. To remedy this, one may impose "structural" contiguity, which has the intended effect of leveraging the secondary terms as a covariate that helps the estimation of the subspace itself, instead of merely serving as a noise residual. We show that the corresponding online estimation procedure can be written as an approximate optimization process on a Grassmannian. We propose an efficient numerical solution, GOSUS, Grassmannian Online Subspace Updates with Structured-sparsity, for this problem. GOSUS is expressive enough in modeling both homogeneous perturbations of the subspace and structural contiguities of outliers, and after certain manipulations, solvable via an alternating direction method of multipliers (ADMM). We evaluate the empirical performance of this algorithm on two problems of interest: online background subtraction and online multiple face tracking, and demonstrate that it achieves competitive performance with the state-of-the-art in near real time. Jia Xu 0011, Vamsi K. Ithapu, Lopamudra Mukherjee, James M. Rehg |
ICCV | 1 |
| 2012 | Random walks based multi-image segmentation: Quasiconvexity results and GPU-based solutionsabstractproblem using Random Walker (RW) segmentation as the core segmentation algorithm, rather than the traditional MRF approach adopted in the literature so far. Our formulation is similar to previous approaches in the sense that it also permits Cosegmentation constraints (which impose consistency between the extracted objects from ≥ 2 images) using a nonparametric model. However, several previous nonparametric cosegmentation methods have the serious limitation that they require adding one auxiliary node (or variable) for every pair of pixels that are similar (which effectively limits such methods to describing only those objects that have high entropy appearance models). In contrast, our proposed model completely eliminates this restrictive dependence -the resulting improvements are quite significant. Our model further allows an optimization scheme exploiting quasiconvexity for model-based segmentation with no dependence on the scale of the segmented foreground. Finally, we show that the optimization can be expressed in terms of linear algebra operations on sparse matrices which are easily mapped to GPU architecture. We provide a highly specialized CUDA library for Cosegmentation exploiting this special structure, and report experimental results showing these advantages. Maxwell D. Collins, Jia Xu 0011, Leo J. Grady |
CVPR | 2 |
| 2012 | Analyzing the Subspace Structure of Related Images: Concurrent Segmentation of Image Sets
Lopamudra Mukherjee, Jia Xu 0011, Maxwell D. Collins |
ECCV (4) | 3 |