VLDB 2026 Research / reviewers in the wild / expert
Jens Rittscher
dblp:11/377
· DBLP profile ↗
38ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-8528-8298ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 12 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Supervised Monocular Depth and Pose Estimation for Endoscopy With Latent PriorsabstractAccurate 3D reconstruction in endoscopy enables quantitative and holistic lesion characterization within the gastrointestinal (GI) tract. To achieve this, reliable depth and pose estimation is required. However, endoscopy systems are monocular, and existing methods relying on synthetic datasets or complex models often lack generalizability in challenging endoscopic conditions. We propose a robust self-supervised monocular depth and pose estimation framework that incorporates a StyleGAN-based generator and a Variational Autoencoder (VAE). The StyleGAN generator leverages extensive depth scenes from natural images to condition the depth network, enhancing realism and robustness of depth predictions through latent feature priors. For pose estimation, we reformulate it within a VAE framework, treating pose transitions as latent variables to regularize scale, stabilize z-axis prominence, and improve x-y sensitivity. To further enhance pose stability and generalizability, we introduce a prior transfer module that distills motion knowledge from natural scene SLAM systems. Specifically, pose priors from a pretrained SLAM model-supervised on large-scale natural scene datasets-are used to guide the latent distribution of pose through a KL-divergence reparameterization. This mechanism effectively transfers structural motion priors into the endoscopic domain, improving trajectory consistency under challenging conditions. This dual refinement pipeline enables accurate depth and pose predictions, effectively addressing the GI tract's complex textures and lighting. Extensive evaluations on SimCol, C3VD, and EndoSLAM datasets confirm our framework's superior performance over published self-supervised methods in endoscopic depth and pose estimation. All data descriptions and code are available at https://github.com/EricXuziang/Self-supervised-with-Latent-Priors.git. Ziang Xu 0001, James East, Sharib Ali, Jens Rittscher |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Unsupervised Discovery of Spatiotypes and Context-Aware Graph Neural Networks for Modeling Clinical Endpoints
Muhammad Dawood, Emily Thomas, Rosalin Cooper, Carlo Pescia, Anna Sozanska, Hosuk Ryou, Daniel Royston, Jens Rittscher |
MICCAI (11) | 8 |
| 2025 | Self-interactive learning: Fusion and evolution of multi-scale histomorphology features for molecular traits prediction in computational pathologyabstractPredicting disease-related molecular traits from histomorphology brings great opportunities for precision medicine. Despite the rich information present in histopathological images, extracting fine-grained molecular features from standard whole slide images (WSI) is non-trivial. The task is further complicated by the lack of annotations for subtyping and contextual histomorphological features that might span multiple scales. This work proposes a novel multiple-instance learning (MIL) framework capable of WSI-based cancer morpho-molecular subtyping by fusion of different-scale features. Our method, debuting as Inter-MIL, follows a weakly-supervised scheme. It enables the training of the patch-level encoder for WSI in a task-aware optimisation procedure, a step normally not modelled in most existing MIL-based WSI analysis frameworks. We demonstrate that optimising the patch-level encoder is crucial to achieving high-quality fine-grained and tissue-level subtyping results and offers a significant improvement over task-agnostic encoders. Our approach deploys a pseudo-label propagation strategy to update the patch encoder iteratively, allowing discriminative subtype features to be learned. This mechanism also empowers extracting fine-grained attention within image tiles (the small patches), a task largely ignored in most existing weakly supervised-based frameworks. With Inter-MIL, we carried out four challenging cancer molecular subtyping tasks in the context of ovarian, colorectal, lung, and breast cancer. Extensive evaluation results show that Inter-MIL is a robust framework for cancer morpho-molecular subtyping with superior performance compared to several recently proposed methods, in small dataset scenarios where the number of available training slides is less than 100. The iterative optimisation mechanism of Inter-MIL significantly improves the quality of the image features learned by the patch embedded and generally directs the attention map to areas that better align with experts' interpretation, leading to the identification of more reliable histopathology biomarkers. Moreover, an external validation cohort is used to verify the robustness of Inter-MIL on molecular trait prediction. Korsuk Sirinukunwattana, Bin Li 0064, Kezia Gaitskell, Enric Domingo, Willem Bonnaffé, Marta Wojciechowska, Ruby Wood, Nasullah Khalid Alham, Stefano Malacrino, Dan J. Woodcock, Clare Verrill, Jens Rittscher |
Medical Image Anal. | 14 |
| 2025 | Editorial for Special Issue on Foundation Models for Medical Image Analysis
Xiaosong Wang 0001, Dequan Wang, Jens Rittscher, Dimitris N. Metaxas, Shaoting Zhang 0001 |
Medical Image Anal. | 4 |
| 2024 | MultiVarNet - Predicting Tumour Mutational Status at the Protein Level
Louis-Oscar Morel, Muhammad Muzammel, Nathan Vinçon, Valentin Derangère, Sylvain Ladoire, Jens Rittscher |
MICCAI (3) | 6 |
| 2024 | SSL-CPCD: Self-Supervised Learning With Composite Pretext-Class Discrimination for Improved Generalisability in Endoscopic Image AnalysisabstractData-driven methods have shown tremendous progress in medical image analysis. In this context, deep learning-based supervised methods are widely popular. However, they require a large amount of training data and face issues in generalisability to unseen datasets that hinder clinical translation. Endoscopic imaging data is characterised by large inter- and intra-patient variability that makes these models more challenging to learn representative features for downstream tasks. Thus, despite the publicly available datasets and datasets that can be generated within hospitals, most supervised models still underperform. While self-supervised learning has addressed this problem to some extent in natural scene data, there is a considerable performance gap in the medical image domain. In this paper, we propose to explore patch-level instance-group discrimination and penalisation of inter-class variation using additive angular margin within the cosine similarity metrics. Our novel approach enables models to learn to cluster similar representations, thereby improving their ability to provide better separation between different classes. Our results demonstrate significant improvement on all metrics over the state-of-the-art (SOTA) methods on the test set from the same and diverse datasets. We evaluated our approach for classification, detection, and segmentation. SSL-CPCD attains notable Top 1 accuracy of 79.77% in ulcerative colitis classification, an 88.62% mean average precision (mAP) for detection, and an 82.32% dice similarity coefficient for polyp segmentation tasks. These represent improvements of over 4%, 2%, and 3%, respectively, compared to the baseline architectures. We demonstrate that our method generalises better than all SOTA methods to unseen datasets, reporting over 7% improvement. Ziang Xu 0001, Jens Rittscher, Sharib Ali |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Joint Prediction of Response to Therapy, Molecular Traits, and Spatial Organisation in Colorectal Cancer Biopsies
Ruby Wood, Enric Domingo, Korsuk Sirinukunwattana, Maxime W. Lafarge, Viktor H. Koelzer, Timothy S. Maughan, Jens Rittscher |
MICCAI (5) | 7 |
| 2023 | FANet: A Feedback Attention Network for Improved Biomedical Image SegmentationabstractThe increase of available large clinical and experimental datasets has contributed to a substantial amount of important contributions in the area of biomedical image analysis. Image segmentation, which is crucial for any quantitative analysis, has especially attracted attention. Recent hardware advancement has led to the success of deep learning approaches. However, although deep learning models are being trained on large datasets, existing methods do not use the information from different learning epochs effectively. In this work, we leverage the information of each training epoch to prune the prediction maps of the subsequent epochs. We propose a novel architecture called feedback attention network (FANet) that unifies the previous epoch mask with the feature map of the current training epoch. The previous epoch mask is then used to provide hard attention to the learned feature maps at different convolutional layers. The network also allows rectifying the predictions in an iterative fashion during the test time. We show that our proposed feedback attention model provides a substantial improvement on most segmentation metrics tested on seven publicly available biomedical imaging datasets demonstrating the effectiveness of FANet. The source code is available at https://github.com/nikhilroxtomar/FANet. Nikhil Kumar Tomar, Debesh Jha, Michael Riegler 0001, Håvard D. Johansen, Dag Johansen, Jens Rittscher, Pål Halvorsen, Sharib Ali |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Predicting Molecular Traits from Tissue Morphology Through Self-interactive Multi-instance Learning
Korsuk Sirinukunwattana, Kezia Gaitskell, Ruby Wood, Clare Verrill, Jens Rittscher |
MICCAI (2) | 6 |
| 2021 | EndoUDA: A Modality Independent Segmentation Approach for Endoscopy Imaging
Numan Celik, Sharib Ali, Soumya Gupta 0001, Barbara Braden, Jens Rittscher |
MICCAI (3) | 5 |
| 2021 | Early Detection of Liver Fibrosis Using Graph Convolutional Networks
Marta Wojciechowska, Stefano Malacrino, Natalia Garcia Martin, Hamid Fehri, Jens Rittscher |
MICCAI (8) | 5 |
| 2021 | Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopyabstractThe Endoscopy Computer Vision Challenge (EndoCV) is a crowd-sourcing initiative to address eminent problems in developing reliable computer aided detection and diagnosis endoscopy systems and suggest a pathway for clinical translation of technologies. Whilst endoscopy is a widely used diagnostic and treatment tool for hollow-organs, there are several core challenges often faced by endoscopists, mainly: 1) presence of multi-class artefacts that hinder their visual interpretation, and 2) difficulty in identifying subtle precancerous precursors and cancer abnormalities. Artefacts often affect the robustness of deep learning methods applied to the gastrointestinal tract organs as they can be confused with tissue of interest. EndoCV2020 challenges are designed to address research questions in these remits. In this paper, we present a summary of methods developed by the top 17 teams and provide an objective comparison of state-of-the-art methods and methods designed by the participants for two sub-challenges: i) artefact detection and segmentation (EAD2020), and ii) disease detection and segmentation (EDD2020). Multi-center, multi-organ, multi-class, and multi-modal clinical endoscopy datasets were compiled for both EAD2020 and EDD2020 sub-challenges. The out-of-sample generalization ability of detection algorithms was also evaluated. Whilst most teams focused on accuracy improvements, only a few methods hold credibility for clinical usability. The best performing teams provided solutions to tackle class imbalance, and variabilities in size, origin, modality and occurrences by exploring data augmentation, data fusion, and optimal class thresholding techniques. Sharib Ali, Mariia Dmitrieva, Noha M. Ghatwary, Sophia Bano, Gorkem Polat, Alptekin Temizel, Adrian Krenzer, Amar Hekalo, Bogdan J. Matuszewski, Mourad Gridach, Irina Voiculescu, Vishnusai Yoganand, Arnav Chavan, Aryan Raj, Nhan T. Nguyen, Dat Q. Tran, Lê Duy Huynh, Nicolas Boutry, Shahadate Rezvy, Haijian Chen, Yoon Ho Choi, Anand Subramanian 0004, Velmurugan Balasubramanian, Xiaohong W. Gao, Hongyu Hu, Yusheng Liao, Danail Stoyanov, Christian Daul, Stefano Realdon, Renato Cannizzaro, Dominique Lamarque, Terry Tran-Nguyen, Adam Bailey, Barbara Braden, James E. East, Jens Rittscher |
Medical Image Anal. | 37 |
| 2021 | A deep learning framework for quality assessment and restoration in video endoscopyabstractEndoscopy is a routine imaging technique used for both diagnosis and minimally invasive surgical treatment. Artifacts such as motion blur, bubbles, specular reflections, floating objects and pixel saturation impede the visual interpretation and the automated analysis of endoscopy videos. Given the widespread use of endoscopy in different clinical applications, robust and reliable identification of such artifacts and the automated restoration of corrupted video frames is a fundamental medical imaging problem. Existing state-of-the-art methods only deal with the detection and restoration of selected artifacts. However, typically endoscopy videos contain numerous artifacts which motivates to establish a comprehensive solution. In this paper, a fully automatic framework is proposed that can: 1) detect and classify six different artifacts, 2) segment artifact instances that have indefinable shapes, 3) provide a quality score for each frame, and 4) restore partially corrupted frames. To detect and classify different artifacts, the proposed framework exploits fast, multi-scale and single stage convolution neural network detector. In addition, we use an encoder-decoder model for pixel-wise segmentation of irregular shaped artifacts. A quality score is introduced to assess video frame quality and to predict image restoration success. Generative adversarial networks with carefully chosen regularization and training strategies for discriminator-generator networks are finally used to restore corrupted frames. The detector yields the highest mean average precision (mAP) of 45.7 and 34.7, respectively for 25% and 50% IoU thresholds, and the lowest computational time of 88 ms allowing for near real-time processing. The restoration models for blind deblurring, saturation correction and inpainting demonstrate significant improvements over previous methods. On a set of 10 test videos, an average of 68.7% of video frames successfully passed the quality score (≥0.9) after applying the proposed restoration framework thereby retaining 25% more frames compared to the raw videos. The importance of artifacts detection and their restoration on improved robustness of image analysis methods is also demonstrated in this work. Sharib Ali, Felix Zhou 0001, Adam Bailey, Barbara Braden, James E. East, Jens Rittscher |
Medical Image Anal. | 7 |
| 2020 | Microscopic Fine-Grained Instance Classification Through Deep Attention
Mengran Fan, Tapabrata Chakraborti, Eric I-Chao Chang, Yan Xu 0001, Jens Rittscher |
MICCAI (5) | 5 |
| 2018 | Improving Whole Slide Segmentation Through Visual Context - A Systematic Study
Korsuk Sirinukunwattana, Nasullah Khalid Alham, Clare Verrill, Jens Rittscher |
MICCAI (2) | 4 |
| 2011 | LPSM: Fitting shape model by linear programmingabstractWe propose a shape model fitting algorithm that uses linear programming optimization. Most shape model fitting approaches (such as ASM, AAM) are based on gradient-descent-like local search optimization and usually suffer from local minima. In contrast, linear programming (LP) techniques achieve globally optimal solution for linear problems. In [1], a linear programming scheme based on successive convexification was proposed for matching static object shape in images among cluttered background and achieved very good performance. In this paper, we rigorously derive the linear formulation of the shape model fitting problem in the LP scheme and propose an LP shape model fitting algorithm (LPSM). In the experiments, we compared the performance of our LPSM with the LP graph matching algorithm(LPGM), ASM, and a CONDENSATION based ASM algorithm on a test set of PUT database. The experiments show that LPSM can achieve higher shape fitting accuracy. We also evaluated its performance on the fitting of some real world face images collected from internet. The results show that LPSM can handle various appearance outliers and can avoid local minima problem very well, as the fitting is carried out by LP optimization with l1norm robust cost function. Jilin Tu, Brandon Laflen, Xiaoming Liu 0002, Musodiq O. Bello, Jens Rittscher, Peter H. Tu |
FG | 5 |
| 2011 | Non-parametric Population Analysis of Cellular Phenotypes
Shantanu Singh, Firdaus Janoos, Thierry Pécot, Enrico Caserta, Kun Huang 0001, Jens Rittscher, Gustavo Leone, Raghu Machiraju |
MICCAI (2) | 6 |
| 2011 | Coupled minimum-cost flow cell tracking for high-throughput quantitative analysis
Dirk Ryan Padfield, Jens Rittscher, Badrinath Roysam |
Medical Image Anal. | 2 |
| 2010 | Automated Training Data Generation for Microscopy Focus Classification
Dirk Ryan Padfield, Jens Rittscher, Richard McKay |
MICCAI (2) | 3 |
| 2009 | A model change detection approach to dynamic scene modelingabstractIn this work we propose a dynamic scene model to provide information about the presence of salient motion in the scene, and that could be used for focusing the attention of a pan/tilt/zoom camera, or for background modeling purposes. Rather than proposing a set of saliency detectors, we define what we mean by salient motion, and propose a precise model for it. Detecting salient motion becomes equivalent to detecting a model change. We derive optimal online procedures to solve this problem, which enable a very fast implementation. Promising results show that our model can effectively detect salient motion even in severely cluttered scenes, and while a camera is panning and tilting. Seon Joo Kim, Gianfranco Doretto, Jens Rittscher, Peter H. Tu, Nils Krahnstoever, Marc Pollefeys |
AVSS | 3 |
| 2009 | Spatio-temporal cell cycle phase analysis using level sets and fast marching methods
Dirk Ryan Padfield, Jens Rittscher, Nick Thomas, Badrinath Roysam |
Medical Image Anal. | 2 |
| 2008 | Unified Crowd Segmentation
Peter H. Tu, Thomas Sebastian, Gianfranco Doretto, Nils Krahnstoever, Jens Rittscher, Ting Yu 0003 |
ECCV (4) | 5 |
| 2007 | View adaptive detection and distributed site wide trackingabstractUsing a detect and track paradigm, we present a surveillance framework where each camera uses local resources to perform real-time person detection. These detections are then processed by a distributed site-wide tracking system. The detectors themselves are based on boosted user-defined exemplars, which capture both appearance and shape information. The detectors take integral images of both intensity and Sobel responses as input. This data representation enables efficient processing without relying on background subtraction or other motion cues. View-specific person detectors are constructed by iteratively presenting the boosting algorithm with training data associated with each individual camera. These detections are then transmitted from a distributed set of tracking clients to a server, which maintains a set of site-wide target tracks. Automatic calibration methods allow for tracking to be performed in a ground plane representation, which enables effective camera hand-off. Factors such as network latencies and scalability will be discussed. Peter H. Tu, Nils Krahnstoever, Jens Rittscher |
AVSS | 3 |
| 2007 | Shape and Appearance Context ModelingabstractIn this work we develop appearance models for computing the similarity between image regions containing deformable objects of a given class in realtime. We introduce the concept of shape and appearance context. The main idea is to model the spatial distribution of the appearance relative to each of the object parts. Estimating the model entails computing occurrence matrices. We introduce a generalization of the integral image and integral histogram frameworks, and prove that it can be used to dramatically speed up occurrence computation. We demonstrate the ability of this framework to recognize an individual walking across a network of cameras. Finally, we show that the proposed approach outperforms several other methods. Gianfranco Doretto, Thomas Sebastian, Jens Rittscher, Peter H. Tu |
ICCV | 4 |
| 2006 | Optimal Pose for Face RecognitionabstractResearchers in psychology have well studied the impact of the pose of a face as perceived by humans, and concluded that the so-called 3/4 view, halfway between the front view and the profile view, is the easiest for face recognition by humans. For face recognition by machines, while much work has been done to create recognition algorithms that are robust to pose variation, little has been done in finding the most representative pose for recognition. In this paper, we use a number of algorithms to evaluate face recognition performance when various poses are used for training. The result, similar to findings in psychology that the 3/4 view is the best, is also justified by the discrimination power of different regions on the face, computed from both the appearance and the geometry of these regions. We believe our study is both scientifically interesting and practically beneficial for many applications. Xiaoming Liu 0002, Jens Rittscher, Tsuhan Chen |
CVPR (2) | 2 |
| 2005 | Detecting and counting people in surveillance applicationsabstractA number of surveillance scenarios require the detection and tracking of people. Although person detection and counting systems are commercially available today, there is need for further research to address the challenges of real world scenarios. The focus of this work is the segmentation of groups of people into individuals. One relevant application of this algorithm is people counting. Experiments document that the presented approach leads to robust people counts. Xiaoming Liu 0002, Peter H. Tu, Jens Rittscher, A. G. Amitha Perera, Nils Krahnstoever |
AVSS | 3 |
| 2005 | Simultaneous Estimation of Segmentation and ShapeabstractThe main focus of this work is the integration of feature grouping and model based segmentation into one consistent framework. The algorithm is based on partitioning a given set of image features using a likelihood function that is parameterized on the shape and location of potential individuals in the scene. Using a variant of the EM formulation, maximum likelihood estimates of both the model parameters and the grouping are obtained simultaneously. The resulting algorithm performs global optimization and generates accurate results even when decisions can not be made using local context alone. An important feature of the algorithm is that the number of people in the scene is not modeled explicitly. As a result no prior knowledge or assumed distributions are required. The approach is shown to be robust with respect to partial occlusion, shadows, clutter, and can operate over a large range of challenging view angles including those that are parallel to the ground plane. Comparisons with existing crowd segmentation systems are made and the utility of coupling crowd segmentation with a temporal tracking system is demonstrated. Jens Rittscher, Peter H. Tu, Nils Krahnstoever |
CVPR (2) | 1 |
| 2003 | Site Calibration for Large Indoor ScenesabstractIndoor video surveillance networks as typically used in retail-networks need to be consistently calibrated with respect to a global coordinate system to provide the spatial context required for automatic recognition. This can be problematic if the fields of view of the cameras are disjoint. The paper presents a strategy to calibrate disjoint views into one coordinate system. The system is based on a device which can calculate its current position and orientation with respect to the global coordinate frame while it traverses the entire site without requiring direct line of sight between all sets of points. The paper explains why more traditional approaches have disadvantages and presents experimental results using the new method. Peter H. Tu, Jens Rittscher, Timothy P. Kelliher |
AVSS | 2 |
| 2003 | Video Content Annotation Using Visual Analysis and a Large Semantic KnowledgebaseabstractWe present a novel approach to automatically annotating broadcast video. To manage the enormous variety of objects, events and scenes in video problem domains such as news video, we couple generic image analysis with a semantic database, WordNet, containing huge amounts of real-world information. Object and event recognition are performed by searching WordNet for concepts jointly supported by image evidence and topic context derived from the video transcript. No object- specific or event-specific training is required, and only a few object models and detection algorithms are required to label much of the significant content of news video. The hierarchical structure of WordNet yields hierarchical recognition, dynamically tailored to the level of supporting image evidence. The potential of the approach is demonstrated by analyzing a wide variety of scenes in news video. Anthony Hoogs, Jens Rittscher, Gees C. Stein, John Schmiederer |
CVPR (2) | 2 |
| 2003 | Enabling video annotation using a semantic database extended with visual knowledgeabstractA semantic database has been extended with visual information to enable video annotation. This paper describes a lexical database, WordNet. We show its limitations with respect to describing visual characteristics, and describe an extension to WordNet that contains specific visual information. Having such a semantic database makes video annotation possible for broadcast news: a domain that can cover any topic and involve a wide variety of events, objects and scenes. Combining basic visual analysis techniques and a semantic database containing visual descriptions avoids the problem developing large numbers of specific object and event detectors. Such a semantic database can be of great value for the analysis of multi-modal information. As far as we know, such a database has not been developed before. Gees C. Stein, Jens Rittscher, Anthony Hoogs |
ICME | 2 |
| 2002 | Towards the automatic analysis of complex human body motions
Jens Rittscher, Andrew Blake 0001, Stephen J. Roberts |
Image Vis. Comput. | 1 |
| 2002 | An HMM-Based Segmentation Method for Traffic Monitoring MoviesabstractShadows of moving objects often obstruct robust visual tracking. We propose an HMM-based segmentation method which classifies in real time each pixel or region into three categories: shadows, foreground, and background objects. In the case of traffic monitoring movies, the effectiveness of the proposed method has been proven through experimental results. Jien Kato, Toyohide Watanabe, Sébastien Joga, Jens Rittscher, Andrew Blake 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2001 | Guiding Random Particles by Deterministic Search
Josephine Sullivan, Jens Rittscher |
ICCV | 2 |
| 2000 | A Probabilistic Background Model for Tracking
Jens Rittscher, Jien Kato, Sébastien Joga, Andrew Blake 0001 |
ECCV (2) | 1 |
| 2000 | Statistical Foreground Modelling for Object Localisation
Josephine Sullivan, Andrew Blake 0001, Jens Rittscher |
ECCV (2) | 3 |
| 2000 | An Integral Criterion for Detecting Boundary Edges and Textured RegionsabstractEdge maps which are computed from textured scenes using existing methods based on local image analysis are not very meaningful. This is because edges at object boundaries are not differentiated from edges in texture. We introduce a real-time algorithm that overcomes this difficulty by computing the Dirichlet integral in a small image patch at different scales. These measurements are combined and interpreted in a probabilistic framework avoiding the need for a threshold. As a result the output of this algorithm can be utilised by a higher level process. Texture does not fit our model of a boundary edge thus its presence is detected by the probabilistic model as an outlier. Convincing results are shown on synthetic as well as images of the natural world. This algorithm is intended to be a fast preprocessing step for localising boundary edges and textured image regions. Jens Rittscher, Josephine Sullivan |
ICPR | 1 |
| 2000 | Learning and Classification of Complex DynamicsabstractStandard, exact techniques based on likelihood maximization are available for learning auto-regressive process models of dynamical processes. The uncertainty of observations obtained from real sensors means that dynamics can be observed only approximately. Learning can still be achieved via "EM-K"-expectation-maximization (EM) based on Kalman filtering. This cannot handle more complex dynamics, however, involving multiple classes of motion. A problem arises also in the case of dynamical processes observed visually: background clutter arising for example, in camouflage, produces non-Gaussian observation noise. Even with a single dynamical class, non-Gaussian observations put the learning problem beyond the scope of EM-K. For those cases, we show here how "EM-C"-based on the CONDENSATION algorithm which propagates random "particle-sets," can solve the learning problem. Here, learning in clutter is studied experimentally using visual observations of a hand moving over a desktop. The resulting learned dynamical model is shown to have considerable predictive value: when used as a prior for estimation of motion, the burden of computation in visual observation is significantly reduced. Multiclass dynamics are studied via visually observed juggling; plausible dynamical models have been found to emerge from the learning process, and accurate classification of motion has resulted. In practice, EM-C learning is computationally burdensome and the paper concludes with some discussion of computational complexity. Ben North, Andrew Blake 0001, Michael Isard, Jens Rittscher |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 1999 | Classification of Human Body MotionabstractThe classification of human body motion is a difficult problem. In particular, the automatic segmentation of image sequences containing more than one class of motion is challenging. An effective approach is to use mixed discrete/continuous states to couple perception with classification. A spline contour is used to track the outline of the person. We show that, for a quasi-periodic human body motion, an autoregressive process is a suitable model for the contour dynamics. This can then be used as a dynamical model for mixed-state "condensation" filtering, switching automatically between different motion classes. We have developed "partial importance sampling" to enhance the efficiency of the mixed-state condensation filter. It is also shown that the importance sampling can be done in linear time, instead of the previous quadratic algorithm. "Tying" of discrete states is used to obtain further efficiency improvements. Automatic segmentation is demonstrated on video sequences of aerobic exercises. The performance is promising, but there remains a residual misclassification rate, and possible explanations for this are discussed. Jens Rittscher, Andrew Blake 0001 |
ICCV | 1 |