EDBT 2026 Demo / reviewers in the wild / expert
Boris Chidlovskii
dblp:c/BChidlovskii
· DBLP profile ↗
47ranked-venue papers
19as first author
15since 2021 · last 2025
0000-0002-2958-2361ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 24 · 14 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Systems, architecture and hardware · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MUSt3R: Multi-view Network for Stereo 3D ReconstructionabstractDUSt3R introduced a novel paradigm in geometric computer vision by proposing a model that can provide dense and unconstrained Stereo 3D Reconstruction of arbitrary image collections with no prior information about camera calibration nor viewpoint poses. Under the hood, however, DUSt3R processes image pairs, regressing local 3D reconstructions that need to be aligned in a global coordinate system. The number of pairs, growing quadratically, is an inherent limitation that becomes especially concerning for robust and fast optimization in the case of large image collections. In this paper, we propose an extension of DUSt3R from pairs to multiple views, that addresses all aforementioned concerns. Indeed, we propose a Multi-view Network for Stereo 3D Reconstruction, or MUSt3R, that modifies the DUSt3R architecture by making it symmetric and extending it to directly predict 3D structure for all views in a common coordinate frame. Second, we entail the model with a multi-layer memory mechanism which allows to reduce the computational complexity and to scale the reconstruction to large collections, inferring thousands of 3D pointmaps at high frame-rates with limited added complexity. The framework is designed to perform 3D reconstruction both offline and online, and hence can be seamlessly applied to SfM and visual SLAM scenarios showing state-of-the-art performance on various 3D downstream tasks, including uncalibrated Visual Odometry, relative camera pose, scale and focal estimation, 3D reconstruction and multi-view depth estimation. Yohann Cabon, Lucas Stoffl, Leonid Antsfeld, Gabriela Csurka, Boris Chidlovskii, Jérôme Revaud, Vincent Leroy 0003 |
CVPR | 5 |
| 2025 | Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems ApproachabstractProgress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-conditioned behavior, but benchmarks are still dominated by simulation. In this work, we focus on the fine-grained behavior of fast-moving real robots and present a large-scale experimental study involving 262 navigation episodes in a real environment with a physical robot, where we analyze the type of reasoning emerging from end-to-end training. In particular, we study the presence of realistic dynamics which the agent learned for open-loop forecasting, and their interplay with sensing. We analyze the way the agent uses latent memory to hold elements of the scene structure and information gathered during exploration. We probe the planning capabilities of the agent, and find in its memory evidence for somewhat precise plans over a limited horizon. Furthermore, we show in a post-hoc analysis that the value function learned by the agent relates to long-term planning. Put together, our experiments paint a new picture on how using tools from computer vision and sequential decision making have led to new capabilities in robotics and control. An interactive tool is available [here]. Steeven Janny, Hervé Poirier, Leonid Antsfeld, Guillaume Bono, Gianluca Monaci, Boris Chidlovskii, Francesco Giuliari, Alessio Del Bue, Christian Wolf 0001 |
CVPR | 6 |
| 2025 | PanSt3R: Multi-View Consistent Panoptic SegmentationabstractPanoptic segmentation of 3D scenes, involving the segmentation and classification of object instances in a dense 3D reconstruction of a scene, is a challenging problem, especially when relying solely on unposed 2D images. Existing approaches typically leverage off-the-shelf models to extract per-frame 2D panoptic segmentations, before optimizing an implicit geometric representation (often based on NeRF) to integrate and fuse the 2D predictions. We argue that relying on 2D panoptic segmentation for a problem inherently 3D and multi-view is likely suboptimal as it fails to leverage the full potential of spatial relationships across views. In addition to requiring camera parameters, these approaches also necessitate computationally expensive test-time optimization for each scene. Instead, in this work, we propose a unified and integrated approach PanSt3R, which eliminates the need for test-time optimization by jointly predicting 3D geometry and multi-view panoptic segmentation in a single forward pass. Our approach builds upon recent advances in 3D reconstruction, specifically upon MUSt3R, a scalable multi-view version of DUSt3R, and enhances it with semantic awareness and multi-view panoptic segmentation capabilities. We additionally revisit the standard post-processing mask merging procedure and introduce a more principled approach for multi-view segmentation. We also introduce a simple method for generating novel-view predictions based on the predictions of PanSt3R and vanilla 3DGS. Overall, the proposed PanSt3R is conceptually simple, yet fast and scalable, and achieves state-of-the-art performance on several benchmarks, while being orders of magnitude faster than existing methods. Lojze Zust, Yohann Cabon, Juliette Marrie, Leonid Antsfeld, Boris Chidlovskii, Jérôme Revaud, Gabriela Csurka |
ICCV | 5 |
| 2024 | Zero-BEV: Zero-shot Projection of Any First-Person Modality to BEV MapsabstractBird’s-eye view (BEV) maps are an important geometrically structured representation widely used in robotics, in particular self-driving vehicles and terrestrial robots. Existing algorithms either require depth information for the geometric projection, which is not always reliably available, or are trained end-to-end in a fully supervised way to map visual first-person observations to BEV representation, and are therefore restricted to the output modality they have been trained for. In contrast, we propose a new model capable of performing zero-shot projections of any modality available in a first person view to the corresponding BEV map. This is achieved by disentangling the geometric inverse perspective projection from the modality transformation, e.g. RGB to occupancy. The method is general and we showcase experiments projecting to BEV three different modalities: semantic segmentation, motion vectors and object bounding boxes detected in first person. We experimentally show that the model outperforms competing methods, in particular the widely used baseline resorting to monocular depth estimation. Gianluca Monaci, Leonid Antsfeld, Boris Chidlovskii, Christian Wolf 0001 |
3DV | 3 |
| 2024 | Learning to Navigate Efficiently and Precisely in Real EnvironmentsabstractIn the context of autonomous navigation of terrestrial robots, the creation of realistic models for agent dynamics and sensing is a widespread habit in the robotics literature and in commercial applications, where they are used for model based control and/or for localization and mapping. The more recent Embodied AI literature, on the other hand, focuses on modular or end-to-end agents trained in simulators like Habitat or AI-Thor, where the emphasis is put on photo-realistic rendering and scene diversity, but high-fidelity robot motion is assigned a less privileged role. The resulting sim2real gap significantly impacts transfer of the trained models to real robotic platforms. In this work we explore end-to-end training of agents in simulation in settings which minimize the sim2real gap both, in sensing and in actuation. Our agent directly predicts (discretized) velocity commands, which are maintained through closed-loop control in the real robot. The behavior of the real robot (including the underlying low-level controller) is identified and simulated in a modified Habitat simulator. Noise models for odometry and localization further contribute in lowering the sim2real gap. We evaluate on real navigation scenarios, explore different localization and point goal calculation methods and report significant gains in performance and robustness compared to prior work. Guillaume Bono, Hervé Poirier, Leonid Antsfeld, Gianluca Monaci, Boris Chidlovskii, Christian Wolf 0001 |
CVPR | 5 |
| 2024 | DUSt3R: Geometric 3D Vision Made EasyabstractMulti-view stereo reconstruction (MVS) in the wild re-quires to first estimate the camera intrinsic and extrinsic parameters. These are usually tedious and cumbersome to obtain, yet they are mandatory to triangulate corresponding pixels in 3D space, which is at the core of all best performing MVS algorithms. In this work, we take an opposite stance and introduce DUSt3R, a radically novel paradigm for Dense and Unconstrained Stereo 3D Reconstruction of arbitrary image collections, operating without prior infor-mation about camera calibration nor viewpoint poses. We cast the pairwise reconstruction problem as a regression of pointmaps, relaxing the hard constraints of usual projective camera models. We show that this formulation smoothly unifies the monocular and binocular reconstruction cases. In the case where more than two images are provided, we fur-ther propose a simple yet effective global alignment strategy that expresses all pairwise pointmaps in a common refer-ence frame. We base our network architecture on standard Transformer encoders and decoders, allowing us to leverage powerful pretrained models. Our formulation directly provides a 3D model of the scene as well as depth information, but interestingly, we can seamlessly recover from it, pixel matches, focal lengths, relative and absolute cameras. Extensive experiments on all these tasks showcase how DUSt3R effectively unifies various 3D vision tasks, setting new performance records on monocular & multi-view depth estimation as well as relative pose estimation. In summary, DUSt3R makes many geometric 3D vision tasks easy. Code and mod-els at https://github.com/naver/dust3r. Shuzhe Wang, Vincent Leroy 0003, Yohann Cabon, Boris Chidlovskii, Jérôme Revaud |
CVPR | 4 |
| 2024 | End-to-End (Instance)-Image Goal Navigation through Correspondence as an Emergent PhenomenonabstractMost recent work in goal oriented visual navigation resorts to large-scale machine learning in simulated environments. The main challenge lies in learning compact representations generalizable to unseen environments and in learning high-capacity perception modules capable of reasoning on high-dimensional input. The latter is particularly difficult when the goal is not given as a category ("ObjectNav") but as an exemplar image ("ImageNav"), as the perception module needs to learn a comparison strategy requiring to solve an underlying visual correspondence problem. This has been shown to be difficult from reward alone or with standard auxiliary tasks. We address this problem through a sequence of two pretext tasks, which serve as a prior for what we argue is one of the main bottleneck in perception, extremely wide-baseline relative pose estimation and visibility prediction in complex scenes. The first pretext task, cross-view completion is a proxy for the underlying visual correspondence problem, while the second task addresses goal detection and finding directly. We propose a new dual encoder with a large-capacity binocular ViT model and show that correspondence solutions naturally emerge from the training signals. Experiments show significant improvements and SOTA performance on the two benchmarks, ImageNav and the Instance-ImageNav variant, where camera intrinsics and height differ between observation and goal. Guillaume Bono, Leonid Antsfeld, Boris Chidlovskii, Philippe Weinzaepfel, Christian Wolf 0001 |
ICLR | 3 |
| 2024 | Self-supervised Pretraining and Finetuning for Monocular Depth and Visual OdometryabstractFor the task of simultaneous monocular depth and visual odometry estimation, we propose learning self-supervised transformer-based models in two steps. Our first step consists in a generic pretraining to learn 3D geometry, using cross-view completion objective (CroCo), followed by self-supervised finetuning on non-annotated videos. We show that our self-supervised models can reach state-of-the-art performance ’without bells and whistles’ using standard components such as visual transformers, dense prediction transformers and adapters. We demonstrate the effectiveness of our proposed method by running evaluations on six benchmark datasets, both static and dynamic, indoor and outdoor, with synthetic and real images. For all datasets, our method outperforms state-of-the-art methods, in particular for depth prediction task. Leonid Antsfeld, Boris Chidlovskii |
ICRA | 2 |
| 2023 | CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowabstractDespite impressive performance for high-level downstream tasks, self-supervised pre-training methods have not yet fully delivered on dense geometric vision tasks such as stereo matching or optical flow. The application of self-supervised concepts, such as instance discrimination or masked image modeling, to geometric tasks is an active area of research. In this work, we build on the recent cross-view completion framework, a variation of masked image modeling that leverages a second view from the same scene which makes it well suited for binocular downstream tasks. The applicability of this concept has so far been limited in at least two ways: (a) by the difficulty of collecting real-world image pairs – in practice only synthetic data have been used – and (b) by the lack of generalization of vanilla transformers to dense downstream tasks for which relative position is more meaningful than absolute position. We explore three avenues of improvement. First, we introduce a method to collect suitable real-world image pairs at large scale. Second, we experiment with relative positional embeddings and show that they enable vision transformers to perform substantially better. Third, we scale up vision transformer based cross-completion architectures, which is made possible by the use of large amounts of data. With these improvements, we show for the first time that state-of-the-art results on stereo matching and optical flow can be reached without using any classical task-specific techniques like correlation volume, iterative estimation, image warping or multi-scale reasoning, thus paving the way towards universal vision models. Philippe Weinzaepfel, Thomas Lucas 0002, Vincent Leroy 0003, Yohann Cabon, Vaibhav Arora, Romain Brégier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, Jérôme Revaud |
ICCV | 9 |
| 2023 | Multi-Object Navigation in real environments using hybrid policiesabstractNavigation has been classically solved in robotics through the combination of SLAM and planning. More recently, beyond waypoint planning, problems involving significant components of (visual) high-level reasoning have been explored in simulated environments, mostly addressed with large-scale machine learning, in particular RL, offline-RL or imitation learning. These methods require the agent to learn various skills like local planning, mapping objects and querying the learned spatial representations. In contrast to simpler tasks like waypoint planning (PointGoal), for these more complex tasks the current state-of-the-art models have been thoroughly evaluated in simulation but, to our best knowledge, not yet in real environments. In this work we focus on sim2real transfer. We target the challenging Multi-Object Navigation (Multi-ON) task [41] and port it to a physical environment containing real replicas of the originally virtual Multi-ON objects. We introduce a hybrid navigation method, which decomposes the problem into two different skills: (1) waypoint navigation is addressed with classical SLAM combined with a symbolic planner, whereas (2) exploration, semantic mapping and goal retrieval are dealt with deep neural networks trained with a combination of supervised learning and RL. We show the advantages of this approach compared to end-to-end methods both in simulation and a real environment and outperform the SOTA for this task [28]. Assem Sadek, Guillaume Bono, Boris Chidlovskii, Atilla Baskurt, Christian Wolf 0001 |
ICRA | 3 |
| 2023 | Learning Whom to Trust in Navigation: Dynamically Switching Between Classical and Neural PlanningabstractNavigation of terrestrial robots is typically addressed either with localization and mapping (SLAM) followed by classical planning on the dynamically created maps, or by machine learning (ML), often through end-to-end training with reinforcement learning (RL) or imitation learning (IL). Recently, modular designs have achieved promising results, and hybrid algorithms that combine ML with classical planning have been proposed. Existing methods implement these combinations with handcrafted functions, which cannot fully exploit the complementary nature of the policies and the complex regularities between scene structure and planning performance. Our work builds on the hypothesis that the strengths and weaknesses of neural planners and classical planners follow some regularities, which can be learned from training data, in particular from interactions. This is grounded on the assumption that, both, trained planners and the mapping algorithms underlying classical planning are subject to failure cases depending on the semantics of the scene and that this dependence is learnable: for instance, certain areas, objects or scene structures can be reconstructed easier than others. We propose a hierarchical method composed of a high-level planner dynamically switching between a classical and a neural planner. We fully train all neural policies in simulation and evaluate the method in both simulation and real experiments with a LoCoBot robot, showing significant gains in performance, in particular in the real environment. We also qualitatively conjecture on the nature of data regularities exploited by the high-level planner. Sombit Dey, Assem Sadek, Gianluca Monaci, Boris Chidlovskii, Christian Wolf 0001 |
IROS | 4 |
| 2022 | PUMP: Pyramidal and Uniqueness Matching Priors for Unsupervised Learning of Local DescriptorsabstractExisting approaches for learning local image descriptors have shown remarkable achievements in a wide range of geometric tasks. However, most of them require perpixel correspondence-level supervision, which is difficult to acquire at scale and in high quality. In this paper, we propose to explicitly integrate two matching priors in a single loss in order to learn local descriptors without supervision. Given two images depicting the same scene, we extract pixel descriptors and build a correlation volume. The first prior enforces the local consistency of matches in this volume via a pyramidal structure iteratively constructed using a non-parametric module. The second prior exploits the fact that each descriptor should match with at most one descriptor from the other image. We combine our unsupervised loss with a standard self-supervised loss trained from synthetic image augmentations. Feature descriptors learned by the proposed approach outperform their fully- and self-supervised counterparts on various geometric benchmarks such as visual localization and image matching, achieving state-of-the-art performance. Project webpage: https://europe.naverlabs.com/research/3d-vision/pump. Jérôme Revaud, Vincent Leroy 0003, Philippe Weinzaepfel, Boris Chidlovskii |
CVPR | 4 |
| 2022 | An in-depth experimental study of sensor usage and visual reasoning of robots navigating in real environmentsabstractVisual navigation by mobile robots is classically tackled through SLAM plus optimal planning, and more recently through end-to-end training of policies implemented as deep networks. While the former are often limited to waypoint planning, but have proven their efficiency even on real physical environments, the latter solutions are most frequently employed in simulation, but have been shown to be able learn more complex visual reasoning, involving complex semantic regularities. Navigation by real robots in physical environments is still an open problem. End-to-end training approaches have been thoroughly tested in simulation only, with experiments involving real robots being restricted to rare performance evaluations in simplified laboratory conditions. In this work we present an in-depth study of the performance and reasoning capacities of real physical agents, trained in simulation and deployed to two different physical environments. Beyond benchmarking, we provide insights into the generalization capabilities of different agents training in different conditions. We visualize sensor usage and the importance of the different types of signals. We show, that for the PointGoal task, an agent pre-trained on wide variety of tasks and fine-tuned on a simulated version of the target environment can reach competitive performance without modelling any sim2real transfer, i.e. by deploying the trained agent directly from simulation to a real physical robot. Assem Sadek, Guillaume Bono, Boris Chidlovskii, Christian Wolf 0001 |
ICRA | 3 |
| 2022 | CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionabstractMasked Image Modeling (MIM) has recently been established as a potent pre-training paradigm. A pretext task is constructed by masking patches in an input image, and this masked content is then predicted by a neural network using visible patches as sole input. This pre-training leads to state-of-the-art performance when finetuned for high-level semantic tasks, e.g. image classification and object detection. In this paper we instead seek to learn representations that transfer well to a wide variety of 3D vision and lower-level geometric downstream tasks, such as depth prediction or optical flow estimation. Inspired by MIM, we propose an unsupervised representation learning task trained from pairs of images showing the same scene from different viewpoints. More precisely, we propose the pretext task of cross-view completion where the first input image is partially masked, and this masked content has to be reconstructed from the visible content and the second image. In single-view MIM, the masked content often cannot be inferred precisely from the visible portion only, so the model learns to act as a prior influenced by high-level semantics. In contrast, this ambiguity can be resolved with cross-view completion from the second unmasked image, on the condition that the model is able to understand the spatial relationship between the two images. Our experiments show that our pretext task leads to significantly improved performance for monocular 3D vision downstream tasks such as depth estimation. In addition, our model can be directly applied to binocular downstream tasks like optical flow or relative camera pose estimation, for which we obtain competitive results without bells and whistles, i.e., using a generic architecture without any task-specific design. Philippe Weinzaepfel, Vincent Leroy 0003, Thomas Lucas 0002, Romain Brégier, Yohann Cabon, Vaibhav Arora, Leonid Antsfeld, Boris Chidlovskii, Gabriela Csurka, Jérôme Revaud |
NeurIPS | 8 |
| 2021 | Magnetic Field Sensing for Pedestrian and Robot Indoor PositioningabstractIn this paper we address the problem of indoor localization using magnetic field data in two setups, when data is collected by (i) human-held mobile phone and (ii) by localization robots that perturb magnetic data with their own electromagnetic field. For the first setup, we revise the state of the art approaches and propose a novel extended pipeline to benefit from the presence of magnetic anomalies in indoor environment created by different ferromagnetic objects. We capture changes of the Earth’s magnetic field due to indoor magnetic anomalies and transform them in multi-variate times series. We then convert temporal patterns into visual ones. We use methods of Recurrence Plots, Gramian Angular Fields and Markov Transition Fields to represent magnetic field time series as image sequences. We regress the continuous values of user position in a deep neural network that combines convolutional and recurrent layers. For the second setup, we analyse how magnetic field data get perturbed by robots’ electromagnetic field. We add an alignment step to the main pipeline, in order to compensate the mismatch between train and test sets obtained by different robots. We test our methods on two public (MagPie [11] and IPIN’20 [4]) and one proprietary (Hyundai department store) datasets. We report evaluation results and show that our methods outperform the state of the art methods by a large margin. Leonid Antsfeld, Boris Chidlovskii |
IPIN | 2 |
| 2020 | Self-Supervised Attention Learning for Depth and Ego-motion EstimationabstractWe address the problem of depth and ego-motion estimation from image sequences. Recent advances in the domain propose to train a deep learning model for both tasks using image reconstruction in a self-supervised manner. We revise the assumptions and the limitations of the current approaches and propose two improvements to boost the performance of the depth and ego-motion estimation. We first use Lie group properties to enforce the geometric consistency between images in the sequence and their reconstructions. We then propose a mechanism to pay attention to image regions where the image reconstruction gets corrupted. We show how to integrate the attention mechanism in the form of attention gates in the pipeline and use attention coefficients as a mask. We evaluate the new architecture on the KITTI datasets and compare it to the previous techniques. We show that our approach improves the state-of-the-art results for ego-motion estimation and achieve comparable results for depth estimation. Assem Sadek, Boris Chidlovskii |
IROS | 2 |
| 2020 | Magnetic sensor based indoor positioning by multi-channel deep regression: poster abstractabstractModern smartphones are equipped with built-in magnetometers that capture disturbances of the Earth's magnetic field induced by ferromagnetic objects. In indoor environment, using magnetic field data turns to be a strong alternative to conventional localization techniques as requiring no special infrastructure. We revise the state of the art methods based on landmark classification [5] and propose a novel approach. We represent magnetic data time series as image sequences and compose multi-channel input to a deep neural network. We use four methods, Recurrence plots, Gramian Angular Fields and Markov Transition Fields, to capture different patterns in magnetic data stream. We complete the landmark-based classification with deep regression on the user's position and combine convolutional and recurrent layers in the deep network. We evaluate our methods on the recently published MagPie dataset [3] and show that they outperform the state of the art methods. Leonid Antsfeld, Boris Chidlovskii, Dmitrii Borisov |
SenSys | 2 |
| 2019 | Semi-supervised Variational Autoencoder for WiFi Indoor LocalizationabstractWe address the problem of indoor localization based on WiFi signal strengths. We develop a semi-supervised deep learning method able to train a prediction model from a small set of annotated WiFi observations and a massive set of non-annotated ones. Our method is based on the variational autoencoder deep network. We complement the network with an additional component of structural projection able to further improve the localization accuracy in a complex, multi-building and multi-floor environment. We consider several different network compositions which combine the classification and regression sub-tasks to achieve optimal performance. We evaluate our method on the public UJI-IndoorLoc dataset and show that the proposed method allows to maintain the state of the art localization accuracy with a very limited amount of annotated data. Boris Chidlovskii, Leonid Antsfeld |
IPIN | 1 |
| 2018 | Mining Smart Card Data for Travellers' Mini ActivitiesabstractIn the context of public transport modeling and simulation, we address the problem of mismatch between simulated transit trips and observed ones. We point to the weakness of the current travel demand modeling process; the trips it generates are overly optimistic and do not reflect the real passenger choices. To explain the deviation of simulated trips from the observed trips, we introduce the notion of mini-activities the travelers do during the trips. We propose to mine the smart card data and identify characteristics that help detect the mini activities. We develop a technique to integrate them in the generated trips and learn such an integration from two available sources, the trip history and trip planner recommendations. For an input travel demand, we build a Markov chain over the trip collection and apply the Monte Carlo Markov Chain algorithm to integrate mini activities in such a way that the trip characteristics converge to the target distributions. We test our method on the trip data set collected in Nancy, France. The evaluation results demonstrate a very important reduction of the trip generation error, and a good capacity to cope with new simulation scenarios. Boris Chidlovskii |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Domain Adaptation in the Absence of Source Domain DataabstractThe overwhelming majority of existing domain adaptation methods makes an assumption of freely available source domain data. An equal access to both source and target data makes it possible to measure the discrepancy between their distributions and to build representations common to both target and source domains. In reality, such a simplifying assumption rarely holds, since source data are routinely a subject of legal and contractual constraints between data owners and data customers. When source domain data can not be accessed, decision making procedures are often available for adaptation nevertheless. These procedures are often presented in the form of classification, identification, ranking etc. rules trained on source data and made ready for a direct deployment and later reuse. In other cases, the owner of a source data is allowed to share a few representative examples such as class means. In this paper we address the domain adaptation problem in real world applications, where the reuse of source domain data is limited to classification rules or a few representative examples. We extend the recent techniques of feature corruption and their marginalization, both in supervised and unsupervised settings. We test and compare them on private and publicly available source datasets and show that significant performance gains can be achieved despite the absence of source data and shortage of labeled target data. Boris Chidlovskii, Stéphane Clinchant, Gabriela Csurka |
KDD | 1 |
| 2015 | Learning urban users' choices to improve trip recommendationsabstractWe analyze the work of urban trip planners and the relevance of trips they recommend upon user queries. We propose to improve the planner recommendations by learning from choices made by travelers who use the transportation network on the daily basis. We analyze a large collection of individual travelers' trips collected from the automated fare collection systems; we convert the trips into pair-wise preferences for traveling from a given origin to a destination at a given time point. We model passenger preferences with a number of smoothed time-dependent latent variables which are used to learn a ranking function for trips. This function can be used to re-rank the top planner's recommendations. Results of tests for cities of Nancy, France and Adelaide, Australia show a considerable increase of the recommendation relevance. Boris Chidlovskii |
DSAA | 1 |
| 2013 | Connecting comments and tags: improved modeling of social tagging systemsabstractCollaborative tagging systems are now deployed extensively to help users share and organize resources. Tag prediction and recommendation can simplify and streamline the user experience, and by modeling user preferences, predictive accuracy can be significantly improved. However, previous methods typically model user behavior based only on a log of prior tags, neglecting other behaviors and information in social tagging systems, e.g., commenting on items and connecting with other users. On the other hand, little is known about the connection and correlations among these behaviors and contexts in social tagging systems. Dawei Yin 0001, Shengbo Guo, Boris Chidlovskii, Brian D. Davison 0001, Cédric Archambeau, Guillaume Bouchard |
WSDM | 3 |
| 2012 | Tag Ranking by Linear Relational Neighbourhood PropagationabstractWe propose a tag recommendation method which can assist users in tagging process by suggesting relevant tags. The method is based on query-based ranking on relational multi-type graphs which capture the annotation relationship between objects and tags, as well as the object similarity and tag correlation. The additional advance consists in extending the linear neighbourhood propagation to the relational graphs with the Laplacian regularization framework. We report evaluation results on a large-scale Flickr data set. Boris Chidlovskii |
ASONAM | 1 |
| 2012 | Learning Multiple Tasks with Boosted Decision Trees
Jean Baptiste Faddoul, Boris Chidlovskii, Rémi Gilleron, Fabien Torre |
ECML/PKDD (1) | 2 |
| 2011 | Local metric learning for tag recommendation in social networksabstractWe address the problem of tag recommendation for media objects, like images, videos, etc in social media sharing systems. We propose a framework that 1) extracts both object features and the social context and 2) uses them to learn recommendation rules. The social context is described by different types of information, such as a user's personal objects, the objects of a user's social contacts, the importance of the user in the social network, etc. Both object features and the social context are first used to guide the k-nearest neighbour method for the tag recommendation. We then enhance the method by the local topology adjustment on how the nearest neighbours are selected. We learn a local transformation of the feature space surrounding a given object which pushes together objects with the same tags and puts apart objects with different tags. We show how to learn the Mahalanobis distance metric on multi-tag objects and adopt it to the tag recommendation problem. Boris Chidlovskii, Aymen Benzarti |
ACM Symposium on Document Engineering | 1 |
| 2011 | Learning Recommendations in Social Media Systems by Weighting Multiple Relations
Boris Chidlovskii |
ECML/PKDD (1) | 1 |
| 2010 | Boosting Multi-Task Weak Learners with Applications to Textual and Social DataabstractLearning multiple related tasks from data simultaneously can improve predictive performance relative to learning these tasks independently. In this paper we propose a novel multi-task learning algorithm called MT-Adaboost: it extends Ada boost algorithm to the multi-task setting; it uses as multi-task weak classifier a multi-task decision stump. This allows to learn different dependencies between tasks for different regions of the learning space. Thus, we relax the conventional hypothesis that tasks behave similarly in the whole learning space. Moreover, MT-Adaboost can learn multiple tasks without imposing the constraint of sharing the same label set and/or examples between tasks. A theoretical analysis is derived from the analysis of the original Adaboost. Experiments for multiple tasks over large scale textual data sets with social context (Enron and Tobacco) give rise to very promising results. Jean Baptiste Faddoul, Boris Chidlovskii, Fabien Torre, Rémi Gilleron |
ICMLA | 2 |
| 2010 | Multi-modality in one-class classificationabstractWe propose a method for improving classification performance in a one-class setting by combining classifiers of different modalities. We apply the method to the problem of distinguishing responsive documents in a corpus of e-mails, like Enron Corpus. We extract the social network of actors which is implicit in a large body of electronic communication and turn it into valuable features for classifying the exchanged documents. Working in a one-class setting we folow a semi-supervised approach based on the Mapping Convergence framework. We propose an alternative interpretation, that allows for broader applicability when positive and negative items are not naturally separable. We propose an extension to the one-class evaluation framework in truly one-case cases when only some positive training examples are available. We extent the one-class setting to the co-training principle that enables us to take advantage of multiple views on the data. We report evaluation results of this extension on three different corpora including Enron Corpus. Matthijs Hovelynck, Boris Chidlovskii |
WWW | 2 |
| 2009 | Scalable Feature Extraction from Noisy DocumentsabstractWe cope with the metadata recognition in layout-oriented documents. We address the problem as a classification task and propose a method for automatic extraction of relevant features, in presence of content and structural noise, caused by scanning, OCR and segmentation problems. The method is based on the automatic analysis of documents and requires no particular preprocessing. The method mines the documents and determines frequent patterns, which are bothliteral patterns and their generalization. We also propose a sampling technique which processes a sample of documents and uses the Chernoff bounds to estimate the pattern frequency in the entire dataset. As a number of frequent patterns as feature candidates grows, the method applies a scalable feature selection method to determine the most relevant features to a given classification task. A series of evaluations on two collections show that the method performs comparably to the manual work on rule writing made by domain experts. Loïc Lecerf, Boris Chidlovskii |
ICDAR | 2 |
| 2008 | Scalable Feature Selection for Multi-class Problems
Boris Chidlovskii, Loïc Lecerf |
ECML/PKDD (1) | 1 |
| 2006 | ALDAI: active learning documents annotation interfaceabstractIn the framework of the LegDoC project at XRCE, we present a document annotation interface with an integrated active learning component. The interface is designed for automating the annotation of layout-oriented documents with semantic labels. We describe core functionalities of the interface; we pay a particular attention to the learning component, including the feature management and different strategies of deploying the active learner in the document annotation. Boris Chidlovskii, Jérôme Fuselier, Loïc Lecerf |
ACM Symposium on Document Engineering | 1 |
| 2006 | Document annotation by active learning techniquesabstractWe present a system for the semantic annotation of layout-oriented documents, with an integrated learning component. We introduce probabilistic learning methods on tree-like documents and we present different active learning techniques for training document annotation models. We report some preliminary results of deploying such active learning techniques on an important case of document collection annotation. Loïc Lecerf, Boris Chidlovskii |
ACM Symposium on Document Engineering | 2 |
| 2006 | Documentum ECI self-repairing wrappers: performance analysisabstractDocumentum Enterprise Content Integration (ECI) services is a content integration middleware that provides one-query access to the Intranet and Internet content resources. The ECI Adapter technology offers an interface to any application for data and metadata extraction from unstructured Web pages. It offers a unique frame-work of wrapper production, automatic recovery and maintenance, developed at Xerox Research Centre Europe and based on state-of-art algorithms from machine learning and grammatical inference. In this presentation we analyze the performance of ECI adapters deployed in current commercial installations. We benefit from accessing reports on daily tests for all ECI commercially deployed adapters collected from June 2003 to September 2005. Using the daily reports, we analyze different aspects of the wrapper technology. Boris Chidlovskii, Bruno Roustant, Marc Brette |
SIGMOD Conference | 1 |
| 2005 | A web-based document harmonization and annotation chain: from PDF to RDFabstractWe propose a demonstration of a Web-based document harmonization and annotation chain developed within the VIKEF integrated project. The chain integrates a combination of Web Services in order to access, harmonize and semantically annotate remote document collections. Annotations are then mapped onto RDF descriptions that serve as a basis for building semantic-enabled services to support community processes. Thierry Jacquin, Olivier Fambon, Boris Chidlovskii |
ACM Symposium on Document Engineering | 3 |
| 2005 | A Probabilistic Learning Method for XML Annotation of Documents
Boris Chidlovskii, Jérôme Fuselier |
IJCAI | 1 |
| 2004 | Supervised learning for the legacy document conversionabstractWe consider the problem of document conversion from the rendering-oriented HTML markup into a semantic-oriented XML annotation defined by user-specific DTDs or XML Schema descriptions. We represent both source and target documents as rooted ordered trees so the conversion can be achieved by applying a set of tree transformations. We apply the supervised learning framework to the conversion task according to which the tree transformations are learned from a set of training examples. %Because of the complexity of tree-to-tree transformations, We develop a two-step approach to the conversion problem, that first labels leaves in the source trees and then recomposes target trees from the leaf labels. We present two solutions based of the leaf classification with the target terminals and paths. Moreover, we develop three methods for the leaf classification. All methods and solutions have been tested on two real collections. Boris Chidlovskii, Jérôme Fuselier |
ACM Symposium on Document Engineering | 1 |
| 2003 | A structural adviser for the XML document authoringabstractSince the XML format became a de facto standard for structured documents, the IT research and industry have developed a number of XML editors to help users produce structured documents in XML format. However, the manual generation of structured documents in XML format remains a tedious and time-consuming process because of the excessive verbosity and length of XML code. In this paper, we design a structural adviser for the XML document authoring. The adviser intervenes at any step of the authoring process to suggest one tag or entire tree-like pattern the user is most likely to use next. Adviser suggestions are based on finding analogies between the currently edited fragment and sample data being either previously generated documents in the collection or the history of the current document authoring. The adviser is beneficial in cases when no schema is provided for XML documents, or schema associated with the document is too general and sample data contain specific patterns not captured in the schema. We design the adviser architecture and develop a method for efficient indexing and retrieval of optimal suggestions at any step of the document authoring. Boris Chidlovskii |
ACM Symposium on Document Engineering | 1 |
| 2003 | Crawling for Domain-Specific Hidden Web ResourcesabstractThe Hidden Web, the part of the Web that remains unavailable for standard crawlers, has become an important research topic during recent years. Its size is estimated to 400 to 500 times larger than that of the publicly indexable Web (PIW). Furthermore, the information on the hidden Web is assumed to be more structured, because it is usually stored in databases. In this paper, we describe a crawler which starting from the PIW finds entry points into the hidden Web. The crawler is domain-specific and is initialized with pre-classified documents and relevant keywords. We describe our approach to the automatic identification of Hidden Web resources among encountered HTML forms. We conduct a series of experiments using the top-level categories in the Google directory and report our analysis of the discovered Hidden Web resources. André Bergholz, Boris Chidlovskii |
WISE | 2 |
| 2002 | Automatic Repairing of Web Wrappers by Combining Redundant ViewsabstractWe address the problem of automatic maintenance of Web wrappers used in data integration systems to encapsulate an access to Web information providers. The maintenance of Web wrappers is critical as providers often changes the page format and/or structure making wrappers inoperable. The solution we propose extends the conventional wrapper architecture with a novel component of automatic maintenance and recovery. We consider the automatic recovery as special type of the classification problem and use ensemble methods of machine learning to build alternative views of provider pages. We combine extraction rules of conventional wrappers with content features of extracted information to accurate recovery from three types of format changes, namely, content, context and structural changes. We report results of the recovery performance for format changes at widely used Web providers. Boris Chidlovskii |
ICTAI | 1 |
| 2001 | Wrapping Web Information Providers by Transducer Induction
Boris Chidlovskii |
ECML | 1 |
| 2000 | Wrapper Generation via Grammar Induction
Boris Chidlovskii, Jon Ragetli, Maarten de Rijke |
ECML | 1 |
| 2000 | Semantic Caching of Web Queries
Boris Chidlovskii, Uwe M. Borghoff |
VLDB J. | 1 |
| 1999 | Approximation Techniques for Indexing Two-Dimensional Constraint DatabasesabstractConstraint databases have recently been proposed as a powerful framework to model and retrieve spatial data. The use of constraint databases should be supported by access data structures that make effective use of secondary storage and reduce query processing time. In this paper, we consider the indexing problem for objects represented by conjunctions of two-variable linear constraints and we analyze the problem of determining all generalized tuples whose extension intersects or is contained in the extension of a given half-plane. In an earlier paper we have shown that both selection problems can be reduced to a point location problem by using a dual transformation. If the angular coefficient of the half-plane belongs to a predefined set, we have proved that a dynamic optimal indexing solution, based on B/sup +/-trees, exists. In this paper we propose two approximation techniques that can be used to find the result when the angular coefficient does not belong to the predefined set. We also experimentally compare the proposed techniques with R-trees. Elisa Bertino, Barbara Catania, Boris Chidlovskii |
DASFAA | 3 |
| 1999 | Indexing Constraint Databases by Using a Dual RepresentationabstractLinear constraint databases are a powerful framework to model spatial and temporal data. The use of constraint databases should be supported by access data structures that make effective use of secondary storage and reduce query processing time. Such structures should be able to store both finite and infinite objects and perform both containment (ALL) and intersection (EXIST) queries. As standard indexing techniques have certain limitations in satisfying such requirements, we employ the concept of geometric duality for designing new indexing techniques. In (Bertino et al., 1997) we have used the dual transformation for polyhedra to develop a dynamic optimal indexing solution based on B/sup +/-trees, to detect all objects contained in or intersecting a given half-plane, when the angular coefficient belongs to a predefined set. We extend the previous solution to allow angular coefficients to take any value. We present two approximation techniques for the dual representation of spatial objects, based on B/sup +/-trees. The techniques handle both finite and infinite objects and process both ALL and EXIST selections in a uniform way. We show the practical applicability of the proposed techniques by an experimental comparison with respect to R/sup +/-trees. Elisa Bertino, Barbara Catania, Boris Chidlovskii |
ICDE | 3 |
| 1999 | Semantic Cache Mechanism for Heterogeneous Web Querying
Boris Chidlovskii, Claudia Roncancio, Marie-Luise Schneider |
Comput. Networks | 1 |
| 1998 | Query Translation for Distributed Information Gathering on the WebabstractThe heterogeneity of Web information services poses new problems in the processing of user queries over distributed federated data. It was recently proven that the query translation between two information services supporting all Boolean operators can be done in an optimal way. On the Web, however, the situation when at least one Boolean operator is not supported is frequent. We study the case where a Boolean user query cannot be directly translated and should be split into sub-queries. We propose two strategies for query subsumption and discuss in detail the strategy which minimizes the number of submitted sub-queries. We derive an appropriate query form for the minimal strategy and demonstrate how both translation strategies are implemented in the Knowledge Brokers system. Boris Chidlovskii, Uwe M. Borghoff |
IDEAS | 1 |
| 1997 | The Constraint-Based Knowledge Broker SystemabstractSummary form only given. The amount of information available from electronic sources on the World Wide Web and other on-line information repositories is highly heterogeneous and increasing dramatically. Tools are needed to extract relevant information from these repositories. The Constraint-Based Knowledge Brokers project (CBKB) at RXRC Grenoble realizes sophisticated facilities for efficient information retrieval, schema integration, and knowledge fusion. The current implementation of the CBKB research prototype involves three kinds of agents: a) users, who input queries and process answers (i.e., ranking, fusion) through a GUI; b) wrappers, capable of interrogating heterogeneous information sources, which can provide answers to elementary queries (essentially various public bibliographic catalogues available on the Web, as well as preprint archives and opera information repositories); c) brokers, which can manage complex queries (i.e., decompose a complex query, recompose the partial answers, synthesize a full answer) and which mediate between the GUI and the different wrappers. The core of the system is given by the brokers, which provide various important services such as intelligent caching, filtering and knowledge combination. Requests, intermediate information, and results are internally represented via feature constraints. Requests do not need to be fully defined; they may correspond to partial specifications of the requested information. Jean-Marc Andreoli, Uwe M. Borghoff, Pierre-Yves Chevalier, Boris Chidlovskii, Remo Pareschi, Jutta Willamowski |
ICDE | 4 |