VLDB 2026 Research / reviewers in the wild / expert
Piotr Mirowski
dblp:27/5853 · also Piotr W. Mirowski
· DBLP profile ↗
23ranked-venue papers
15as first author
8since 2021 · last 2026
0000-0002-8685-0932ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-authorHuman-computer interaction and ubiquitous computing · 5 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spontaneous Spectacles: AR- and LLM-Enabled Improvised Theatre PracticeabstractWe employ augmented reality (AR) glasses combined with automated speech recognition, and optionally, our bespoke system based on large language models (LLMs), to enable novice (and challenge veteran) improvisers to join us in improvisational theatre scenes. Our AR+LLM improv system has been validated in 51 live theatre performances and 3 workshops. This interactive demo session is an applied improvisation workshop open to participants without theatre experience, and is facilitated by two professional performers who help “get you out of your head”. Piotr Mirowski, Boyd Branch, Kory W. Mathewson |
Creativity & Cognition | 1 |
| 2025 | AI and Non-Western Art Worlds: Reimagining Critical AI Futures through Artistic Inquiry and Situated Dialogue
Rida Qadri, Piotr Mirowski, Remi Denton |
CHI | 2 |
| 2024 | Designing and Evaluating Dialogue LLMs for Co-Creative Improvised Theatre
Boyd Branch, Piotr Mirowski, Kory W. Mathewson, Sophia Ppali, Alexandra Covaci |
ICCC | 2 |
| 2024 | Visual theatrical improvisation alongside Artificial Intelligence image generators
Piotr Mirowski, Boyd Branch, Kory W. Mathewson |
ICCC | 1 |
| 2023 | Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry ProfessionalsabstractLanguage models are increasingly attracting interest from writers. However, such models lack long-range semantic coherence, limiting their usefulness for longform creative writing. We address this limitation by applying language models hierarchically, in a system we call Dramatron. By building structural context via prompt chaining, Dramatron can generate coherent scripts and screenplays complete with title, characters, story beats, location descriptions, and dialogue. We illustrate Dramatron’s usefulness as an interactive co-creative system with a user study of 15 theatre and film industry professionals. Participants co-wrote theatre scripts and screenplays with Dramatron and engaged in open-ended interviews. We report reflections both from our interviewees and from independent reviewers who critiqued performances of several of the scripts to illustrate how both Dramatron and hierarchical text generation could be useful for human-machine co-creativity. Finally, we discuss the suitability of Dramatron for co-creativity, ethical considerations—including plagiarism and bias—and participatory models for the design and deployment of such tools. Piotr Mirowski, Kory W. Mathewson, Jaylen Pittman, Richard Evans 0001 |
CHI | 1 |
| 2022 | CLIP-CLOP: CLIP-Guided Collage and Photomontage
Piotr Mirowski, Dylan Banarse, Mateusz Malinowski, Simon Osindero, Chrisantha Fernando |
ICCC | 1 |
| 2021 | Tele-Immersive Improv: Effects of Immersive Visualisations on Rehearsing and Performing Theatre OnlineabstractPerformers acutely need but lack tools to remotely rehearse and create live theatre, particularly due to global restrictions on social interactions during the Covid-19 pandemic. No studies, however, have heretofore examined how remote video-collaboration affects performance. This paper presents the findings of a field study with 16 domain experts over six weeks investigating how tele-immersion affects the rehearsal and performance of improvisational theatre. To conduct the study, an original media server was developed for co-locating remote performers into shared virtual 3D environments which were accessed through popular video conferencing software. The results of this qualitative study indicate that tele-immersive environments uniquely provide performers with a strong sense of co- presence, feelings of physical connection, and an increased ability to enter the social-flow states required for improvisational theatre. Based on our observations, we put forward design recommendations for video collaboration tools tailored to the unique demands of live performance. Boyd Branch, Christos Efstratiou, Piotr Mirowski, Kory W. Mathewson, Paul Allain |
CHI | 3 |
| 2021 | Collaborative Storytelling with Human Actors and AI Narrators
Boyd Branch, Piotr Mirowski, Kory W. Mathewson |
ICCC | 2 |
| 2020 | Learning to Follow Directions in Street ViewabstractNavigating and understanding the real world remains a key challenge in machine learning and inspires a great variety of research in areas such as language grounding, planning, navigation and computer vision. We propose an instruction-following task that requires all of the above, and which combines the practicality of simulated environments with the challenges of ambiguous, noisy real world data. StreetNav is built on top of Google Street View and provides visually accurate environments representing real places. Agents are given driving instructions which they must learn to interpret in order to successfully navigate in this environment. Since humans equipped with driving instructions can readily navigate in previously unseen cities, we set a high bar and test our trained agents for similar cognitive capabilities. Although deep reinforcement learning (RL) methods are frequently evaluated only on data that closely follow the training distribution, our dataset extends to multiple cities and has a clean train/test separation. This allows for thorough testing of generalisation ability. This paper presents the StreetNav environment and tasks, models that establish strong baselines, and extensive analysis of the task and the trained agents. Karl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath, Keith Anderson, Raia Hadsell |
AAAI | 3 |
| 2020 | Do Digital Agents Do Dada?
Gunter Lösel, Piotr Mirowski, Kory W. Mathewson |
ICCC | 2 |
| 2020 | Rosetta Code: Improv in Any Language
Piotr Mirowski, Kory W. Mathewson, Boyd Branch, Thomas Winters, Ben Verhoeven, Jenny Elfving |
ICCC | 1 |
| 2019 | Human Improvised Theatre Augmented with Artificial IntelligenceabstractImprovisational theatre (improv) has been proposed as a grand challenge for general artificial intelligence (AI)~\citemartin2016improvisational. Current state-of-the-art conversational intelligence models lack proper grounding, language understanding, and generate meaningless meandering responses~\citedziri2018augmenting. Utilizing them as improvised comedy partners (improvisors) is doomed to fail - curiously, this limitation makes their use particularly appealing. Improv theatre celebrates risk taking and failure by inviting performers to express themselves without hesitation or fear of being judged~\citejohnstone1979impro. Our installation is an interactive improv workshop for a group of interested participants, culminating in a live public performance. Attendees are invited to observe and interact with AI-based improvisational theatre technology. The workshop is facilitated by two improv theatre professionals with a combined 30 years of experience in teaching, training, and touring. The performance features various AI tools for augmented creativity. Piotr Mirowski, Kory W. Mathewson |
Creativity & Cognition | 1 |
| 2019 | Cross-View Policy Learning for Street NavigationabstractThe ability to navigate from visual observations in unfamiliar environments is a core component of intelligent agents and an ongoing challenge for Deep Reinforcement Learning (RL). Street View can be a sensible testbed for such RL agents, because it provides real-world photographic imagery at ground level, with diverse street appearances; it has been made into an interactive environment called StreetLearn and used for research on navigation. However, goal-driven street navigation agents have not so far been able to transfer to unseen areas without extensive retraining, and relying on simulation is not a scalable solution. Since aerial images are easily and globally accessible, we propose instead to transfer a ground view policy, from training areas to unseen (target) parts of the city, by utilizing aerial view observations. Our core idea is to pair the ground view with an aerial view and to learn a joint policy that is transferable across views. We achieve this by learning a similar embedding space for both views, distilling the policy across views and dropping out visual modalities. We further reformulate the transfer learning paradigm into three stages: 1) cross-modal training, when the agent is initially trained on multiple city regions, 2) aerial view-only adaptation to a new area, when the agent is adapted to a held-out region using only the easily obtainable aerial view, and 3) ground view-only transfer, when the agent is tested on navigation tasks on unseen ground views, without aerial imagery. Our experimental results suggest that the proposed cross-view policy learning enables better generalization of the agent and allows for more effective transfer to unseen environments.The ability to navigate from visual observations in unfamiliar environments is a core component of intelligent agents and an ongoing challenge for Deep Reinforcement Learning (RL). Street View can be a sensible testbed for such RL agents, because it provides real-world photographic imagery at ground level, with diverse street appearances; it has been made into an interactive environment called StreetLearn and used for research on navigation. However, goal-driven street navigation agents have not so far been able to transfer to unseen areas without extensive retraining, and relying on simulation is not a scalable solution. Since aerial images are easily and globally accessible, we propose instead to train a multi-modal policy on ground and aerial views, then transfer the ground view policy to unseen (target) parts of the city by utilizing aerial view observations. Our core idea is to pair the ground view with an aerial view and to learn a joint policy that is transferable across views. We achieve this by learning a similar embedding space for both views, distilling the policy across views and dropping out visual modalities. We further reformulate the transfer learning paradigm into three stages: 1) cross-modal training, when the agent is initially trained on multiple city regions, 2) aerial view-only adaptation to a new area, when the agent is adapted to a held-out region using only the easily obtainable aerial view, and 3) ground view-only transfer, when the agent is tested on navigation tasks on unseen ground views, without aerial imagery. Experimental results suggest that the proposed cross-view policy learning enables better generalization of the agent and allows for more effective transfer to unseen environments. Huiyi Hu, Piotr Mirowski, Mehrdad Farajtabar |
ICCV | 3 |
| 2018 | Learning to Navigate in Cities Without a MapabstractNavigating through unstructured environments is a basic capability of intelligent creatures, and thus is of fundamental interest in the study and development of artificial intelligence. Long-range navigation is a complex cognitive task that relies on developing an internal representation of space, grounded by recognisable landmarks and robust visual processing, that can simultaneously support continuous self-localisation ("I am here") and a representation of the goal ("I am going there"). Building upon recent research that applies deep reinforcement learning to maze navigation problems, we present an end-to-end deep reinforcement learning approach that can be applied on a city scale. Recognising that successful navigation relies on integration of general policies with locale-specific knowledge, we propose a dual pathway architecture that allows locale-specific features to be encapsulated, while still enabling transfer to multiple cities. A key contribution of this paper is an interactive navigation environment that uses Google Street View for its photographic content and worldwide coverage. Our baselines demonstrate that deep reinforcement learning agents can learn to navigate in multiple cities and to traverse to target destinations that may be kilometres away. A video summarizing our research and showing the trained agent in diverse city environments as well as on the transfer task is available at: https://sites.google.com/view/learn-navigate-cities-nips18 Piotr Mirowski, Matthew Koichi Grimes, Mateusz Malinowski, Karl Moritz Hermann, Keith Anderson, Denis Teplyashin, Karen Simonyan, Koray Kavukcuoglu, Andrew Zisserman, Raia Hadsell |
NeurIPS | 1 |
| 2017 | Learning to Navigate in Complex Environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J. Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, Dharshan Kumaran, Raia Hadsell |
ICLR (Poster) | 1 |
| 2014 | Building Optimal Radio-Frequency Signal MapsabstractA popular way for using radio-frequency (RF) signals (e.g. WiFi) to position people or device indoors is by matching received radio signal strength (RSS) to fingerprints that are spatial signatures of such measures. Traditionally such signal maps are built by manual collection of repeated measurements at predefined locations following a spatial sampling scheme. Recently, such labor intensive processes are being replaced by robot-based automation or crowd-sourced simultaneous localization and mapping (SLAM). These new approaches produce time-stamped trajectories along with time-stamped RSS as the human or robot moves freely about the building. However, they require an additional procedure to segment the continuous RF samples into fingerprint cells to produce a robust signal map. In this paper, we explore several strategies for building optimal signal maps from RSS collected along robotic or pedestrian trajectories. We compare two clustering algorithms with a baseline strategy that divides the trajectories into a hierarchy of fixed-size grids. We study the trade-off between the spatial extent of the fingerprint cells and the differentiability of the RSS distribution in each cell, as well as their impact on localization accuracy and on fingerprint storage. We experimented with traces collected by an autonomous robot exploring a large multi-floor office building. Piotr Mirowski, Tin Kam Ho, Phil Whiting |
ICPR | 1 |
| 2014 | Pose Invariant Activity Classification for Multi-floor Indoor LocalizationabstractSmartphone based indoor localization caught massive interest of the localization community in recent years. Combining pedestrian dead reckoning obtained using the phone's inertial sensors with the Graph SLAM (Simultaneous Localization and Mapping) algorithm is one of the most effective approaches to reconstruct the entire pedestrian trajectory given a set of visited landmarks during movement. A key to Graph SLAM-based localization is the detection of reliable landmarks, which are typically identified using visual cues or via NFC tags or QR codes. Alternatively, human activity can be classified to detect organic landmarks such as visits to stairs and elevators while in movement. We provide a novel human activity classification framework that is invariant to the pose of the smartphone. Pose invariant features allow robust observation no matter how a user puts the phone in the pocket. In addition, activity classification obtained by an SVM (Support Vector Machine) is used in a Bayesian framework with an HMM (Hidden Markov Model) that improves the activity inference based on temporal smoothness. Furthermore, the HMM jointly infers activity and floor information, thus providing multi-floor indoor localization. Our experiments show that the proposed framework detects landmarks accurately and enables multi-floor indoor localization from the pocket using Graph SLAM. Saehoon Yi, Piotr Mirowski, Tin Kam Ho, Vladimir Pavlovic 0001 |
ICPR | 2 |
| 2013 | SignalSLAM: Simultaneous localization and mapping with mixed WiFi, Bluetooth, LTE and magnetic signalsabstractIndoor localization typically relies on measuring a collection of RF signals, such as Received Signal Strength (RSS) from WiFi, in conjunction with spatial maps of signal fingerprints. A new technology for localization could arise with the use of 4G LTE telephony small cells, with limited range but with rich signal strength information, namely Reference Signal Received Power (RSRP). In this paper, we propose to combine an ensemble of available sources of RF signals to build multi-modal signal maps that can be used for localization or for network deployment optimization. We primarily rely on Simultaneous Localization and Mapping (SLAM), which provides a solution to the challenge of building a map of observations without knowing the location of the observer. SLAM has recently been extended to incorporate signal strength from WiFi in the so-called WiFi-SLAM. In parallel to WiFi-SLAM, other localization algorithms have been developed that exploit the inertial motion sensors and a known map of either WiFi RSS or of magnetic field magnitude. In our study, we use all the measurements that can be acquired by an off-the-shelf smartphone and crowd-source the data collection from several experimenters walking freely through a building, collecting time-stamped WiFi and Bluetooth RSS, 4G LTE RSRP, magnetic field magnitude, GPS reference points when outdoors, Near-Field Communication (NFC) readings at specific landmarks and pedestrian dead reckoning based on inertial data. We resolve the location of all the users using a modified version of Graph-SLAM optimization of the users poses with a collection of absolute location and pairwise constraints that incorporates multi-modal signal similarity. We demonstrate that we can recover the user positions and thus simultaneously generate dense signal maps for each WiFi access point and 4G LTE small cell, “from the pocket”. Finally, we demonstrate the localization performance using selected single modalities, such as only WiFi and the WiFi signal maps that we generated. Piotr Mirowski, Tin Kam Ho, Saehoon Yi, Michael MacDonald |
IPIN | 1 |
| 2011 | KL-divergence kernel regression for non-Gaussian fingerprint based localizationabstractVarious methods have been developed for indoor localization using WLAN signals. Algorithms that fingerprint the Received Signal Strength Indication (RSSI) of WiFi for different locations can achieve tracking accuracies of the order of a few meters. RSSI fingerprinting suffers though from two main limitations: first, as the signal environment changes, so does the fingerprint database, which requires regular updates; second, it has been reported that, in practice, certain devices record more complex (e.g bimodal) distributions of WiFi signals, precluding algorithms based on the mean RSSI. In this article, we propose a simple methodology that takes into account the full distribution for computing similarities among fingerprints using Kullback-Leibler divergence, and that performs localization through kernel regression. Our method provides a natural way of smoothing over time and trajectories. Moreover, we propose unsupervised KL-divergence-based recalibration of the training fingerprints. Finally, we apply our method to work with histograms of WiFi connections to access points, ignoring RSSI distributions, and thus removing the need for recalibration. We demonstrate that our results outperform nearest neighbors or Kalman and Particle Filters, achieving up to 1m accuracy in office environments. We also show that our method generalizes to non-Gaussian RSSI distributions. Piotr Mirowski, Harald Steck, Phil Whiting, Ravishankar Palaniappan, Michael MacDonald, Tin Kam Ho |
IPIN | 1 |
| 2010 | Feature-rich continuous language models for speech recognitionabstractState-of-the-art probabilistic models of text such as n-grams require an exponential number of examples as the size of the context grows, a problem that is due to the discrete word representation. We propose to solve this problem by learning a continuous-valued and low-dimensional mapping of words, and base our predictions for the probabilities of the target word on non-linear dynamics of the latent space representation of the words in context window. We build on neural networks-based language models; by expressing them as energy-based models, we can further enrich the models with additional inputs such as part-of-speech tags, topic information and graphs of word similarity. We demonstrate a significantly lower perplexity on different text corpora, as well as improved word accuracy rate on speech recognition tasks, as compared to Kneser-Ney back-off n-gram-based language models. Piotr Mirowski, Sumit Chopra, Suhrid Balakrishnan, Srinivas Bangalore |
SLT | 1 |
| 2009 | Dynamic Factor Graphs for Time Series Modeling
Piotr Mirowski, Yann LeCun |
ECML/PKDD (2) | 1 |
| 2008 | Retrieving scale from quasi-stationary images
Piotr Mirowski, Daniel M. Tetzlaff |
Pattern Recognit. Lett. | 1 |
| 2007 | Time-Delay Neural Networks and Independent Component Analysis for EEG-Based Prediction of Epileptic Seizures Propagation
Piotr Mirowski, Deepak Madhavan, Yann LeCun |
AAAI | 1 |