Yizi Chen

dblp:250/4440 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-1637-0092ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Sketch2Terrain: AI-Driven Real-Time Terrain Sketch Mapping in Augmented Reality
Tianyi Xiao, Yizi Chen, Sailin Zhong, Peter Kiefer, Jakub Krukar, Kevin Gonyop Kim, Lorenz Hurni, Angela Schwering, Martin Raubal
CHI2
2025 Unsupervised Urban Land Use Mapping with Street View Contrastive Clustering and a Geographical Prior
abstract
Urban land use classification and mapping are critical for urban planning, resource management, and environmental monitoring. Existing remote sensing techniques often lack precision in complex urban environments due to the absence of ground-level details. Unlike aerial perspectives, street view images provide a ground-level view that captures more human and social activities relevant to land use in complex urban scenes. Existing street view-based methods primarily rely on supervised classification, which is challenged by the scarcity of high-quality labeled data and the difficulty of generalizing across diverse urban landscapes. This study introduces an unsupervised contrastive clustering model for street view images with a built-in geographical prior, to enhance clustering performance. When combined with a simple visual assignment of the clusters, our approach offers a flexible and customizable solution to land use mapping, tailored to the specific needs of urban planners. We experimentally show that our method can generate land use maps from geotagged street view image datasets of two cities. As our methodology relies on the universal spatial coherence of geospatial data ("Tobler's law"), it can be adapted to various settings where street view images are available, to enable scalable, unsupervised land use mapping and updating. The code is available at https://github.com/lin102/CCGP.
Lin Che 0001, Yizi Chen, Tanhua Jin, Martin Raubal, Konrad Schindler, Peter Kiefer
SIGSPATIAL/GIS2
2025 Generative AI in Map-Making: A Technical Exploration and Its Implications for Cartographers
abstract
Traditional map-making relies heavily on Geographic Information Systems (GIS), requiring domain expertise and being time-consuming, especially for repetitive tasks. Recent advances in generative AI (GenAI), particularly image diffusion models, offer new opportunities for automating and democratizing the map-making process. However, these models struggle with accurate map creation due to limited control over spatial composition and semantic layout. To address this, we integrate vector data to guide map generation in different styles, specified by the textual prompts. Our model is the first to generate accurate maps in controlled styles, and we have integrated it into a web application to improve its usability and accessibility. We conducted a user study with professional cartographers to assess the fidelity of generated maps, the usability of the web application, and the implications of ever-emerging GenAI in map-making. The findings have suggested the potential of our developed application and, more generally, the GenAI models in helping both non-expert users and professionals in creating maps more efficiently. We have also outlined further technical improvements and emphasized the new role of cartographers to advance the paradigm of AI-assisted map-making.
Claudio Affolter, Sidi Wu 0001, Yizi Chen, Lorenz Hurni
SIGSPATIAL/GIS3
2025 Integrating With Multimodal Information for Enhancing Robotic Grasping With Vision-Language Models
abstract
As robots grow increasingly intelligent and utilize data from various sensors, relying solely on unimodal data sources is becoming inadequate for their operational needs. Consequently, integrating multimodal data has emerged as a critical area of focus. However, the effective combination of different data modalities poses a considerable challenge, especially in complex and dynamic settings where accurate object recognition and manipulation are essential. In this paper, we introduce a novel framework integrating with Multimodal Information for Grasping Synthesis with vision-language models (MIG) designed to improve robotic grasping capabilities. This framework incorporates visual data, textual information, and human-derived prior knowledge. We start by creating target object masks based on this prior knowledge, which are then used to segregate the target objects from their surroundings in the image. Subsequently, we employ language cues to refine the visual representations of these objects. Finally, our system executes precise grasping actions using visual and textual data synthesis, thus facilitating more effective and contextually aware robotic grasping. We carry out experiments using the OCID-VLG dataset. We observe that our methodology surpasses current state-of-the-art (SOTA) techniques, delivering improvements of 9.91% and 5.70% for top-1 and top-5 predictions in grasp accuracy. Moreover, when apply to the reconstructed Grasp-MultiObject dataset, our approach demonstrates even more substantial enhancements, achieving gains of 17.63% and 22.76% over SOTA methods for top-1 and top-5 predictions, respectively. Note to Practitioners—As robotic systems evolve, the challenge of enabling them to function effectively in complex environments has become increasingly apparent. This paper introduces a solution that integrates multiple sources of data—visual, textual, and human knowledge—to enhance robotic grasping capabilities. The practical problems addressed include the limitations of current unimodal systems that struggle with accurate object recognition and manipulation in dynamic settings, such as warehouses or assembly lines. Our framework, MIG, demonstrates significant improvements in grasp accuracy, making it suitable for tasks where precision is critical. While our results show promise, particularly in controlled experiments, there are limitations to consider. The framework’s performance may vary in unstructured real-world environments due to factors like occlusion or varying lighting conditions. Future work should focus on refining the system for real-time application and exploring additional sensory inputs to enhance robustness. By addressing these challenges, we aim to make this approach more applicable across industries, paving the way for smarter, more adaptable robotic solutions in everyday tasks.
Dongyuan Zheng, Yizi Chen, Jing Luo 0005, Panfeng Huang, Chenguang Yang 0001
IEEE Trans Autom. Sci. Eng.3
2024 StegoGAN: Leveraging Steganography for Non-Bijective Image-to-Image Translation
abstract
Most image-to-image translation models postulate that a unique correspondence exists between the semantic classes of the source and target domains. However, this assumption does not always hold in real-world scenarios due to divergent distributions, different class sets, and asymmet- rical information representation. As conventional GANs attempt to generate images that match the distribution of the target domain, they may hallucinate spurious instances of classes absent from the source domain, thereby dimin- ishing the usefulness and reliability of translated images. CycleGAN-based methods are also known to hide the mis- matched information in the generated images to bypass cy- cle consistency objectives, a process known as steganogra- phy. In response to the challenge of non-bijective image translation, we introduce StegoGAN, a novel model that leverages steganography to prevent spurious features in generated images. Our approach enhances the semantic consistency of the translated images without requiring ad- ditional postprocessing or supervision. Our experimental evaluations demonstrate that StegoGAN outperforms existing GAN-based models across various non-bijective image- to-image translation tasks, both qualitatively and quantita- tively. Our code and pretrained models are accessible at https://github.com/sian-wusidi/StegoGAN.
Sidi Wu 0001, Yizi Chen, Samuel Mermet, Lorenz Hurni, Konrad Schindler, Nicolas Gonthier, Loïc Landrieu
CVPR2
2023 Cross-attention Spatio-temporal Context Transformer for Semantic Segmentation of Historical Maps
abstract
Historical maps provide useful spatio-temporal information on the Earth's surface before modern earth observation techniques came into being. To extract information from maps, neural networks, which gain wide popularity in recent years, have replaced hand-crafted map processing methods and tedious manual labor. However, aleatoric uncertainty, known as data-dependent uncertainty, inherent in the drawing/scanning/fading defects of the original map sheets and inadequate contexts when cropping maps into small tiles considering the memory limits of the training process, challenges the model to make correct predictions. As aleatoric uncertainty cannot be reduced even with more training data collected, we argue that complementary spatio-temporal contexts can be helpful. To achieve this, we propose a U-Net-based network that fuses spatio-temporal features with cross-attention transformers (U-SpaTem), aggregating information at a larger spatial range as well as through a temporal sequence of images. Our model achieves a better performance than other state-or-art models that use either temporal or spatial contexts. Compared with pure vision transformers, our model is more lightweight and effective. To the best of our knowledge, leveraging both spatial and temporal contexts have been rarely explored before in the segmentation task. Even though our application is on segmenting historical maps, we believe that the method can be transferred into other fields with similar problems like temporal sequences of satellite images. Our code is freely accessible at https://github.com/chenyizi086/wu.2023.sigspatial.git.
Sidi Wu 0001, Yizi Chen, Konrad Schindler, Lorenz Hurni
SIGSPATIAL/GIS2
2021 Introducing the Boundary-Aware loss for deep image segmentation
Minh On Vu Ngoc, Yizi Chen, Nicolas Boutry, Joseph Chazalon, Edwin Carlinet, Clément Mallet, Thierry Géraud
BMVC2
2021 ICDAR 2021 Competition on Historical Map Segmentation
Joseph Chazalon, Edwin Carlinet, Yizi Chen, Julien Perret, Bertrand Dumenieu, Clément Mallet, Thierry Géraud, Vincent Nguyen 0001, Josef Baloun, Ladislav Lenc, Pavel Král
ICDAR (4)3
2021 Vectorization of Historical Maps Using Deep Edge Filtering and Closed Shape Extraction
Yizi Chen, Edwin Carlinet, Joseph Chazalon, Clément Mallet, Bertrand Dumenieu, Julien Perret
ICDAR (4)1
2019 Enhanced Pix2pix Dehazing Network
abstract
In this paper, we reduce the image dehazing problem to an image-to-image translation problem, and propose Enhanced Pix2pix Dehazing Network (EPDN), which generates a haze-free image without relying on the physical scattering model. EPDN is embedded by a generative adversarial network, which is followed by a well-designed enhancer. Inspired by visual perception global-first theory, the discriminator guides the generator to create a pseudo realistic image on a coarse scale, while the enhancer following the generator is required to produce a realistic dehazing image on the fine scale. The enhancer contains two enhancing blocks based on the receptive field model, which reinforces the dehazing effect in both color and details. The embedded GAN is jointly trained with the enhancer. Extensive experiment results on synthetic datasets and real-world datasets show that the proposed EPDN is superior to the state-of-the-art methods in terms of PSNR, SSIM, PI, and subjective visual effect.
Yanyun Qu, Yizi Chen, Jingying Huang, Yuan Xie 0006
CVPR2