EDBT 2026 Demo / reviewers in the wild / expert
Li Mi
dblp:129/6296
· DBLP profile ↗
17ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-4886-2430ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SemCo: Toward Semantic Coherent Visual Relationship ForecastingabstractVisual Relationship Forecasting (VRF) in video aims to anticipate relations among objects without observing future visual content. The task relies on capturing and modeling the semantic coherence in object interactions, as it underpins the evolution of events and scenes in videos. However, existing VRF datasets provide limited support for learning such coherence. Their noisy annotations fail to reflect distinct action or scene changes, and weak correlations exist between different actions and relational transitions in subject-object pairs. Furthermore, existing methods struggle to distinguish similar relationships and overfit to unchanging relationships in consecutive frames rather than reflecting the semantic coherence of object interactions. To address these challenges, we present SemCoBench, a benchmark that emphasizes semantic coherence for visual relationship forecasting in video. Based on action labels and short-term subject-object pairs, SemCoBench decomposes relationship categories and dynamics by cleaning and reorganizing video datasets to ensure predicting semantic coherence in object interactions. In addition, we propose the Semantic Coherent Transformer (SemCoFormer), which consists of a Relationship Augmented Module (RAM) and a Coherent Reasoning Module (CRM). The RAM bidirectionally enhances cross-modal features to better distinguish similar relationships, while the CRM performs sparse encoding of cross-frame relation transitions to model dynamic relational changes. The experimental results on SemCoBench demonstrate that modeling the semantic coherence is a key step toward reasonable, fine-grained, and diverse visual relationship forecasting, contributing to a more comprehensive understanding of video scenes. Our project page: https://lyao-61.github.io/ SemCo/. Yangjun Ou, Li Mi, Zhenzhong Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | VinaBench: Benchmark for Faithful and Consistent Visual NarrativesabstractVisual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge constraints used for planning the stories. In this work, we propose a new benchmark, VinaBench, to address this challenge. Our benchmark annotates the underlying commonsense and discourse constraints in visual narrative samples, offering systematic scaffolds for learning the implicit strategies of visual storytelling. Based on the incorporated narrative constraints, we further propose novel metrics to closely evaluate the consistency of generated narrative images and the alignment of generations with the input textual narrative. Our results across three generative vision models demonstrate that learning with VinaBench’s knowledge constraints effectively improves the faithfulness and cohesion of generated visual narratives.1 Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, Antoine Bosselut |
CVPR | 3 |
| 2025 | GeoExplorer: Active Geo-Localization with Curiosity-Driven ExplorationabstractActive Geo-localization (AGL) is the task of localizing a goal, represented in various modalities (e.g., aerial images, ground-level images, or text), within a predefined search area. Current methods approach AGL as a goal-reaching reinforcement learning (RL) problem with a distance-based reward. They localize the goal by implicitly learning to minimize the relative distance from it. However, when distance estimation becomes challenging or when encountering unseen targets and environments, the agent exhibits reduced robustness and generalization ability due to the less reliable exploration strategy learned during training. In this paper, we propose GeoExplorer, an AGL agent that incorporates curiosity-driven exploration through intrinsic rewards. Unlike distance-based rewards, our curiosity-driven reward is goal-agnostic, enabling robust, diverse, and contextually relevant exploration based on effective environment modeling. These capabilities have been proven through extensive experiments across four AGL benchmarks, demonstrating the effectiveness and generalization ability of GeoExplorer in diverse settings, particularly in localizing unfamiliar targets and environments. Li Mi, Manon Béchaz, Zeming Chen 0001, Antoine Bosselut, Devis Tuia |
ICCV | 1 |
| 2025 | Unsupervised Multiview UAV Image Geolocalization via Iterative RenderingabstractUnmanned Aerial Vehicle (UAV) Cross-View Geo-Localization (CVGL) poses significant challenges due to the substantial view discrepancies between oblique UAV images and overhead satellite images. Existing methods heavily rely on supervised learning with labeled datasets to extract viewpoint-invariant features for cross-view retrieval. However, these approaches are computationally expensive, prone to overfitting region-specific cues, and exhibit limited generalizability to new regions. To overcome this issue, we propose an unsupervised solution that lifts the scene representation to 3D space from UAV observations for satellite image generation, providing a robust representation against view distortion. By generating orthogonal images that closely resemble satellite views, our method reduces view discrepancies in feature representation and mitigates shortcuts in region-specific image pairing. To further align the perspective of the rendered image with the real one, we design an iterative camera pose updating mechanism that progressively modulates the rendered query image with potential satellite targets, eliminating spatial offsets relative to the reference images. Additionally, this iterative refinement strategy enhances cross-view feature invariance through view-consistent fusion across iterations. As such, our unsupervised paradigm naturally avoids the problem of region-specific overfitting, enabling generic CVGL for UAV images without feature fine-tuning or data-driven training. Experiments on the University-1652 and SUES-200 datasets demonstrate that our approach significantly improves geo-localization accuracy while maintaining robustness across diverse regions. Notably, without model fine-tuning or paired training, our method achieves competitive performance with recent supervised methods. Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Li Mi, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | HARG: Hierarchical Adaptive Reasoning Graph for Activity ParsingabstractAs a video understanding task, activity parsing aims at encompassing actions into multiple levels of activity components, including activity, sub-activity and atomic action, enabling understanding of complex video scenes within multimedia systems. Existing methods form activity parsing as a multi-task learning problem to predict multi-granular activity labels simultaneously, which ignores modeling the hierarchical structure and the fine-grained transitions of activity components at different levels. In this paper, we propose a Hierarchical Adaptive Reasoning Graph (HARG) to model the hierarchical structure (i.e., object level$\rightarrow$atomic action level$\rightarrow$activity level) dynamically and precisely. To achieve that, an object reasoning graph (ORG) and an atomic action reasoning graph (ARG) are designed to reason fine-grained information transitions between multiple actors at different levels. In addition, an adaptive segmentation module (ASM) is investigated for bridging the gap among different levels, permitting step-by-step reasoning from the object level to the atomic action level. Experimental results show our method outperforms state-of-the-art methods on two activity parsing datasets, achieving hierarchical modeling and fine-grained reasoning for activity understanding. The code is available on GitHub:https://github.com/whuoyj/HARG. Yangjun Ou, Li Mi, Zhenzhong Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceabstractAsking questions about visual environments is a crucial way for intelligent agents to understand rich multi-faceted scenes, raising the importance of Visual Question Generation (VQG) systems. Apart from being grounded to the image, existing VQG systems can use textual constraints, such as expected answers or knowledge triplets, to generate focused questions. These constraints allow VQG systems to specify the question content or leverage external commonsense knowledge that can not be obtained from the image content only. However, generating focused questions using textual constraints while enforcing a high relevance to the image content remains a challenge, as VQG systems often ignore one or both forms of grounding. In this work, we propose Contrastive Visual Question Generation (ConVQG), a method using a dual contrastive objective to discriminate questions generated using both modalities from those based on a single one. Experiments on both knowledge-aware and standard VQG benchmarks demonstrate that ConVQG outperforms the state-of-the-art methods and generates image-grounded, text-guided, and knowledge-rich questions. Our human evaluation results also show preference for ConVQG questions compared to non-contrastive baselines. Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai, Antoine Bosselut, Devis Tuia |
AAAI | 1 |
| 2024 | ConGeo: Robust Cross-View Geo-Localization Across Ground View Variations
Li Mi, Chang Xu 0027, Javiera Castillo-Navarro, Syrielle Montariol, Wen Yang 0001, Antoine Bosselut, Devis Tuia |
ECCV (14) | 1 |
| 2024 | Knowledge-Aware Visual Question Generation for Remote Sensing ImagesabstractWith the rapid development of remote sensing image archives, asking questions about images has become an effective way of gathering specific information or performing image retrieval. However, automatically generated image-based questions tend to be simplistic and template-based, which hinders the real deployment of question answering or visual dialogue systems. To enrich and diversify the questions, we propose a knowledge-aware remote sensing visual question generation model, KRSVQG, that incorporates external knowledge related to the image content to improve the quality and contextual understanding of the generated questions. The model takes an image and a related knowledge triplet from external knowledge sources as inputs and leverages image captioning as an intermediary representation to enhance the image grounding of the generated questions. To assess the performance of KRSVQG, we utilized two datasets that we manually annotated: NWPU-300 and TextRS-300. Results on these two datasets demonstrate that KRSVQG outperforms existing methods and leads to knowledge-enriched questions, grounded in both image and domain knowledge. Li Mi, Javiera Castillo-Navarro, Devis Tuia |
IGARSS | 2 |
| 2024 | Knowledge-Aware Text-Image Retrieval for Remote Sensing ImagesabstractImage-based retrieval in large Earth observation archives is challenging because one needs to navigate across thousands of candidate matches only with the query image as a guide. By using text as information supporting the visual query, the retrieval system gains in usability, but at the same time faces difficulties due to the diversity of visual signals that cannot be summarized by a short caption only. For this reason, as a matching-based task, cross-modal text–image retrieval often suffers from information asymmetry between text and images. To address this challenge, we propose a Knowledge-aware Text–Image Retrieval (KTIR) method for remote sensing images. By mining relevant information from an external knowledge graph, KTIR enriches the text scope available in the search query and alleviates the information gaps between text and images for better matching. Moreover, by integrating domain-specific knowledge, KTIR also enhances the adaptation of pretrained vision–language models to remote sensing applications. Experimental results on three commonly used remote sensing text–image retrieval benchmarks show that the proposed knowledge-aware method leads to varied and consistent retrievals, outperforming state-of-the-art retrieval methods. Li Mi, Xianjie Dai, Javiera Castillo-Navarro, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Object-Relation Reasoning Graph for Action RecognitionabstractAction recognition is a challenging task since the attributes of objects as well as their relationships change constantly in the video. Existing methods mainly use object-level graphs or scene graphs to represent the dynamics of objects and relationships, but ignore modeling the fine-grained relationship transitions directly. In this paper, we propose an Object-Relation Reasoning Graph (OR2G) for reasoning about action in videos. By combining an object-level graph (OG) and a relation-level graph (RG), the proposed OR2G catches the attribute transitions of objects and reasons about the relationship transitions between objects simultaneously. In addition, a graph aggregating module (GAM) is investigated by applying the multi-head edge-to-node message passing operation. GAM feeds back the information from the relation node to the object node and enhances the coupling between the object-level graph and the relation-level graph. Experiments in video action recognition demonstrate the effectiveness of our approach when compared with the state-of-the-art methods. Yangjun Ou, Li Mi, Zhenzhong Chen 0001 |
CVPR | 2 |
| 2022 | Effects of Lossy Compression on Remote Sensing Image Classification Based on Convolutional Sparse CodingabstractLossy compression causes the degradation of the classification accuracy of remote sensing (RS) images due to the introduced distortion by compression. In this letter, a convolutional sparse coding (CSC)-based method is proposed to quantitatively measure such an effect. In detail, the filters used in CSC are learned by online convolutional dictionary learning (OCDL) to construct the dictionary. Thereafter, the sparse coefficient maps are obtained based on the alternating direction method of multipliers (ADMM) algorithm. In addition, multiple kernel learning (MKL) is used to estimate the corresponding classification accuracy. The experimental results demonstrate that our method performs better in predicting the classification accuracy of RS images compared with the other state-of-the-art algorithms. Jingru Wei, Li Mi, Jing Ling, Zhenzhong Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Predicate Correlation Learning for Scene Graph GenerationabstractFor a typical Scene Graph Generation (SGG) method in image understanding, there usually exists a large gap in the performance of the predicates’ head classes and tail classes. This phenomenon is mainly caused by the semantic overlap between different predicates as well as the long-tailed data distribution. In this paper, a Predicate Correlation Learning (PCL) method for SGG is proposed to address the above problems by taking the correlation between predicates into consideration. To measure the semantic overlap between highly correlated predicate classes, a Predicate Correlation Matrix (PCM) is defined to quantify the relationship between predicate pairs, which is dynamically updated to remove the matrix’s long-tailed bias. In addition, PCM is integrated into a predicate correlation loss function (LPC) to reduce discouraging gradients of unannotated classes. The proposed method is evaluated on several benchmarks, where the performance of the tail classes is significantly improved when built on existing methods. Leitian Tao, Li Mi, Nannan Li 0004, Xianhang Cheng, Yaosi Hu, Zhenzhong Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Analysis on Medication Rules of Guide to Clinical Medical Records from the Prescriptions Containing Poria Cocos Based on Data MiningabstractAt present, there have been numerous studies in Ye Tianshi and his book guide to clinical medical records. Poria Cocos is the most commonly used traditional Chinese medicine in guide to clinical medical records and the most commonly used traditional Chinese medicine by Ye Tianshi in clinical practice. However, there is not any study on the law of using Poria Cocos in Ye Tianshi. In order to solve the problem, data mining methods such as association rules and complex network analysis are used in this study. Association rules and complex network analysis were used to analyze the prescription that containing Poria Cocos in guide to clinical medical records, and the frequency of single traditional Chinese medicine, the frequency of combination of traditional Chinese medicine, the distribution of core traditional Chinese medicine, and indications were obtained. The research will provide reference to the further study of Ye Tianshi’s clinical thought and Poria Cocos. Mei Zhao, Li Mi |
BIBM | 2 |
| 2021 | AFNet: Adaptive Fusion Network for Remote Sensing Image Semantic SegmentationabstractSemantic segmentation of remote sensing images plays an important role in many applications. However, a remote sensing image typically comprises a complex and heterogenous urban landscape with objects in various sizes and materials, which causes challenges to the task. In this work, a novel adaptive fusion network (AFNet) is proposed to improve the performance of very high resolution (VHR) remote sensing image segmentation. To coherently label size-varied ground objects from different categories, we design multilevel architecture with the scale-feature attention module (SFAM). By SFAM, at the location of small objects, low-level features from the shallow layers of convolutional neural network (CNN) are enhanced, whilst for large objects, high-level features from deep layers are enhanced. Thus, the features of size-varied objects could be preserved during fusing features from different levels, which helps to label size-varied objects. As for labeling the category with high intra-class difference and varied scales, the multiscale structure with a scale-layer attention module (SLAM) is utilized to learn representative features, where an adjacent score map refinement module (ACSR) is employed as the classifier. By SLAM, when fusing multiscale features, based on the interested objects scale, feature map from appropriate scale is given greater weights. With such a scale-aware strategy, the learned features can be more representative, which is helpful to distinguish objects for semantic segmentation. Besides, the performance is further improved by introducing several nonlinear layers to the ACSR. Extensive experiments conducted on two well-known public high-resolution remote sensing image data sets show the effectiveness of our proposed model. Code and predictions are available at https://github.com/athauna/AFNet/ Li Mi, Zhenzhong Chen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Hierarchical Graph Attention Network for Visual Relationship DetectionabstractVisual Relationship Detection (VRD) aims to describe the relationship between two objects by providing a structural triplet shown as. Existing graph-based methods mainly represent the relationships by an object-level graph, which ignores to model the triplet-level dependencies. In this work, a Hierarchical Graph Attention Network (HGAT) is proposed to capture the dependencies on both object-level and triplet-level. Object-level graph aims to capture the interactions between objects, while the triplet-level graph models the dependencies among relation triplets. In addition, prior knowledge and attention mechanism are introduced to fix the redundant or missing edges on graphs that are constructed according to spatial correlation. With these approaches, nodes are allowed to attend over their spatial and semantic neighborhoods' features based on the visual or semantic feature correlation. Experimental results on the well-known VG and VRD datasets demonstrate that our model significantly outperforms the state-of-the-art methods. Li Mi |
CVPR | 1 |
| 2020 | Exploiting the local temporal information for video captioning
Li Mi, Yaosi Hu, Zhenzhong Chen 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Scene Context-Driven Vehicle Detection in High-Resolution Aerial ImagesabstractAs the spatial resolution of remote sensing images is improving gradually, it is feasible to realize “scene-object” collaborative image interpretation. Unfortunately, this idea is not fully utilized in vehicle detection from high-resolution aerial images, and most of the existing methods may be promoted by considering the variability of vehicle spatial distribution in different image scenes and treating vehicle detection tasks scene-specific. With this motivation, a scene context-driven vehicle detection method is proposed in this paper. At first, we perform scene classification using the deep learning method and, then, detect vehicles in roads and parking lots separately through different vehicle detectors. Afterward, we further optimize the detection results using different postprocessing rules according to different scene types. Experimental results show that the proposed approach outperforms the state-of-the-art algorithms in terms of higher detection accuracy rate and lower false alarm rate. Chao Tao 0001, Li Mi, Yansheng Li 0001, Ji Qi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |