EDBT 2026 Demo / reviewers in the wild / expert
Guoming Wang
dblp:22/6650
· DBLP profile ↗
20ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-3131-6916ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Theory of computation · 2Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolving Generalist Virtual Agents with Generative and Associative MemoryabstractGeneralist Virtual Agents (GVAs) powered by Multimodal Large Language Models (MLLMs) exhibit impressive capabilities. However, their long-term learning is hampered by a core limitation: a failure to evolve beyond existing trajectories. This stems from memory systems that treat experiences as isolated fragments and rely on brittle semantic retrieval, preventing the synthesis of novel solutions from disparate knowledge. To address this, we introduce CA3Mem, a framework inspired by the human hippocampus that organizes experiences into a structured memory graph. Leveraging this graph, CA3Mem features two key innovations: 1) a generative memory recombination mechanism that synthesizes novel solutions to drive agent evolution, and 2) an associative retrieval algorithm that employs spreading activation to recall a comprehensive and contextually-aware set of experiences. Experiments on OSWorld and WebArena demonstrate that CA3Mem significantly enhances agent capabilities, leading to marked improvements in long-horizon planning, compositional generalization for novel tasks, and continuous adaptation from experience. Zhenkui Zhang, Wendong Bu, Kaihang Pan, Bingchen Miao, Wenqiao Zhang, Guoming Wang, Wei Ji 0008, Juncheng Li 0006, Siliang Tang |
AAAI | 6 |
| 2025 | Choice is what matters after AttentionabstractThe decoding strategies widely used in large language models (LLMs) today are Top-$p$ Sampling and Top-$k$ Sampling, both of which are methods situated between greedy decoding and random sampling. Inspired by the concept of loss aversion from prospect theory in behavioral economics, and the endowment effect as highlighted by Richard H. Thaler, the 2017 Nobel Memorial Prize in Economic Sciences — particularly the principle that "the negative utility of an equivalent loss is approximately twice the positive utility of a comparable gain" — we have developed a new decoding strategy called Loss Sampling. We have demonstrated the effectiveness and validity of our method on several LLMs, including Llama-2, Llama-3 and Mistral. Our approach improves text quality by 4-30% across four pure text tasks while maintaining diversity in text generation. Furthermore, we also extend our method to multimodal large models (LMs) and Beam Search, demonstrating the effectiveness and versatility of Loss Sampling with improvements ranging from 1-10%. Chenhan Fu, Guoming Wang, Juncheng Li 0006, Rongxing Lu, Siliang Tang |
AISTATS | 2 |
| 2025 | ITERATE: Image-Text Enhancement, Retrieval, and Alignment for Transmodal Evolution with LLMsabstractInspired by human cognitive behavior, we introduce visual modality to enhance the performance of pure text-based question-answering tasks with the development of multimodal models. However, obtaining corresponding images through manual annotation often entails high costs. Faced with this challenge, an intuitive strategy is to use search engines or use web scraping techniques to automatically obtain relevant image information. However, the images obtained by this strategy may be of low quality and may not match the context of the original task, which could fail to improve or even decrease performance on downstream tasks. In this paper, we propose a novel framework named “ITERATE”, aimed at retrieving and optimizing the quality of images to improve the alignment between text and images. Inspired by evolutionary algorithms in reinforcement learning and driven by the synergy of large language models (LLMs) and multimodal models, ITERATE employs a series of strategic actions such as filtering, optimizing, and retrieving to acquire higher quality images, and repeats this process over multiple generations to enhance the quality of the entire image cluster. Our experimental results on the ScienceQA, ARC-Easy, and OpenDataEval datasets also verify the effectiveness of our method, showing improvements of 3.5%, 5%, and 7%, respectively. Chenhan Fu, Guoming Wang, Juncheng Li 0006, Wenqiao Zhang, Rongxing Lu, Siliang Tang |
COLING | 2 |
| 2025 | Global Discovery: A Global Graph-RAG Approach for Query-Focused Multimodal Summarization Across Multiple PDF Papers
Chenhan Fu, Guoming Wang, Rongxing Lu, Siliang Tang |
KSEM (5) | 2 |
| 2025 | MedQuery: A Graph-Driven Medical Literature-Enhanced Query Answering SystemabstractIn the fields of medicine and science, the volume of specialized literature has grown exponentially, containing vast multimodal data-text, images, and tables-that is essential for conveying in-depth scientific insights. However, effectively retrieving, processing, and answering high-level, complex queries from this data remains a significant challenge. In this study, we introduce MedQuery, a multimodal medical knowledge query-answering system that integrates query-based literature retrieval with a response generation module capable of reasoning across multimodal data. Our system begins with literature retrieval from PubMed, using keyword extraction and query alignment to improve document accuracy and relevance. Next, the multimodal processing module processes source documents, extracting images, tables, and text, converting all into a unified textual format. This data is then structured into a global graph capturing relationships among document elements, allowing our system to support a more integrated, in-depth understanding of complex medical queries beyond basic fact retrieval. Extensive evaluations across multiple datasets, including PubMedQA, PubMed-Summarization, and our own MedInquiry dataset, demonstrate that MedQuery outperforms traditional methods and existing commercial AI systems, achieving around 90% win rates in answer quality and a 13-36% improvement in accuracy. Chenhan Fu, Yu Xia 0028, Guoming Wang, Rongxing Lu, Siliang Tang |
ICMR | 3 |
| 2025 | LLAUS: A High-Quality Instruction-Tuned Large Vision Language Assistant for UltraSoundabstractIn recent years, multimodal large models in the medical field have garnered widespread attention. However, this focus has primarily been on CT and MRI imaging, inadvertently neglecting the needs of economically underdeveloped regions and specific populations, such as pregnant women. These groups are often unable to utilize CT and MRI due to their prohibitive costs and potential harm to the body. Meanwhile, ultrasound, an economically viable and very low side effects medical imaging technique, has been largely overlooked by researchers. This study introduces a high-quality instruction-tuned Large vision Language Assistant for UltraSound (LLAUS), designed to answer questions about medical ultrasound images, aiming to assist clinicians in impoverished areas to improve the provision of healthcare services. To address the challenge of missing high-quality ultrasound data, we propose the Adaptive Caption Enhancement(ACE) and Adaptive Caption Optimization (ACO) strategies and have developed a high-quality instruction-following dataset. Subsequently, we fine-tune a Large Vision-Language Model (LVLM) using a novel Zoom-In method. By training on high-quality instruction-following datas, LLAUS demonstrates exceptional multimodal ultrasound communication capabilities, assisting in querying ultrasound images based on open-ended instructions. On tasks related to question-answering and caption generation for ultrasound images, LLAUS exhibits strong performance. Junhao Guo, XueFeng Shan, Guoming Wang, Dong Chen 0017, Rongxing Lu, Siliang Tang |
ICMR | 3 |
| 2025 | MedAI Hub: A Multimodal Medical Data Platform with Evolutionary Image Enhancement and Graph-Driven Literature RetrievalabstractWe present MedAI Hub, an integrated multimodal medical platform designed to bridge clinical practice and research by transforming patient-doctor interactions into structured scientific data. This platform supports comprehensive management of multimodal medical records-including clinical notes, medical images, and patient-reported outcomes-while implementing privacy-preserving data sharing mechanisms. Building upon this infrastructure, we introduce two novel AI-driven modules:(1) ITERATE (Image-Text Enhancement, Retrieval, and Alignment): An evolutionary algorithm inspired by Visual Genome that optimizes medical image-text alignment through iterative cross-modal refinement. Leveraging LLM-guided ''DNA evolution'' and multimodal feedback, ITERATE enhances ultrasound image quality for diagnostic tasks, achieving 3.5-7% accuracy gains on ScienceQA and ARC-Easy benchmarks.(2) MedQuery: A graph-driven literature retrieval system that constructs multimodal knowledge graphs from medical literature (text, figures, tables). By aligning PubMed documents with complex clinical queries through semantic relationship modeling, it achieves >90% answer quality win rates and 13-36% accuracy improvements on PubMedQA and MedInquiry datasets. MedAI Hub demonstrates that synergistic integration of clinical data platforms with evolutionary vision-language optimization and multimodal knowledge graphs significantly advances medical AI capabilities, enabling more accurate diagnostics and research insights. The platform and algorithms are publicly available to accelerate innovation in medical AI. Guoming Wang |
ACM Multimedia | 1 |
| 2025 | MM-CARP: Multimodal Model with Cross-Modal Retrieval-Augmented and Visual Region Perception
Junhao Guo, Chenhan Fu, Guoming Wang, Rongxing Lu, Dong Chen 0017, Siliang Tang |
MMM (2) | 3 |
| 2025 | Combining Non-numerical Text and Numerical Sequences in LLM-Based Survival Prediction
Guoqing Qian, Guoming Wang, Rongxing Lu, Siliang Tang |
PRICAI | 4 |
| 2025 | SkyNet: An Extensible Edge-Cloud Collaborative Framework for Robots in Long-Horizon TasksabstractLarge language models (LLMs) have shown promise in empowering robotics, but their widespread real-world application remains challenging due to two main issues: (1) existing research is "out-of-the-box", struggling to generalize to new robots and tasks, especially long-horizon tasks, and (2) deploying more general and powerful LLMs exceeds the capabilities of commodity hardware. To address these challenges, we propose the edge-cloud collaborative framework, i.e., SkyNet. We deploy LLMs in the cloud to create initial plans, select executable skills from a predefined library, and send them to the edge-based robot. The robot integrates multiple modules to form a policy network to complete the skills and update the feedback history. Based on the feedback history, the cloud LLMs determine whether to replan. The edge-cloud collaborative approach alleviates the pressure of deploying LLMs on commodity hardware, while the modular design enables easy extension to different tasks or robots without reconfiguring everything. To address the lack of standardized real-world experimental setups, we set up two easily replicable long-horizon tasks on a mobile robot equipped with commodity hardware, analyze the performance of various modules, and demonstrated the effectiveness of SkyNet. Guoming Wang, Rongxing Lu, Siliang Tang |
SMC | 2 |
| 2025 | Improving Vision Anomaly Detection With the Guidance of Language ModalityabstractRecent years have seen a surge of interest in anomaly detection. However, existing unsupervised anomaly detectors, particularly those for the vision modality, face significant challenges due to redundant information and sparse latent space. In contrast, anomaly detectors demonstrate superior performance in the language modality due to the unimodal nature of the data. This paper tackles the aforementioned challenges for vision modality from a multimodal point of view. Specifically, we propose Cross-modal Guidance (CMG), comprising of Cross-modal Entropy Reduction (CMER) and Cross-modal Linear Embedding (CMLE), to address the issues of redundant information and sparse latent space, respectively. CMER involves masking portions of the raw image and computing the matching score with the corresponding text. Essentially, CMER eliminates irrelevant pixels to direct the detector's focus towards critical content. To learn a more compact latent space for the vision anomaly detection, CMLE learns a correlation structure matrix from the language modality. Then, the acquired matrix compels the distribution of images to resemble that of texts in the latent space. Extensive experiments demonstrate the effectiveness of the proposed methods. Particularly, compared to the baseline that only utilizes images, the performance of CMG has been improved by 16.81%. Ablation experiments further confirm the synergy among the proposed CMER and CMLE, as each component depends on the other to achieve optimal performance. Dong Chen 0017, Kaihang Pan, Guangyu Dai, Guoming Wang, Yueting Zhuang, Siliang Tang, Mingliang Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | DIEM: Decomposition-Integration Enhancing Multimodal InsightsabstractIn image question answering, due to the abundant and sometimes redundant information, precisely matching and integrating the information from both text and images is a challenge. In this paper, we propose the Decomposition-Integration Enhancing Multimodal Insight (DIEM) which initially decomposes the given question and image into multiple subquestions and several sub-images aiming to isolate specific elements for more focused analysis. We then in-tegrate these sub-elements by matching each subquestion with its relevant sub-images, while also retaining the original image, to construct a comprehensive answer to the original question without losing sight of the overall context. This strategy mirrors the human cognitive process of simplifying complex problems into smaller components for individual analysis, followed by an integration of these insights. We implement DIEM on the LLaVA-v1.5 model, and evaluate its performance on ScienceQA and MM-Vet. Ex-perimental results indicate that our method boosts accu-racy in most question classes of the ScienceQA (+2.03% in average), especially in the image modality (+3.40%). On MM-Vet, our method achieves an improvement in MM-Vet scores, increasing from 31.1 to 32.4. These findings high-light DIEM's effectiveness in harmonizing the complexities of multimodal data, demonstrating its ability to enhance accuracy and depth in image question answering through its decomposition-integration process. Guoming Wang, Junhao Guo, Juncheng Li 0006, Wenqiao Zhang, Rongxing Lu, Siliang Tang |
CVPR | 2 |
| 2024 | De-fine: Decomposing and Refining Visual Programs with Auto-FeedbackabstractVisual programming, a modular paradigm, integrates different modules and Python operators to solve various vision-language tasks. Unlike end-to-end models that need task-specific data, it performs visual processing and inference in an unsupervised manner. Current visual programming methods generate programs in a single pass where the ability to evaluate and optimize based on feedback, unfortunately, is lacking, which consequentially limits their effectiveness for complex, multi-step problems. Inspired by benders decomposition, we introduce De-fine, a training-free framework that automatically decomposes complex tasks into simpler subtasks and refines programs through auto-feedback. This model-agnostic approach can improve logical reasoning performance by integrating the strengths of multiple models. Our experiments across various visual tasks show that De-fine creates more robust programs. Moreover, viewing each feedback module as an independent agent will yield fresh prospects for the field of agent research. Minghe Gao, Juncheng Li 0006, Hao Fei 0001, Liang Pang 0001, Wei Ji 0008, Guoming Wang, Zheqi Lv, Wenqiao Zhang, Siliang Tang, Yueting Zhuang |
ACM Multimedia | 6 |
| 2024 | WorldGPT: Empowering LLM as Multimodal World ModelabstractWorld models are progressively being employed across diverse fields, extending from basic environment simulation to complex scenario construction. However, existing models are mainly trained on domain-specific states and actions, and confined to single-modality state representations. In this paper, We introduce WorldGPT, a generalist world model built upon Multimodal Large Language Model (MLLM). WorldGPT acquires an understanding of world dynamics through analyzing millions of videos across various domains. To further enhance WorldGPT's capability in specialized scenarios and long-term tasks, we have integrated it with a novel cognitive architecture that combines memory offloading, knowledge retrieval, and context reflection. As for evaluation, we build WorldNet, a multimodal state transition prediction benchmark encompassing varied real-life scenarios. Conducting evaluations on WorldNet directly demonstrates WorldGPT's capability to accurately model state transition patterns, affirming its effectiveness in understanding and predicting the dynamics of complex scenarios. We further explore WorldGPT's emerging potential in serving as a world simulator, helping multimodal agents generalize to unfamiliar domains through efficiently synthesising multimodal instruction instances which are proved to be as reliable as authentic data for fine-tuning purposes. The code and dataset are available on the https://github.com/DCDmllm/WorldGPT Zhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li 0006, Guoming Wang, Siliang Tang, Yueting Zhuang |
ACM Multimedia | 5 |
| 2023 | FedAA: Using Non-sensitive Modalities to Improve Federated Learning while Preserving Image PrivacyabstractFederated learning aims to train a better global model without sharing the sensitive training samples (usually images) of local clients. Since the sample distributions in local clients tend to be different from each other (i.e., non-IID), one of the major challenges for federated learning is to alleviate model degradation when aggregating local models. The degradation can be attributed to the weight divergence that quantifies the difference of local models from different training processes. Furthermore, non-IID also results in feature space heterogeneity during local training, making neurons of local models in the same location have different functions and further exacerbating weight divergence. In this paper, we demonstrate that the problem can be solved by sharing information from the non-sensitive modality (e.g., metadata, non-sensitive descriptions, etc.) while keeping the sensitive information of images protected. In particular, we propose Federated Learning with Adversarial Example and Adversarial Identifier (FedAA) that trains adversarial examples based on the shared non-sensitive modality to fine-tune local models before global aggregation. The training of local models is enhanced by client identifiers that discriminate the source of inputs to force different local models to get similar outputs and be more homogeneous during the local training. Experiments show that FedAA significantly outperforms recent non-IID federated learning algorithms while preserving image privac, by sharing information from non-sensitive modalities. Dong Chen 0017, Siliang Tang, Zijin Shen, Guoming Wang, Jun Xiao 0001, Yueting Zhuang, Carl Yang 0001 |
ACM Multimedia | 4 |
| 2023 | Cross-Modal Data Augmentation for Tasks of Different ModalitiesabstractData augmentation has become one of the keys to alleviating the over-fitting of models on training data and improving the generalization capabilities on testing data. Most existing data augmentation methods only focus on one modality, which is incapable when facing multiple data modalities. Some prior works try to interpolate with random coefficients in the latent space to generate new samples, which can generically work for any data modality. However, these works ignore the extra information conveyed by multimodality data. In fact, the extra information in one modality can provide semantic directions to generate more meaningful samples in another modality. This paper proposes Cross-modal Data Augmentation (CMDA), a simple yet effective data augmentation method to alleviate the over-fitting issue and improve the generalization performance. We evaluate CMDA on unsupervised and supervised tasks of different modalities, on which CMDA consistently and significantly outperforms baselines. For instance, CMDA improves the unsupervised anomaly detection baseline in vision modality from the AUROC$76.46\%, 73.07\%$and 64.36% to$83.25\%, 76.22\%$and 70.57% on three different datasets, respectively. Besides, extensive experiments demonstrate that CMDA is applicable to various neural network architectures. Furthermore, prior methods that interpolate in the latent space need to work with downstream tasks to construct the latent space. In contrast, CMDA can work with or without downstream tasks, which makes the applicability of CMDA more extensive. The source code is publicly available for non-commercial or research use athttps://github.com/Anfeather/CMDA Dong Chen 0017, Yueting Zhuang, Zijin Shen, Carl Yang 0001, Guoming Wang, Siliang Tang, Yi Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2019 | An efficient and privacy-Preserving pre-clinical guide scheme for mobile eHealthcare
Guoming Wang, Rongxing Lu, Cheng Huang 0001, Yong Liang Guan 0001 |
J. Inf. Secur. Appl. | 1 |
| 2015 | PGuide: An Efficient and Privacy-Preserving Smartphone-Based Pre-Clinical Guidance SchemeabstractWith the pervasiveness of smartphones, mobile e-Healthcare has attracted considerable attention in recent years. Disease risk prediction, as it can assist in predicting user's disease with big data analytics techniques, has become one of important topics in the field of e-Healthcare. However, if the privacy issue is not well addressed, disease risk predication cannot step into its flourish. Aiming at addressing this challenge, in this paper, we propose a new efficient and privacy- preserving pre-clinical guidance scheme, called PGuide, which offers self-diagnosis service to medical users in a privacy-preserving way. In specific, to motivate medical users to provide more detailed health profile for accurate disease risk prediction, we introduce a privacy-preserving comparison protocol PPCP in the PGuide scheme. As a result, with enough health profile information offered by the medical users, the accuracy of disease risk prediction can be improved. Detailed security analysis shows that our proposed PGuide scheme ensures the privacy-preservation for both medical users and service provider. In addition, the performance evaluation via extensive experiments also demonstrates that our proposed PPCP protocol is much efficient in terms of low computational cost and communication overhead. Guoming Wang, Rongxing Lu, Cheng Huang 0001 |
GLOBECOM | 1 |
| 2013 | Shared Randomness and Quantum Communication in the Multi-party ModelabstractWe study shared randomness in the context of multi-party number-in-hand communication protocols in the simultaneous message passing model. We show that with three or more players, shared randomness exhibits new interesting properties that have no direct analogues in the two-party case. First, we demonstrate a hierarchy of modes of shared randomness, with the usual shared randomness where all parties access the same random string as the strongest form in the hierarchy. We show exponential separations between its levels, and some of our bounds may be of independent interest. For example, we show that the equality function can be solved by a protocol of constant length using the weakest form of shared randomness, which we call XOR-shared randomness. Second, we show that quantum communication cannot replace shared randomness in the k-party case, where k ≥ 3 is any constant. We demonstrate a promise function GPkthat can be computed by a classical protocol of constant length when (the strongest form of) shared randomness is available, but any quantum protocol without shared randomness must send nΩ(1)qubits to compute it. Moreover, the quantum complexity of GPk remains nΩ(1)even if the “second strongest” mode of shared randomness is available. While a somewhat similar separation was already known in the two-party case, in the multi-party case our statement is qualitatively stronger: · In the two-party case, only a relational communication problem with similar properties is known. · In the two-party case, the gap between the two complexities of a problem can be at most exponential, as it is known that 2O(c)log n qubits can always replace shared randomness in any c-bit protocol. Our bounds imply that with quantum communication alone, in general, it is not possible to simulate efficiently even a three-bit three-party classical protocol that uses shared randomness. Dmitry Gavinsky, Tsuyoshi Ito, Guoming Wang |
CCC | 3 |
| 2008 | Parameter Estimation of Quantum ChannelsabstractThe efficiency of parameter estimation of quantum channels is studied in this paper. We introduce the concept of programmable parameters to the theory of estimation. It is found that programmable parameters obey the standard quantum limit strictly; hence, no speedup is possible in its estimation. We also construct a class of nonunitary quantum channels whose parameter can be estimated in a way that the standard quantum limit is broken. The study of estimation of general quantum channels also enables an investigation of the effect of noises on quantum estimation. Zheng-Feng Ji, Guoming Wang, Runyao Duan, Yuan Feng 0001, Mingsheng Ying |
IEEE Trans. Inf. Theory | 2 |