VLDB 2026 Research / reviewers in the wild / expert
Yimu Wang
dblp:140/7766
· DBLP profile ↗
19ranked-venue papers
12as first author
14since 2021 · last 2026
0000-0003-4255-9104ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 7 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias EstimationabstractExisting vision-language model (VLM)-based methods for out-of-distribution (OOD) detection typically rely on similarity scores between input images and in-distribution (ID) text prototypes. However, the modality gap between image and text often results in high false positive rates, as OOD samples can exhibit high similarity to ID text prototypes. To mitigate the impact of this modality gap, we propose incorporating ID image prototypes along with ID text prototypes. We present theoretical and empirical evidence indicating that this approach enhances VLM-based OOD detection performance without any additional training. To further reduce the gap between image and text, we introduce a novel few-shot tuning framework, suPreMe, comprising biased prompt generation (BPG) and image-text consistency (ITC) modules. BPG enhances image-text fusion and improves generalization (prevents overfitting on the training data) by conditioning ID text prototypes on the Gaussian-based estimated image domain bias; ITC reduces the modality gap by minimizing intra- and inter-modal distances. Moreover, inspired by our theoretical and empirical findings, we introduce a novel OOD score SGMP, leveraging uni- and cross-modal similarities. Finally, extensive experiments demonstrate that suPreMe consistently outperforms existing VLM-based OOD detection methods. Yimu Wang, Evelien Riddell, Adrian Chow, Sean Sedwards, Krzysztof Czarnecki 0001 |
WACV | 1 |
| 2025 | LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal ExpertsabstractRedundancy of visual tokens in multi-modal large language models (MLLMs) significantly reduces their computational efficiency.Recent approaches, such as resamplers and summarizers, have sought to reduce the number of visual tokens, but at the cost of visual reasoning ability.To address this, we propose LEO-MINI, a novel MLLM that significantly reduces the number of visual tokens and simultaneously boosts visual reasoning capabilities.For efficiency, LEO-MINI incorporates COTR, a novel token reduction module to consolidate a large number of visual tokens into a smaller set of tokens, using the similarity between visual tokens, text tokens, and a compact learnable query.For effectiveness, to scale up the model's ability with minimal computational overhead, LEO-MINI employs MMOE, a novel mixture of multi-modal experts module.MMOE employs a set of LoRA experts with a novel router to switch between them based on the input text and visual tokens instead of only using the input hidden state.MMOE also includes a general LoRA expert that is always activated to learn general knowledge for LLM reasoning.For extracting richer visual features, MMOE employs a set of vision experts trained on diverse domain-specific data.To demonstrate LEO-MINI's improved efficiency and performance, we evaluate it against existing efficient MLLMs on various benchmark visionlanguage tasks. Yimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki 0001 |
EMNLP | 1 |
| 2025 | OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection
Adrian Chow, Evelien Riddell, Yimu Wang, Sean Sedwards, Krzysztof Czarnecki 0001 |
ICCV | 3 |
| 2025 | DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation ModelsabstractYimu Wang, Shuai Yuan, Bo Xue, Xiangru Jian, Wei Pang, Mushi Wang, Ning Yu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yimu Wang, Bo Xue 0004, Xiangru Jian, Mushi Wang |
NAACL (Long Papers) | 1 |
| 2025 | Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language ModelsabstractImproving the visual understanding ability of vision-language models (VLMs) is crucial for enhancing their performance across various tasks. While using multiple pretrained visual experts has shown great promise, it often incurs significant computational costs during training and inference. To address this challenge, we propose HAWAII, a novel framework that distills knowledge from multiple visual experts into a single vision encoder, enabling it to inherit the complementary strengths of several experts with minimal computational overhead. To mitigate conflicts among different teachers and switch between different teacher-specific knowledge, instead of using a fixed set of adapters for multiple teachers, we propose to use teacher-specific Low-Rank Adaptation (LoRA) adapters with a corresponding router. Each adapter is aligned with a specific teacher, avoiding noisy guidance during distillation. To enable efficient knowledge distillation, we propose fine-grained and coarse-grained distillation. At the fine-grained level, token importance scores are employed to emphasize the most informative tokens from each teacher adaptively. At the coarse-grained level, we summarize the knowledge from multiple teachers and transfer it to the student using a set of general-knowledge LoRA adapters with a router. Extensive experiments on various vision-language tasks demonstrate the superiority of HAWAII, compared to the popular open-source VLMs. Yimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki 0001 |
NeurIPS | 1 |
| 2025 | AiDe: Improving 3D Open-Vocabulary Semantic Segmentation by Aligned Vision-Language Learningabstract3D open-vocabulary semantic segmentation aims at recognizing countless categories beyond the limited set of annotations used in traditional settings. Due to the lack of large-scale 3D-vision-language segmentation data, instead of training models from scratch, the current solutions distill knowledge from pre-trained 2D vision-language models (VLMs) into 3D models. However, this distillation is supervised by misaligned 3D-scene-image-to-text data pairs, consequently leading to suboptimal performance. Moreover, as 2D VLMs are trained on 2D datasets, text encoders of VLMs, which serve as the bridge between 3D models and an unbounded set of categories, lack 3D semantics. In this paper, to address these issues and improve generalization performance, we propose an Aligned 3D Open-Vocabulary SEmantic Segmentation framework, called AiDe, with two novel modules. To collect high-quality and well-aligned 3D-scene-image-to-text pairs, our CLIP-rewarded alignment module (i) generates diverse captions of multi-view images of 3D scenes to capture details by varying the temperatures and then (ii) samples captions based on their similarity to corresponding images for rich and accurate associations. Next, to adapt 2D VLMs to 3D contexts, our adaptive segmentation module introduces (iii) trainable tokens within the input space and each layer of the text encoder, while freezing the text encoder to avoid catastrophic forgetting. Extensive experiments show that AiDe outperforms previous methods by a large margin on three representative benchmarks, demonstrating its effectiveness. Yimu Wang, Krzysztof Czarnecki 0001 |
WACV | 1 |
| 2025 | Lexicographic Lipschitz Bandits: New Algorithms and a Lower BoundabstractThis paper studies a multiobjective bandit problem under lexicographic ordering, wherein the learner aims to maximize $m$ objectives, each with different levels of importance. First, we introduce the local trade-off, $\lambda_*$, which depicts the trade-off between different objectives. For the case when an upper bound of $\lambda_*$ is known, i.e., $\lambda\geq\lambda_*$, we develop an algorithm that achieves a general regret bound of $\widetilde{O}(\Lambda^i(\lambda)T^{(d_z^i+1)/(d_z^i+2)})$ for the $i$-th objective, where $i\in\{1,2,\ldots,m\}$, $\Lambda^i(\lambda)=1+\lambda+\cdots+\lambda^{i-1}$, $d_z^i$ is the zooming dimension for the $i$-th objective, and $T$ is the time horizon. Next, we provide a matching lower bound for the lexicographic Lipschitz bandit problem, proving that our algorithm is optimal in terms of $\lambda_*$ and $T$. Finally, for the case where $m=2$, we remove the dependence on the knowledge about $\lambda_*$, albeit at the cost of increasing the regret bound to $\widetilde{O}(\Lambda^i(\lambda_*)T^{(3d_z^i+4)/(3d_z^i+6)})$, which remains optimal in terms of $\lambda_*$. Compared to existing work on lexicographic multi-armed bandits, our approach improves the current regret bound of $\widetilde{O}(T^{2/3})$ and extends the number of arms to infinity. Numerical experiments confirm the effectiveness of our algorithms. Bo Xue 0004, Ji Cheng 0001, Fei Liu 0044, Yimu Wang, Lijun Zhang 0005, Qingfu Zhang 0001 |
J. Mach. Learn. Res. | 4 |
| 2024 | Lost Domain Generalization Is a Natural Consequence of Lack of Training DomainsabstractWe show a hardness result for the number of training domains required to achieve a small population error in the test domain. Although many domain generalization algorithms have been developed under various domain-invariance assumptions, there is significant evidence to indicate that out-of-distribution (o.o.d.) test accuracy of state-of-the-art o.o.d. algorithms is on par with empirical risk minimization and random guess on the domain generalization benchmarks such as DomainBed. In this work, we analyze its cause and attribute the lost domain generalization to the lack of training domains. We show that, in a minimax lower bound fashion, any learning algorithm that outputs a classifier with an ε excess error to the Bayes optimal classifier requires at least poly(1/ε) number of training domains, even though the number of training data sampled from each training domain is large. Experiments on the DomainBed benchmark demonstrate that o.o.d. test accuracy is monotonically increasing as the number of training domains increases. Our result sheds light on the intrinsic hardness of domain generalization and suggests benchmarking o.o.d. algorithms by the datasets with a sufficient number of training domains. Yimu Wang, Hongyang Zhang 0001 |
AAAI | 1 |
| 2024 | Multiobjective Lipschitz Bandits under Lexicographic OrderingabstractThis paper studies the multiobjective bandit problem under lexicographic ordering, wherein the learner aims to simultaneously maximize ? objectives hierarchically. The only existing algorithm for this problem considers the multi-armed bandit model, and its regret bound is O((KT)^(2/3)) under a metric called priority-based regret. However, this bound is suboptimal, as the lower bound for single objective multi-armed bandits is Omega(KlogT). Moreover, this bound becomes vacuous when the arm number K is infinite. To address these limitations, we investigate the multiobjective Lipschitz bandit model, which allows for an infinite arm set. Utilizing a newly designed multi-stage decision-making strategy, we develop an improved algorithm that achieves a general regret bound of O(T^((d_z^i+1)/(d_z^i+2))) for the i-th objective, where d_z^i is the zooming dimension for the i-th objective, with i in {1,2,...,m}. This bound matches the lower bound of the single objective Lipschitz bandit problem in terms of T, indicating that our algorithm is almost optimal. Numerical experiments confirm the effectiveness of our algorithm. Bo Xue 0004, Ji Cheng 0001, Fei Liu 0044, Yimu Wang, Qingfu Zhang 0001 |
AAAI | 4 |
| 2023 | Cooperation or Competition: Avoiding Player Domination for Multi-Target Robustness via Adaptive BudgetsabstractDespite incredible advances, deep learning has been shown to be susceptible to adversarial attacks. Numerous approaches have been proposed to train robust networks both empirically and certifiably. However, most of them defend against only a single type of attack, while recent work takes steps forward in defending against multiple attacks. In this paper, to understand multi-target robustness, we view this problem as a bargaining game in which different players (adversaries) negotiate to reach an agreement on a joint direction of parameter updating. We identify a phenomenon named player domination in the bargaining game, namely that the existing max-based approaches, such as MAX and MSD, do not converge. Based on our theoretical analysis, we design a novel framework that adjusts the budgets of different adversaries to avoid any player dominance. Experiments on standard benchmarks show that employing the proposed framework to the existing approaches significantly advances multi-target robustness. Yimu Wang, Dinghuai Zhang, Heng Huang 0001, Hongyang Zhang 0001 |
CVPR | 1 |
| 2023 | Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery BanksabstractIn this work, we present a post-processing solution to address the hubness problem in crossmodal retrieval, a phenomenon where a small number of gallery data points are frequently retrieved, resulting in a decline in retrieval performance.We first theoretically demonstrate the necessity of incorporating both the gallery and query data for addressing hubness as hubs always exhibit high similarity with gallery and query data.Second, building on our theoretical results, we propose a novel framework, Dual Bank Normalization (DBNORM).While previous work has attempted to alleviate hubness by only utilizing the query samples, DB-NORM leverages two banks constructed from the query and gallery samples to reduce the occurrence of hubs during inference.Next, to complement DBNORM, we introduce two novel methods, dual inverted softmax and dual dynamic inverted softmax, for normalizing similarity based on the two banks.Specifically, our proposed methods reduce the similarity between hubs and queries while improving the similarity between non-hubs and queries.Finally, we present extensive experimental results on diverse language-grounded benchmarks, including text-image, text-video, and text-audio, demonstrating the superior performance of our approaches compared to previous methods in addressing hubness and boosting retrieval performance.Our code is available at https://github.com/yimuwangcs/ Better_Cross_Modal_Retrieval. Yimu Wang, Xiangru Jian, Bo Xue 0004 |
EMNLP | 1 |
| 2023 | Multimodal Federated Learning via Contrastive Representation Ensemble
Qiying Yu, Yang Liu 0165, Yimu Wang |
ICLR | 3 |
| 2023 | Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed RewardsabstractThis paper investigates the problem of generalized linear bandits with heavy-tailed rewards, whose $(1+\epsilon)$-th moment is bounded for some $\epsilon\in (0,1]$. Although there exist methods for generalized linear bandits, most of them focus on bounded or sub-Gaussian rewards and are not well-suited for many real-world scenarios, such as financial markets and web-advertising. To address this issue, we propose two novel algorithms based on truncation and mean of medians. These algorithms achieve an almost optimal regret bound of $\widetilde{O}(dT^{\frac{1}{1+\epsilon}})$, where $d$ is the dimension of contextual information and $T$ is the time horizon. Our truncation-based algorithm supports online learning, distinguishing it from existing truncation-based approaches. Additionally, our mean-of-medians-based algorithm requires only $O(\log T)$ rewards and one estimator per epoch, making it more practical. Moreover, our algorithms improve the regret bounds by a logarithmic factor compared to existing algorithms when $\epsilon=1$. Numerical experimental results confirm the merits of our algorithms. Bo Xue 0004, Yimu Wang, Yuanyu Wan, Jinfeng Yi, Lijun Zhang 0005 |
NeurIPS | 2 |
| 2021 | Deep Unified Cross-Modality Hashing by Pairwise Data AlignmentabstractWith the increasing amount of multimedia data, cross-modality hashing has made great progress as it achieves sub-linear search time and low memory space. However, due to the huge discrepancy between different modalities, most existing cross-modality hashing methods cannot learn unified hash codes and functions for modalities at the same time. The gap between separated hash codes and functions further leads to bad search performance. In this paper, to address the issues above, we propose a novel end-to-end Deep Unified Cross-Modality Hashing method named DUCMH, which is able to jointly learn unified hash codes and unified hash functions by alternate learning and data alignment. Specifically, to reduce the discrepancy between image and text modalities, DUCMH utilizes data alignment to learn an auxiliary image to text mapping under the supervision of image-text pairs. For text data, hash codes can be obtained by unified hash functions, while for image data, DUCMH first maps images to texts by the auxiliary mapping, and then uses the mapped texts to obtain hash codes. DUCMH utilizes alternate learning to update unified hash codes and functions. Extensive experiments on three representative image-text datasets demonstrate the superiority of our DUCMH over several state-of-the-art cross-modality hashing methods. Yimu Wang, Bo Xue 0004, Quan Cheng 0001, Yuhui Chen, Lijun Zhang 0005 |
IJCAI | 1 |
| 2020 | Nearly Optimal Regret for Stochastic Linear Bandits with Heavy-Tailed PayoffsabstractIn this paper, we study the problem of stochastic linear bandits with finite action sets. Most of existing work assume the payoffs are bounded or sub-Gaussian, which may be violated in some scenarios such as financial markets. To settle this issue, we analyze the linear bandits with heavy-tailed payoffs, where the payoffs admit finite 1+epsilon moments for some epsilon in (0,1]. Through median of means and dynamic truncation, we propose two novel algorithms which enjoy a sublinear regret bound of widetilde{O}(d^(1/2)T^(1/(1+epsilon))), where d is the dimension of contextual information and T is the time horizon. Meanwhile, we provide an Omega(d^(epsilon/(1+epsilon))T^(1/(1+epsilon))) lower bound, which implies our upper bound matches the lower bound up to polylogarithmic factors in the order of d and T when epsilon=1. Finally, we conduct numerical experiments to demonstrate the effectiveness of our algorithms and the empirical results strongly support our theoretical guarantees. Bo Xue 0004, Guanghui Wang 0006, Yimu Wang, Lijun Zhang 0005 |
IJCAI | 3 |
| 2020 | Searching Privately by Imperceptible Lying: A Novel Private Hashing Method with Differential PrivacyabstractIn the big data era, with the increasing amount of multi-media data, approximate nearest neighbor~(ANN) search has been an important but challenging problem. As a widely applied large-scale ANN search method, hashing has made great progress, and achieved sub-linear search time with low memory space. However, the advances in hashing are based on the availability of large and representative datasets, which often contain sensitive information. Typically, the privacy of this individually sensitive information is compromised. In this paper, we tackle this valuable yet challenging problem and formulate a task termed as private hashing, which takes into account both searching performance and privacy protection. Specifically, we propose a novel noise mechanism, i.e., Random Flipping, and two private hashing algorithms, i.e., PHashing and PITQ, with the refined analysis within the framework of differential privacy, since differential privacy is a well-established technique to measure the privacy leakage of an algorithm. Random Flipping targets binary scenarios and leverages the "Imperceptible Lying" idea to guarantee ε-differential privacy by flipping each datum of the binary matrix (noise addition). To preserve ε-differential privacy, PHashing perturbs and adds noise to the hash codes learned by non-private hashing algorithms using Random Flipping. However, the noise addition for privacy in PHashing will cause severe performance drops. To alleviate this problem, PITQ leverages the power of alternative learning to distribute the noise generated by Random Flipping into each iteration while preserving ε-differential privacy. Furthermore, to empirically evaluate our algorithms, we conduct comprehensive experiments on the image search task and demonstrate that proposed algorithms achieve equal performance compared with non-private hashing methods. Yimu Wang, Shiyin Lu, Lijun Zhang 0005 |
ACM Multimedia | 1 |
| 2020 | Piecewise Hashing: A Deep Hashing Method for Large-Scale Fine-Grained Search
Yimu Wang, Xiu-Shen Wei, Bo Xue 0004, Lijun Zhang 0005 |
PRCV (2) | 1 |
| 2020 | An Adversarial Domain Adaptation Network For Cross-Domain Fine-Grained RecognitionabstractIn this paper, we tackle a valuable yet very challenging visual recognition task, where the instances are within a subordinate category, and the target domain undergoes a shift with the source domain. This task, termed as cross-domain fine-grained recognition, relates closely to many real-life scenarios, e.g., recognizing retail products in storage racks by models trained with images collected in controlled environments. To deal with this problem, we design a new algorithm and propose a corresponding fine-grained domain adaptation dataset. Firstly, we propose a novel end-to-end CNN architecture that integrates two specialized modules: an adversarial module for domain alignment and a self-attention module for fine-grained recognition. The adversarial module is used to handle domain shift by gradually aligning the different domains with domain-level and class-level alignments, and strive to help the classifier learn with domain-invariant features generated by nets. The self-attention module is designed to capture discriminative image regions which are crucial for fine-grained visual recognition. Secondly, we collect a large-scale fine-grained domain adaptation dataset of retail products, which contains 52,011 images of 263 classes from 3 domains. Thirdly, we validate the effectiveness of our method on three datasets, showing that the proposed method can yield significant improvements over baseline methods on fine-grained datasets. Besides, we also evaluate the effectiveness of the self-attention module by performing visualization, which can capture the discriminative image regions in both source and target domains. Yimu Wang, Renjie Song, Xiu-Shen Wei, Lijun Zhang 0005 |
WACV | 1 |
| 2014 | SoC processor for real-time object labeling in life camera streams with low line level latencyabstractImage recognition systems implement a number of processing stages: preprocessing, segmentation and classification. In camera based video processing chains, usually several frame delays are incurred between the moment of capture and the actual availability of the classification results. Hardware architectures for stream based video processing have already been widely employed. In this paper, a new hardware architecture for accelerating the generic task of connected component analysis and object labeling in the segmentation step is presented. The architecture is specifically optimized for very low latency between image component capture by a camera and the detection in hardware. This latency constitutes only a few delay lines, thereby shortening the response time by a few orders of magnitude in comparison to traditional frame-buffer based methods. Zhengqiang Yu, Luc Claesen, Andy Motten, Yimu Wang, Xiaolang Yan |
ISCAS | 5 |