EDBT 2026 Demo / reviewers in the wild / expert
Yizhou Wang 0006
dblp:71/3387-6
· DBLP profile ↗
18ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-1601-9649ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revealing the Seen, Imagining the Beyond: A Survey of Image-Grounded Chain-of-Thought Reasoning in Multimodal LLMsabstractQihua Dong, Yitian Zhang, Huimin Zeng, Yizhou Wang, Jianglin Lu, Kuo Yang, Yun Fu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qihua Dong, Yizhou Wang 0006, Jianglin Lu, Yun Fu 0001 |
ACL (1) | 4 |
| 2026 | From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual SegmentationabstractVisual segmentation, the task of segmenting an image into semantically meaningful regions, is a cornerstone in machine learning and has widespread applications in industry. Nevertheless, visual segmentation with instruction has been a challenging task for many years. This largely stems from the cross-modal discrepancy between language and image domains, resulting in difficulty in relating the instruction semantics and the pixel-level predictions. In recent years, the remarkable reasoning capabilities of Large Language Models (LLMs) and Large Multimodal Models (LMMs) have spurred a new wave of research aiming to bridge the disparity between natural language instructions and pixel-level understanding. This survey offers the first comprehensive overview of the rapidly evolving field of LLM-driven visual segmentation. We categorize existing approaches based on their core objectives and methodologies, including reasoning-based segmentation, open-vocabulary segmentation, grounding techniques connecting language to pixels, and extensions to video domains. We review recent seminal works in LLM-based visual segmentation, analyzing their architectural innovations, training strategies, and benchmark performance. Furthermore, we discuss the common datasets, evaluation metrics, and identify key challenges and promising future directions at the intersection of language and visual segmentation. We hope this survey serves as a valuable resource for researchers and practitioners seeking to understand the current landscape and future directions of leveraging LLMs for sophisticated visual segmentation tasks and applications. The resource summary is available at https://github.com/wyzjack/Awesome-LLM-Visual-Segmentation. Yizhou Wang 0006, Mang Tik Chiu, Lingzhi Zhang, Xuan Shen, Sohrab Amirghodsi, Yun Fu 0001 |
ACL (1) | 1 |
| 2025 | D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question DecompositionabstractVideo large language models (Vid-LLMs), which excel in diverse video-language tasks, can be effectively constructed by adapting image-pretrained vision-language models (VLMs).However, this adaptation remains challenging, as it requires processing dense and temporally extended visual inputs that exceed the capacity of image-based models.This paper identifies the perception bottleneck and token overload as key challenges in extending image-based VLMs to the video domain.To address these issues, we propose D-CoDe, a training-free adaptation framework that incorporates dynamic compression and question decomposition.Specifically, dynamic compression alleviates the perception bottleneck through adaptive selection of representative frames and content-aware aggregation of spatial tokens, thereby reducing redundancy while preserving informative content.In parallel, question decomposition mitigates token overload by reformulating the original query into sub-questions, guiding the model to focus on distinct aspects of the video and enabling more comprehensive understanding.Experiments demonstrate that D-CoDe effectively improves video understanding across various benchmarks.Furthermore, strong performance on the challenging long-video benchmark highlights the potential of D-CoDe in handling complex video-language tasks. Yiyang Huang 0001, Yizhou Wang 0006, Yun Fu 0001 |
EMNLP | 2 |
| 2025 | Representation Potentials of Foundation Models for Multimodal Alignment: A SurveyabstractFoundation models learn highly transferable representations through large-scale pretraining on diverse data.An increasing body of research indicates that these representations exhibit a remarkable degree of similarity across architectures and modalities.In this survey, we investigate the representation potentials of foundation models, defined as the latent capacity of their learned representations to capture task-specific information within a single modality while also providing a transferable basis for alignment and unification across modalities.We begin by reviewing representative foundation models and the key metrics that make alignment measurable.We then synthesize empirical evidence of representation potentials from studies in vision, language, speech, multimodality, and neuroscience.The evidence suggests that foundation models often exhibit structural regularities and semantic consistencies in their representation spaces, positioning them as strong candidates for cross-modal transfer and alignment.We further analyze the key factors that foster representation potentials, discuss open questions, and highlight potential challenges. Jianglin Lu, Yi Xu 0005, Yizhou Wang 0006, Yun Fu 0001 |
EMNLP | 4 |
| 2025 | Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language ModelsabstractRecent work has revealed that large language models (LLMs) can exhibit emergent theory-of-mind (ToM) capabilities—inferring human beliefs, desires, and intentions from text alone. Yet, everyday social reasoning often unfolds visually in dynamic contexts. This paper investigates whether multimodal LLMs can similarly demonstrate ToM skills in video-based tasks. Concretely, we propose a pipeline that fuses video and text signals, retrieves the most relevant frames for each query, and answers questions requiring spatio-temporal social understanding. We introduce a new frame localization benchmark, Theory of Mind Localization (ToMLoc), and show that finetuning a state-of-the-art Video-ChatGPT model on ToMLoc significantly improves performance on Social-Iq 2.0. Our results suggest that bridging textual and visual modalities is essential for capturing complex mental states in real-world scenarios. Moreover, retrieving key frames enhances interpretability by revealing how the model arrives at its inferences. These findings highlight the promise of video-based approaches for achieving more human-like social intelligence in LLMs. Zhanwen Chen, Tianchun Wang, Yizhou Wang 0006, Michal Kosinski, Xiang Zhang 0001, Yun Fu 0001, Sheng Li 0001 |
IJCNN | 3 |
| 2025 | Towards Zero-shot 3D Anomaly Localizationabstract3D anomaly detection and localization is of great significancefor industrial inspection. Prior 3D anomaly detection and localization methods focus on the setting that the testing data share the same category as the training data which is normal. However, in real-world applications, the normal training data for the target 3D objects can be unavailable due to issues like data privacy or export control regulation. To tackle these challenges, we identify a new task -zero-shot 3D anomaly detection and localization, where the training and testing classes do not overlap. To this end, we design 3DzAL,a novel patch-level contrastive learning framework based on pseudo anomalies generated using the inductive bias from task-irrelevant 3D xyz data to learn more representative feature representations. Fur-thermore, we train a normalcy classifier network to classify the normal patches and pseudo anomalies and utilize the classification result jointly with feature distance to design anomaly scores. Instead of directly using the patch point clouds, we introduce adversarial perturbations to the input patch xyz data before feeding into the 3D normalcy classifier for the classification-based anomaly score. We show that 3DzAL outperforms the state-of-the-art anomaly detection and localization performance. Yizhou Wang 0006, Kuan-Chuan Peng, Yun Fu 0001 |
WACV | 1 |
| 2024 | Rewrite the StarsabstractRecent studies have drawn attention to the untapped potential of the “star operation” (element-wise multiplication) in network design. While intuitive explanations abound, the foundational rationale behind its application remains largely unexplored. Our study attempts to reveal the star operation's ability of mapping inputs into high-dimensional, non-linear feature spaces-akin to kernel tricks-without widening the network. We further introduce StarNet, a simple yet pow-erful prototype, demonstrating impressive performance and low latency under compact network structure and efficient budget. Like stars in the sky, the star operation appears unremarkable but holds a vast universe of potential. Our work encourages further exploration across tasks, with codes available at https://github.com/ma-xu/Rewrite-the-Stars. Xu Ma 0005, Xiyang Dai, Yizhou Wang 0006, Yun Fu 0001 |
CVPR | 4 |
| 2024 | Don't Judge by the Look: Towards Motion Coherent Video RepresentationabstractCurrent training pipelines in object recognition neglect Hue Jittering when doing data augmentation as it not only brings appearance changes that are detrimental to classification, but also the implementation is inefficient in practice. In this study, we investigate the effect of hue variance in the context of video understanding and find this variance to be beneficial since static appearances are less important in videos that contain motion information. Based on this observation, we propose a data augmentation method for video understanding, named Motion Coherent Augmentation (MCA), that introduces appearance variation in videos and implicitly encourages the model to prioritize motion patterns, rather than static appearances. Concretely, we propose an operation SwapMix to efficiently modify the appearance of video samples, and introduce Variation Alignment (VA) to resolve the distribution shift caused by SwapMix, enforcing the model to learn appearance invariant representations. Comprehensive empirical evaluation across various architectures and different datasets solidly validates the effectiveness and generalization ability of MCA, and the application of VA in other augmentation methods. Code is available at https://github.com/BeSpontaneous/MCA-pytorch. Huan Wang 0014, Yizhou Wang 0006, Yun Fu 0001 |
ICLR | 4 |
| 2024 | SLA$^{{\text{2}}}$2P: Self-Supervised Anomaly Detection With Adversarial PerturbationabstractAnomaly detection is a foundational yet difficult problem in machine learning. In this work, we propose a new and effective framework, dubbed as SLA2P, for unsupervised anomaly detection. Following the extraction of delegate embeddings from raw data, we implement random projections on the features and consider features transformed by disparate projections as being associated with separate pseudo-classes. We then train a neural network for classification on these transformed features to conduct self-supervised learning. Subsequently, we introduce adversarial disturbances to the modified attributes, and we develop anomaly scores built on the classifier's predictive uncertainties concerning these disrupted features. Our approach is motivated by the fact that as anomalies are relatively rare and decentralized, 1) the training of the pseudo-label classifier concentrates more on acquiring the semantic knowledge of regular data instead of anomalous data; 2) the altered attributes of the normal data exhibit greater resilience to disturbances compared to those of the anomalous data. Therefore, the disrupted modified attributes of anomalies can not be well classified and correspondingly tend to attain lesser anomaly scores. The results of experiments on various benchmark datasets for images, text, and inherently tabular data demonstrate that SLA2P achieves state-of-the-art performance consistently. Yizhou Wang 0006, Can Qin, Rongzhe Wei, Yi Xu 0005, Yun Fu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Momentum is All You Need for Data-Driven Adaptive OptimizationabstractAdaptive gradient methods, e.g., ADAM, have achieved tremendous success in data-driven machine learning, especially deep learning. Employing adaptive learning rates according to the gradients, such methods are able to attain rapid training of modern deep neural networks. Nevertheless, they are observed to suffer from compromised generalization capacity compared with stochastic gradient descent (SGD) and tend to be trapped in local minima at an early stage during the training process. Intriguingly, we discover that the issue can be resolved by substituting the gradient in the second raw moment estimate term with its exponential moving average version in ADAM. The intuition is that the gradient with momentum contains more accurate directional information, and therefore its second-moment estimation is a more preferable option for learning rate scaling than that of the raw gradient. Thereby we propose ADAM$^{3}$ as a new optimizer reaching the goal of training quickly while generalizing much better. Extensive experiments on a variety of tasks and models demonstrate that ADAM$^{3}$ exhibits state-of-the-art performance and superior training stability consistently. Considering the simplicity and effectiveness of ADAM$^{3}$, we believe it has the potential to become a new standard method in deep learning. Code is provided at https://github.com/wyzjack/AdaM3. Yizhou Wang 0006, Yue Kang 0002, Can Qin, Huan Wang 0014, Yi Xu 0005, Yulun Zhang 0001, Yun Fu 0001 |
ICDM | 1 |
| 2023 | Concentric Ring Loss for Face Forgery DetectionabstractThe issue of detecting face forgeries has garnered significant interest in the field of computer vision, primarily driven by the growing social concerns of indistinguishable deepfake images. One of the primary obstacles encountered in the field of deepfake detection is enhancing the discriminative power of learned features. In this paper, we propose a Concentric Ring Loss (CRL) that aims to promote the learning of compressed intra-class features and separated inter-class features inside a model. Specifically, we apply margin penalties in both Euclidean and angular space separately, which serve to increase the separation between real and fake images. Moreover, we introduce a frequency-aware triplet network with a self-developed sample generation strategy, which provides efficient hard triplets for model training. Extensive experiments demonstrate the superiority of our methods over multiple datasets. We show that CRL consistently outperforms the state-of-the-art by a large margin. Yu Yin 0001, Yizhou Wang 0006, Yun Fu 0001 |
ICDM | 3 |
| 2023 | An Unrolled Implicit Regularization Network for Joint Image and Sensitivity Estimation in Parallel MR Imaging with Convergence GuaranteeabstractAbstract. Parallel imaging (PI), relying on multicoils to sense [Formula: see text]-space data, is an effective technique to accelerate magnetic resonance imaging by exploiting spatial sensitivity coding of multiple coils, with an integrated compressive sensing (CS) technology to achieve higher acceleration. In this paper, we propose a novel nonconvex reconstruction model and its proximal alternating linearized minimization (PALM) algorithm for PI in a blind setting that MR image and multichannel sensitivity maps are jointly estimated, regularized by image and sensitivity regularizers. Instead of hand-crafting the image and sensitivity regularizers, we propose unrolling the PALM algorithm to be a deep network for Blind Parallel MRI, dubbed as BPMRI-Net, with two learnable subnetworks to substitute the proximal operators of the image and sensitivity regularizers. We theoretically prove the linear convergence of BPMRI-Net as an iterative algorithm, which alternately updates two variables based on the learnable proximal operators. The learned BPMRI-Net can simultaneously output the MR image and sensitivity maps from undersampled multichannel [Formula: see text]-space data even when the number of low-frequency sampling lines in the center of [Formula: see text]-space is small. Numerical results demonstrate the effectiveness of our method with state-of-the-art reconstruction accuracy. Yan Yang 0007, Yizhou Wang 0006, Jiazhen Wang, Jian Sun 0009, Zongben Xu |
SIAM J. Imaging Sci. | 2 |
| 2022 | Self-supervision Meets Adversarial Perturbation: A Novel Framework for Anomaly DetectionabstractAnomaly detection is a fundamental yet challenging problem in machine learning due to the lack of label information. In this work, we propose a novel and powerful framework, dubbed as SLA2P, for unsupervised anomaly detection. After extracting representative embeddings from raw data, we apply random projections to the features and regard features transformed by different projections as belonging to distinct pseudo-classes. We then train a classifier network on these transformed features to perform self-supervised learning. Next, we add adversarial perturbation to the transformed features to decrease their softmax scores of the predicted labels and design anomaly scores based on the predictive uncertainties of the classifier on these perturbed features. Our motivation is that because of the relatively small number and the decentralized modes of anomalies, 1) the pseudo label classifier's training concentrates more on learning the semantic information of normal data rather than anomalous data; 2) the transformed features of the normal data are more robust to the perturbations than those of the anomalies. Consequently, the perturbed transformed features of anomalies fail to be classified well and accordingly have lower anomaly scores than those of the normal samples. Extensive experiments on image, text, and inherently tabular benchmark datasets back up our findings and indicate that SLA2 achieves state-of-the-art anomaly detection performance consistently. Our code is made publicly available at https://github.com/wyzjack/SLA2P Yizhou Wang 0006, Can Qin, Rongzhe Wei, Yi Xu 0005, Yun Fu 0001 |
CIKM | 1 |
| 2022 | Robust Semi-supervised Domain Adaptation against Noisy LabelsabstractBuilt upon clean/correct labels, semi-supervised domain adaptation (SSDA) is a well-explored task, which, however, may not be easily obtained. This paper considers a challenging but practical scenario, i.e., the noisy SSDA with polluted labels. Specifically, it is observed that abnormal samples appear to have more randomness and inconsistency among the various views. To this end, we have devised an anomaly score function to detect noisy samples based on the similarity of differently augmented instances. The noisy labeled target samples are re-weighted according to such anomaly scores where the abnormal data contribute less to model training. Moreover, pseudo labeling usually suffers from confirmation bias. To remedy it, we have introduced the adversarial disturbance to raise the divergence across differently augmented views. The experimental results on the contaminated SSDA benchmarks demonstrate the effectiveness of our method over the baselines in both robustness and accuracy. Can Qin, Yizhou Wang 0006, Yun Fu 0001 |
CIKM | 2 |
| 2022 | Adaptive Trajectory Prediction via Transferable GNNabstractPedestrian trajectory prediction is an essential component in a wide range of AI applications such as autonomous driving and robotics. Existing methods usually assume the training and testing motions follow the same pattern while ignoring the potential distribution differences (e.g., shopping mall and street). This issue results in inevitable performance decrease. To address this issue, we propose a novel Transferable Graph Neural Network (TGNN) frame-work, which jointly conducts trajectory prediction as well as domain alignment in a unified framework. Specifically, a domain-invariant GNN is proposed to explore the structural motion knowledge where the domain-specific knowledge is reduced. Moreover, an attention-based adaptive knowledge learning module is further proposed to explore fine-grained individual-level feature representations for knowledge transfer. By this way, disparities across different trajectory domains will be better alleviated. More challenging while practical trajectory prediction experiments are designed, and the experimental results verify the superior performance of our proposed model. To the best of our knowledge, our work is the pioneer which fills the gap in benchmarks and techniques for practical pedestrian trajectory prediction across different domains. Yi Xu 0005, Lichen Wang, Yizhou Wang 0006, Yun Fu 0001 |
CVPR | 3 |
| 2022 | Making Reconstruction-based Method Great Again for Video Anomaly DetectionabstractAnomaly detection in videos is a significant yet challenging problem. Previous approaches based on deep neural networks employ either reconstruction-based or prediction-based approaches. Nevertheless, existing reconstruction-based methods 1) rely on old-fashioned convolutional autoencoders and are poor at modeling temporal dependency; 2) are prone to overfit the training samples, leading to indistinguishable reconstruction errors of normal and abnormal frames during the inference phase. To address such issues, firstly, we get inspiration from transformer and propose Spatio-Temporal Auto-Trans-Encoder, dubbed as STATE, as a new autoencoder model for enhanced consecutive frame reconstruction. Our STATE is equipped with a specifically designed learnable convolutional attention module for efficient temporal learning and reasoning. Secondly, we put forward a novel reconstruction-based input perturbation technique during testing to further differentiate anomalous frames. With the same perturbation magnitude, the testing reconstruction error of the normal frames lowers more than that of the abnormal frames, which contributes to mitigating the overfitting problem of reconstruction. Owing to the high relevance of the frame abnormality and the objects in the frame, we conduct object-level reconstruction using both the raw frame and the corresponding optical flow patches. Finally, the anomaly score is designed based on the combination of the raw and motion reconstruction errors using perturbed inputs. Extensive experiments on benchmark video anomaly detection datasets demonstrate that our approach outperforms previous reconstruction-based methods by a notable margin, and achieves state-of-the-art anomaly detection performance consistently. The code is available at https://github.com/wyzjack/MRMGA4VAD. Yizhou Wang 0006, Can Qin, Yi Xu 0005, Xu Ma 0005, Yun Fu 0001 |
ICDM | 1 |
| 2022 | MemREIN: Rein the Domain Shift for Cross-Domain Few-Shot LearningabstractFew-shot learning aims to enable models generalize to new categories (query instances) with only limited labeled samples (support instances) from each category. Metric-based mechanism is a promising direction which compares feature embeddings via different metrics. However, it always fail to generalize to unseen domains due to the considerable domain gap challenge. In this paper, we propose a novel framework, MemREIN, which considers Memorized, Restitution, and Instance Normalization for cross-domain few-shot learning. Specifically, an instance normalization algorithm is explored to alleviate feature dissimilarity, which provides the initial model generalization ability. However, naively normalizing the feature would lose fine-grained discriminative knowledge between different classes. To this end, a memorized module is further proposed to separate the most refined knowledge and remember it. Then, a restitution module is utilized to restitute the discrimination ability from the learned knowledge. A novel reverse contrastive learning strategy is proposed to stabilize the distillation process. Extensive experiments on five popular benchmark datasets demonstrate that MemREIN well addresses the domain shift challenge, and significantly improves the performance up to 16.43% compared with state-of-the-art baselines. Yi Xu 0005, Lichen Wang, Yizhou Wang 0006, Can Qin, Yulun Zhang 0001, Yun Fu 0001 |
IJCAI | 3 |
| 2020 | On Computation and Generalization of Generative Adversarial Imitation Learning
Minshuo Chen, Yizhou Wang 0006, Zhuoran Yang, Xingguo Li, Zhaoran Wang 0001, Tuo Zhao |
ICLR | 2 |