VLDB 2026 Research / reviewers in the wild / expert
Chunyu Xie
dblp:187/1594
· DBLP profile ↗
11ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0002-6607-8209ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Vision and language · 45% Image recognition and object detection · 20% Representation and self-supervised learning · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Software maintenance and evolution · 100% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.7 | 2 | 2025 | LMM-Det: Make Large Multimodal Models Excel in Object Detection · ICCV 2025 IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities · AAAI 2025 |
Computer vision › Image recognition and object detection
object detection |
1.7 | 2 | 2025 | Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection · ICLR 2025 LMM-Det: Make Large Multimodal Models Excel in Object Detection · ICCV 2025 |
Computer vision › Vision and language
vision-language pretraining |
1.5 | 2 | 2025 | FG-CLIP: Fine-Grained Visual and Textual Alignment · ICML 2025 CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023 |
Computer vision › Vision and language › cross-modal alignment
fine-grained alignment |
0.9 | 1 | 2025 | FG-CLIP: Fine-Grained Visual and Textual Alignment · ICML 2025 |
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection |
0.9 | 1 | 2025 | FG-CLIP: Fine-Grained Visual and Textual Alignment · ICML 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities · AAAI 2025 |
Machine learning › Representation and self-supervised learning
structured representation |
0.9 | 1 | 2025 | Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection · ICLR 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
0.8 | 1 | 2024 | Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling · IEEE Trans. Multim. 2024 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.8 | 1 | 2024 | Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling · IEEE Trans. Multim. 2024 |
Computer vision › Vision and language
image-text retrieval |
0.7 | 1 | 2023 | CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023 |
Software maintenance and evolution
log analysis |
0.4 | 1 | 2019 | Robust log-based anomaly detection on unstable log data · ESEC/SIGSOFT FSE 2019 |
Software maintenance and evolution › log analysis
log-based anomaly detection |
0.4 | 1 | 2019 | Robust log-based anomaly detection on unstable log data · ESEC/SIGSOFT FSE 2019 |
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition |
0.3 | 1 | 2018 | Memory Attention Networks for Skeleton-based Action Recognition · IJCAI 2018 |
Medical and health informatics › medical imaging
medical image analysis |
0.3 | 1 | 2025 | Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.2 | 1 | 2024 | Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling · IEEE Trans. Multim. 2024 |
Machine learning › Trustworthy machine learning › interpretability › representation probing
linear probing |
0.2 | 1 | 2024 | Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling · IEEE Trans. Multim. 2024 |
Computer vision › Vision and language
image captioning |
0.2 | 1 | 2023 | CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.2 | 1 | 2023 | CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023 |
Machine learning › Deep learning architectures and training › attention mechanism
attention network |
0.1 | 1 | 2018 | Memory Attention Networks for Skeleton-based Action Recognition · IJCAI 2018 |
Methods — techniques the papers use, named apart from their topics
knowledge bank · 1.7cross-modal feature fusion · 1.7knowledge distillation · 1.4visual grounding · 0.9multimodal adapter · 0.9instruction tuning · 0.9inference optimization · 0.9data distribution adjustment · 0.9contrastive learning · 0.9bounding box annotation · 0.9semantic vector representation · 0.4attention-based Bi-LSTM · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal CapabilitiesabstractIn the field of multimodal large language models (MLLMs), common methods typically involve unfreezing the language model during training to foster profound visual understanding. However, the fine-tuning of such models with vision-language data often leads to a diminution of their natural language processing (NLP) capabilities. To avoid this performance degradation, a straightforward solution is to freeze the language model while developing multimodal competencies. Unfortunately, previous works have not attained satisfactory outcomes. Building on the strategy of freezing the language model, we conduct thorough structural exploration and introduce the Inner-Adaptor Architecture (IAA). Specifically, the architecture incorporates multiple multimodal adaptors at varying depths within the large language model to facilitate direct interaction with the inherently text-oriented transformer layers, thereby enabling the frozen language model to acquire multimodal capabilities. Unlike previous approaches of freezing language models that require large-scale aligned data, our proposed architecture is able to achieve superior performance on small-scale datasets. We conduct extensive experiments to improve the general multimodal capabilities and visual grounding abilities of the MLLM. Our approach remarkably outperforms previous state-of-the-art methods across various vision-language benchmarks without sacrificing performance on NLP tasks. Code and models will be released. Bin Wang 0071, Chunyu Xie, Dawei Leng, Yuhui Yin |
AAAI | 2 |
| 2025 | LMM-Det: Make Large Multimodal Models Excel in Object DetectionabstractLarge multimodal models (LMMs) have garnered wide-spread attention and interest within the artificial intelligence research and industrial communities, owing to their remarkable capability in multimodal understanding, reasoning, and in-context learning, among others. While LMMs have demonstrated promising results in tackling multimodal tasks like image captioning, visual question answering, and visual grounding, the object detection capabilities of LMMs exhibit a significant gap compared to specialist detectors. To bridge the gap, we depart from the conventional methods of integrating heavy detectors with LMMs and propose LMM-Det, a simple yet effective approach that leverages a Large Multimodal Model for vanilla object Detection without relying on specialized detection modules. Specifically, we conduct a comprehensive exploratory analysis when a large multimodal model meets with object detection, revealing that the recall rate degrades significantly compared with specialist detection models. To mitigate this, we propose to increase the recall rate by introducing data distribution adjustment and inference optimization tailored for object detection. We re-organize the instruction conversations to enhance the object detection capabilities of large multimodal models. We claim that a large multimodal model possesses detection capability without any extra detection modules. Extensive experiments support our claim and show the effectiveness of the versatile LMM-Det. The datasets, models, and codes are available at https://github.com/360CVGroup/LMM-Det. Jincheng Li 0002, Chunyu Xie, Ji Ao, Dawei Leng, Yuhui Yin |
ICCV | 2 |
| 2025 | Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detectionabstractZero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Recent studies leverage grounded vision-language models (GLIP) to achieve this by using detailed disease descriptions as prompts for the target disease name during the inference phase.
However, these methods typically treat prompts as equivalent context to the target name, making it difficult to assign specific disease knowledge based on visual information, leading to a coarse alignment between images and target descriptions. In this paper, we propose StructuralGLIP, which introduces an auxiliary branch to encode prompts into a latent knowledge bank layer-by-layer, enabling more context-aware and fine-grained alignment. Specifically, in each layer, we select highly similar features from both the image representation and the knowledge bank, forming structural representations that capture nuanced relationships between image patches and target descriptions. These features are then fused across modalities to further enhance detection performance.
Extensive experiments demonstrate that StructuralGLIP achieves a +4.1\% AP improvement over prior state-of-the-art methods across seven zero-shot medical detection benchmarks, and consistently improves fine-tuned models by +3.2\% AP on endoscopy image datasets. Yuguang Yang 0007, Tongfei Chen, Linlin Yang 0001, Chunyu Xie, Dawei Leng, Xianbin Cao 0001, Baochang Zhang 0001 |
ICLR | 5 |
| 2025 | FG-CLIP: Fine-Grained Visual and Textual AlignmentabstractContrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions. To address this, we propose Fine-Grained CLIP (FG-CLIP), which enhances fine-grained understanding through three key innovations. First, we leverage large multimodal models to generate 1.6 billion long caption-image pairs for capturing global-level semantic details. Second, a high-quality dataset is constructed with 12 million images and 40 million region-specific bounding boxes aligned with detailed captions to ensure precise, context-rich representations. Third, 10 million hard fine-grained negative samples are incorporated to improve the model’s ability to distinguish subtle semantic differences. We construct a comprehensive dataset, termed FineHARD, by integrating high-quality region-specific annotations with challenging fine-grained negative samples. Corresponding training methods are meticulously designed for these data. Extensive experiments demonstrate that FG-CLIP outperforms the original CLIP and other state-of-the-art methods across various downstream tasks, including fine-grained understanding, open-vocabulary object detection, image-text retrieval, and general multimodal benchmarks. These results highlight FG-CLIP’s effectiveness in capturing fine-grained image details and improving overall model performance. The data, code, and models are available at https://github.com/360CVGroup/FG-CLIP. Chunyu Xie, Bin Wang 0071, Fanjing Kong, Jincheng Li 0002, Dawei Liang, Gengshen Zhang, Dawei Leng, Yuhui Yin |
ICML | 1 |
| 2024 | Disjoint Masking With Joint Distillation for Efficient Masked Image ModelingabstractMasked image modeling (MIM) has shown great promise for self-supervised learning (SSL) yet been criticized for learning inefficiency. We believe the insufficient utilization of training signals should be responsible. To alleviate this issue, we introduce a conceptually simple yet learning-efficient MIM training scheme, termedDisjointMasking withJointDistillation (DMJD). For disjoint masking (DM), we sequentially sample multiple masked views per image in a mini-batch with the disjoint regulation to raise the usage of tokens for reconstruction in each image while keeping the masking rate of each view. For joint distillation (JD), we adopt a dual branch architecture to respectively predict invisible (masked) and visible (unmasked) tokens with superior learning targets. Rooting in orthogonal perspectives for training efficiency improvement, DM and JD cooperatively accelerate the training convergence yet not sacrificing the model generalization ability. Concretely, DM can train ViT with less effective training epochs (at most$3.7\times$less time-consuming) to report competitive performance. With JD, our DMJD clearly improves the linear probing classification accuracy, up to 3.4$\%$. On fine-grained downstream tasks like semantic segmentation, object detection,etc., our DMJD also presents superior generalization compared with state-of-the-art SSL methods. Xin Ma 0019, Chang Liu 0047, Chunyu Xie, Long Ye, Yafeng Deng, Xiangyang Ji |
IEEE Trans. Multim. | 3 |
| 2023 | CCMB: A Large-scale Chinese Cross-modal BenchmarkabstractVision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets and downstream datasets with Chinese corpus remain largely unexplored. In this work, we build a large-scale high-quality Chinese Cross-Modal Benchmark named CCMB for the research community, which contains the currently largest public pre-training dataset Zero and five human-annotated fine-tuning datasets for downstream tasks. Zero contains 250 million images paired with 750 million text descriptions, plus two of the five fine-tuning datasets are also currently the largest ones for Chinese cross-modal downstream tasks. Along with the CCMB, we also develop a VLP framework named R2D2, applying a pre-Ranking + Ranking strategy to learn powerful vision-language representations and a two-way distillation method (i.e., target-guided Distillation and feature-guided Distillation) to further enhance the learning capability. With the Zero and the R2D2 VLP framework, we achieve state-of-the-art performance on twelve downstream datasets from five broad categories of tasks including image-text retrieval, image-text matching, image caption, text-to-image generation, and zero-shot image classification. The datasets, models, and codes are available at https://github.com/yuxie11/R2D2 Chunyu Xie, Heng Cai, Jincheng Li 0002, Fanjing Kong, Jianfei Song, Henrique Morimitsu, Lin Yao 0003, Xiangzheng Zhang, Dawei Leng, Baochang Zhang 0001, Xiangyang Ji, Yafeng Deng |
ACM Multimedia | 1 |
| 2022 | Memory Attention Networks for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition has been extensively studied, but it remains an unsolved problem because of the complex variations of skeleton joints in 3-D spatiotemporal space. To handle this issue, we propose a newly temporal-then-spatial recalibration method named memory attention networks (MANs) and deploy MANs using the temporal attention recalibration module (TARM) and spatiotemporal convolution module (STCM). In the TARM, a novel temporal attention mechanism is built based on residual learning to recalibrate frames of skeleton data temporally. In the STCM, the recalibrated sequence is transformed or encoded as the input of CNNs to further model the spatiotemporal information of skeleton sequence. Based on MANs, a new collaborative memory fusion module (CMFM) is proposed to further improve the efficiency, leading to the collaborative MANs (C-MANs), trained with two streams of base MANs. TARM, STCM, and CMFM form a single network seamlessly and enable the whole network to be trained in an end-to-end fashion. Comparing with the state-of-the-art methods, MANs and C-MANs improve the performance significantly and achieve the best results on six data sets for action recognition. The source code has been made publicly available at https://github.com/memory-attention-networks. Ce Li 0002, Chunyu Xie, Baochang Zhang 0001, Jungong Han, Xiantong Zhen, Jie Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Robust log-based anomaly detection on unstable log dataabstractLogs are widely used by large and complex software-intensive systems for troubleshooting. There have been a lot of studies on log-based anomaly detection. To detect the anomalies, the existing methods mainly construct a detection model using log event data extracted from historical logs. However, we find that the existing methods do not work well in practice. These methods have the close-world assumption, which assumes that the log data is stable over time and the set of distinct log events is known. However, our empirical study shows that in practice, log data often contains previously unseen log events or log sequences. The instability of log data comes from two sources: 1) the evolution of logging statements, and 2) the processing noise in log data. In this paper, we propose a new log-based anomaly detection approach, called LogRobust. LogRobust extracts semantic information of log events and represents them as semantic vectors. It then detects anomalies by utilizing an attention-based Bi-LSTM model, which has the ability to capture the contextual information in the log sequences and automatically learn the importance of different log events. In this way, LogRobust is able to identify and handle unstable log events and sequences. We have evaluated LogRobust using logs collected from the Hadoop system and an actual online service system of Microsoft. The experimental results show that the proposed approach can well address the problem of log instability and achieve accurate and robust results on real-world, ever-changing log data. Xu Zhang 0024, Yong Xu 0010, Qingwei Lin, Bo Qiao 0001, Hongyu Zhang 0002, Yingnong Dang, Chunyu Xie, Xinsheng Yang, Ze Li 0005, Junjie Chen 0003, Xiaoting He 0003, Randolph Yao, Jian-Guang Lou, Murali Chintalapati, Furao Shen, Dongmei Zhang 0001 |
ESEC/SIGSOFT FSE | 7 |
| 2019 | Hierarchical residual stochastic networks for time series recognition
Chunyu Xie, Ce Li 0002, Baochang Zhang 0001, Lili Pan 0003, Qixiang Ye, Wei Chen 0016 |
Inf. Sci. | 1 |
| 2018 | Memory Attention Networks for Skeleton-based Action RecognitionabstractSkeleton-based action recognition task is entangled with complex spatio-temporal variations of skeleton joints, and remains challenging for Recurrent Neural Networks (RNNs). In this work, we propose a temporal-then-spatial recalibration scheme to alleviate such complex variations, resulting in an end-to-end Memory Attention Networks (MANs) which consist of a Temporal Attention Recalibration Module (TARM) and a Spatio-Temporal Convolution Module (STCM). Specifically, the TARM is deployed in a residual learning module that employs a novel attention learning network to recalibrate the temporal attention of frames in a skeleton sequence. The STCM treats the attention calibrated skeleton joint sequences as images and leverages the Convolution Neural Networks (CNNs) to further model the spatial and temporal information of skeleton data. These two modules (TARM and STCM) seamlessly form a single network architecture that can be trained in an end-to-end fashion. MANs significantly boost the performance of skeleton-based action recognition and achieve the best results on four challenging benchmark datasets: NTU RGB+D, HDM05, SYSU-3D and UT-Kinect. Chunyu Xie, Ce Li 0002, Baochang Zhang 0001, Chen Chen 0001, Jungong Han, Jianzhuang Liu |
IJCAI | 1 |
| 2018 | Deep Fisher discriminant learning for mobile hand gesture recognition
Ce Li 0002, Chunyu Xie, Baochang Zhang 0001, Chen Chen 0001, Jungong Han |
Pattern Recognit. | 2 |