Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chunyu Xie

dblp:187/1594 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0002-6607-8209ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Vision and language · 45% Image recognition and object detection · 20% Representation and self-supervised learning · 13%
Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.722025
LMM-Det: Make Large Multimodal Models Excel in Object Detection · ICCV 2025
IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities · AAAI 2025
Computer vision › Image recognition and object detection
object detection
1.722025
Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection · ICLR 2025
LMM-Det: Make Large Multimodal Models Excel in Object Detection · ICCV 2025
Computer vision › Vision and language
vision-language pretraining
1.522025
FG-CLIP: Fine-Grained Visual and Textual Alignment · ICML 2025
CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023
Computer vision › Vision and language › cross-modal alignment
fine-grained alignment
0.912025
FG-CLIP: Fine-Grained Visual and Textual Alignment · ICML 2025
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
0.912025
FG-CLIP: Fine-Grained Visual and Textual Alignment · ICML 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities · AAAI 2025
Machine learning › Representation and self-supervised learning
structured representation
0.912025
Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection · ICLR 2025
Computer vision › Vision and language
vision-language model
0.912025
Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.812024
Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling · IEEE Trans. Multim. 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.812024
Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling · IEEE Trans. Multim. 2024
Computer vision › Vision and language
image-text retrieval
0.712023
CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023
Software maintenance and evolution
log analysis
0.412019
Robust log-based anomaly detection on unstable log data · ESEC/SIGSOFT FSE 2019
Software maintenance and evolution › log analysis
log-based anomaly detection
0.412019
Robust log-based anomaly detection on unstable log data · ESEC/SIGSOFT FSE 2019
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.312018
Memory Attention Networks for Skeleton-based Action Recognition · IJCAI 2018
Medical and health informatics › medical imaging
medical image analysis
0.312025
Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection · ICLR 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.212024
Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling · IEEE Trans. Multim. 2024
Machine learning › Trustworthy machine learning › interpretability › representation probing
linear probing
0.212024
Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling · IEEE Trans. Multim. 2024
Computer vision › Vision and language
image captioning
0.212023
CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.212023
CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023
Machine learning › Deep learning architectures and training › attention mechanism
attention network
0.112018
Memory Attention Networks for Skeleton-based Action Recognition · IJCAI 2018

Methods — techniques the papers use, named apart from their topics

knowledge bank · 1.7cross-modal feature fusion · 1.7knowledge distillation · 1.4visual grounding · 0.9multimodal adapter · 0.9instruction tuning · 0.9inference optimization · 0.9data distribution adjustment · 0.9contrastive learning · 0.9bounding box annotation · 0.9semantic vector representation · 0.4attention-based Bi-LSTM · 0.4
YearPublicationVenuePosition
2025 IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities
abstract
In the field of multimodal large language models (MLLMs), common methods typically involve unfreezing the language model during training to foster profound visual understanding. However, the fine-tuning of such models with vision-language data often leads to a diminution of their natural language processing (NLP) capabilities. To avoid this performance degradation, a straightforward solution is to freeze the language model while developing multimodal competencies. Unfortunately, previous works have not attained satisfactory outcomes. Building on the strategy of freezing the language model, we conduct thorough structural exploration and introduce the Inner-Adaptor Architecture (IAA). Specifically, the architecture incorporates multiple multimodal adaptors at varying depths within the large language model to facilitate direct interaction with the inherently text-oriented transformer layers, thereby enabling the frozen language model to acquire multimodal capabilities. Unlike previous approaches of freezing language models that require large-scale aligned data, our proposed architecture is able to achieve superior performance on small-scale datasets. We conduct extensive experiments to improve the general multimodal capabilities and visual grounding abilities of the MLLM. Our approach remarkably outperforms previous state-of-the-art methods across various vision-language benchmarks without sacrificing performance on NLP tasks. Code and models will be released.
Bin Wang 0071, Chunyu Xie, Dawei Leng, Yuhui Yin
AAAI2
2025 LMM-Det: Make Large Multimodal Models Excel in Object Detection
abstract
Large multimodal models (LMMs) have garnered wide-spread attention and interest within the artificial intelligence research and industrial communities, owing to their remarkable capability in multimodal understanding, reasoning, and in-context learning, among others. While LMMs have demonstrated promising results in tackling multimodal tasks like image captioning, visual question answering, and visual grounding, the object detection capabilities of LMMs exhibit a significant gap compared to specialist detectors. To bridge the gap, we depart from the conventional methods of integrating heavy detectors with LMMs and propose LMM-Det, a simple yet effective approach that leverages a Large Multimodal Model for vanilla object Detection without relying on specialized detection modules. Specifically, we conduct a comprehensive exploratory analysis when a large multimodal model meets with object detection, revealing that the recall rate degrades significantly compared with specialist detection models. To mitigate this, we propose to increase the recall rate by introducing data distribution adjustment and inference optimization tailored for object detection. We re-organize the instruction conversations to enhance the object detection capabilities of large multimodal models. We claim that a large multimodal model possesses detection capability without any extra detection modules. Extensive experiments support our claim and show the effectiveness of the versatile LMM-Det. The datasets, models, and codes are available at https://github.com/360CVGroup/LMM-Det.
Jincheng Li 0002, Chunyu Xie, Ji Ao, Dawei Leng, Yuhui Yin
ICCV2
2025 Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection
abstract
Zero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Recent studies leverage grounded vision-language models (GLIP) to achieve this by using detailed disease descriptions as prompts for the target disease name during the inference phase. However, these methods typically treat prompts as equivalent context to the target name, making it difficult to assign specific disease knowledge based on visual information, leading to a coarse alignment between images and target descriptions. In this paper, we propose StructuralGLIP, which introduces an auxiliary branch to encode prompts into a latent knowledge bank layer-by-layer, enabling more context-aware and fine-grained alignment. Specifically, in each layer, we select highly similar features from both the image representation and the knowledge bank, forming structural representations that capture nuanced relationships between image patches and target descriptions. These features are then fused across modalities to further enhance detection performance. Extensive experiments demonstrate that StructuralGLIP achieves a +4.1\% AP improvement over prior state-of-the-art methods across seven zero-shot medical detection benchmarks, and consistently improves fine-tuned models by +3.2\% AP on endoscopy image datasets.
Yuguang Yang 0007, Tongfei Chen, Linlin Yang 0001, Chunyu Xie, Dawei Leng, Xianbin Cao 0001, Baochang Zhang 0001
ICLR5
2025 FG-CLIP: Fine-Grained Visual and Textual Alignment
abstract
Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions. To address this, we propose Fine-Grained CLIP (FG-CLIP), which enhances fine-grained understanding through three key innovations. First, we leverage large multimodal models to generate 1.6 billion long caption-image pairs for capturing global-level semantic details. Second, a high-quality dataset is constructed with 12 million images and 40 million region-specific bounding boxes aligned with detailed captions to ensure precise, context-rich representations. Third, 10 million hard fine-grained negative samples are incorporated to improve the model’s ability to distinguish subtle semantic differences. We construct a comprehensive dataset, termed FineHARD, by integrating high-quality region-specific annotations with challenging fine-grained negative samples. Corresponding training methods are meticulously designed for these data. Extensive experiments demonstrate that FG-CLIP outperforms the original CLIP and other state-of-the-art methods across various downstream tasks, including fine-grained understanding, open-vocabulary object detection, image-text retrieval, and general multimodal benchmarks. These results highlight FG-CLIP’s effectiveness in capturing fine-grained image details and improving overall model performance. The data, code, and models are available at https://github.com/360CVGroup/FG-CLIP.
Chunyu Xie, Bin Wang 0071, Fanjing Kong, Jincheng Li 0002, Dawei Liang, Gengshen Zhang, Dawei Leng, Yuhui Yin
ICML1
2024 Disjoint Masking With Joint Distillation for Efficient Masked Image Modeling
abstract
Masked image modeling (MIM) has shown great promise for self-supervised learning (SSL) yet been criticized for learning inefficiency. We believe the insufficient utilization of training signals should be responsible. To alleviate this issue, we introduce a conceptually simple yet learning-efficient MIM training scheme, termedDisjointMasking withJointDistillation (DMJD). For disjoint masking (DM), we sequentially sample multiple masked views per image in a mini-batch with the disjoint regulation to raise the usage of tokens for reconstruction in each image while keeping the masking rate of each view. For joint distillation (JD), we adopt a dual branch architecture to respectively predict invisible (masked) and visible (unmasked) tokens with superior learning targets. Rooting in orthogonal perspectives for training efficiency improvement, DM and JD cooperatively accelerate the training convergence yet not sacrificing the model generalization ability. Concretely, DM can train ViT with less effective training epochs (at most$3.7\times$less time-consuming) to report competitive performance. With JD, our DMJD clearly improves the linear probing classification accuracy, up to 3.4$\%$. On fine-grained downstream tasks like semantic segmentation, object detection,etc., our DMJD also presents superior generalization compared with state-of-the-art SSL methods.
Xin Ma 0019, Chang Liu 0047, Chunyu Xie, Long Ye, Yafeng Deng, Xiangyang Ji
IEEE Trans. Multim.3
2023 CCMB: A Large-scale Chinese Cross-modal Benchmark
abstract
Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets and downstream datasets with Chinese corpus remain largely unexplored. In this work, we build a large-scale high-quality Chinese Cross-Modal Benchmark named CCMB for the research community, which contains the currently largest public pre-training dataset Zero and five human-annotated fine-tuning datasets for downstream tasks. Zero contains 250 million images paired with 750 million text descriptions, plus two of the five fine-tuning datasets are also currently the largest ones for Chinese cross-modal downstream tasks. Along with the CCMB, we also develop a VLP framework named R2D2, applying a pre-Ranking + Ranking strategy to learn powerful vision-language representations and a two-way distillation method (i.e., target-guided Distillation and feature-guided Distillation) to further enhance the learning capability. With the Zero and the R2D2 VLP framework, we achieve state-of-the-art performance on twelve downstream datasets from five broad categories of tasks including image-text retrieval, image-text matching, image caption, text-to-image generation, and zero-shot image classification. The datasets, models, and codes are available at https://github.com/yuxie11/R2D2
Chunyu Xie, Heng Cai, Jincheng Li 0002, Fanjing Kong, Jianfei Song, Henrique Morimitsu, Lin Yao 0003, Xiangzheng Zhang, Dawei Leng, Baochang Zhang 0001, Xiangyang Ji, Yafeng Deng
ACM Multimedia1
2022 Memory Attention Networks for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has been extensively studied, but it remains an unsolved problem because of the complex variations of skeleton joints in 3-D spatiotemporal space. To handle this issue, we propose a newly temporal-then-spatial recalibration method named memory attention networks (MANs) and deploy MANs using the temporal attention recalibration module (TARM) and spatiotemporal convolution module (STCM). In the TARM, a novel temporal attention mechanism is built based on residual learning to recalibrate frames of skeleton data temporally. In the STCM, the recalibrated sequence is transformed or encoded as the input of CNNs to further model the spatiotemporal information of skeleton sequence. Based on MANs, a new collaborative memory fusion module (CMFM) is proposed to further improve the efficiency, leading to the collaborative MANs (C-MANs), trained with two streams of base MANs. TARM, STCM, and CMFM form a single network seamlessly and enable the whole network to be trained in an end-to-end fashion. Comparing with the state-of-the-art methods, MANs and C-MANs improve the performance significantly and achieve the best results on six data sets for action recognition. The source code has been made publicly available at https://github.com/memory-attention-networks.
Ce Li 0002, Chunyu Xie, Baochang Zhang 0001, Jungong Han, Xiantong Zhen, Jie Chen 0001
IEEE Trans. Neural Networks Learn. Syst.2
2019 Robust log-based anomaly detection on unstable log data
abstract
Logs are widely used by large and complex software-intensive systems for troubleshooting. There have been a lot of studies on log-based anomaly detection. To detect the anomalies, the existing methods mainly construct a detection model using log event data extracted from historical logs. However, we find that the existing methods do not work well in practice. These methods have the close-world assumption, which assumes that the log data is stable over time and the set of distinct log events is known. However, our empirical study shows that in practice, log data often contains previously unseen log events or log sequences. The instability of log data comes from two sources: 1) the evolution of logging statements, and 2) the processing noise in log data. In this paper, we propose a new log-based anomaly detection approach, called LogRobust. LogRobust extracts semantic information of log events and represents them as semantic vectors. It then detects anomalies by utilizing an attention-based Bi-LSTM model, which has the ability to capture the contextual information in the log sequences and automatically learn the importance of different log events. In this way, LogRobust is able to identify and handle unstable log events and sequences. We have evaluated LogRobust using logs collected from the Hadoop system and an actual online service system of Microsoft. The experimental results show that the proposed approach can well address the problem of log instability and achieve accurate and robust results on real-world, ever-changing log data.
Xu Zhang 0024, Yong Xu 0010, Qingwei Lin, Bo Qiao 0001, Hongyu Zhang 0002, Yingnong Dang, Chunyu Xie, Xinsheng Yang, Ze Li 0005, Junjie Chen 0003, Xiaoting He 0003, Randolph Yao, Jian-Guang Lou, Murali Chintalapati, Furao Shen, Dongmei Zhang 0001
ESEC/SIGSOFT FSE7
2019 Hierarchical residual stochastic networks for time series recognition
Chunyu Xie, Ce Li 0002, Baochang Zhang 0001, Lili Pan 0003, Qixiang Ye, Wei Chen 0016
Inf. Sci.1
2018 Memory Attention Networks for Skeleton-based Action Recognition
abstract
Skeleton-based action recognition task is entangled with complex spatio-temporal variations of skeleton joints, and remains challenging for Recurrent Neural Networks (RNNs). In this work, we propose a temporal-then-spatial recalibration scheme to alleviate such complex variations, resulting in an end-to-end Memory Attention Networks (MANs) which consist of a Temporal Attention Recalibration Module (TARM) and a Spatio-Temporal Convolution Module (STCM). Specifically, the TARM is deployed in a residual learning module that employs a novel attention learning network to recalibrate the temporal attention of frames in a skeleton sequence. The STCM treats the attention calibrated skeleton joint sequences as images and leverages the Convolution Neural Networks (CNNs) to further model the spatial and temporal information of skeleton data. These two modules (TARM and STCM) seamlessly form a single network architecture that can be trained in an end-to-end fashion. MANs significantly boost the performance of skeleton-based action recognition and achieve the best results on four challenging benchmark datasets: NTU RGB+D, HDM05, SYSU-3D and UT-Kinect.
Chunyu Xie, Ce Li 0002, Baochang Zhang 0001, Chen Chen 0001, Jungong Han, Jianzhuang Liu
IJCAI1
2018 Deep Fisher discriminant learning for mobile hand gesture recognition
Ce Li 0002, Chunyu Xie, Baochang Zhang 0001, Chen Chen 0001, Jungong Han
Pattern Recognit.2