EDBT 2026 Demo / reviewers in the wild / expert
Yin Xie
dblp:77/2706
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 47% Vision and language · 27% Video understanding and tracking · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 50% Computational photography and imaging · 50% |
Topics — the 14 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
multimodal representation learning |
1.2 | 2 | 2026 | RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm · ACM Multimedia 2025 ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs · AAAI 2026 |
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining |
0.9 | 1 | 2025 | RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm · ACM Multimedia 2025 |
Computer vision › Video understanding and tracking
event clustering |
0.9 | 1 | 2025 | UniViT: Unifying Image and Video Understanding in One Vision Encoder · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
region representation learning |
0.9 | 1 | 2025 | Region-based Cluster Discrimination for Visual Representation Learning · ICCV 2025 |
Robotics › Robot navigation and mapping › spatial cognition › spatial knowledge
spatial semantics |
0.9 | 1 | 2025 | UniViT: Unifying Image and Video Understanding in One Vision Encoder · NeurIPS 2025 |
Computer vision › Video understanding and tracking › temporal modeling
temporal dynamics modeling |
0.9 | 1 | 2025 | UniViT: Unifying Image and Video Understanding in One Vision Encoder · NeurIPS 2025 |
Computer vision › Vision and language
vision-language pretraining |
0.9 | 1 | 2025 | Region-based Cluster Discrimination for Visual Representation Learning · ICCV 2025 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.9 | 1 | 2025 | Region-based Cluster Discrimination for Visual Representation Learning · ICCV 2025 |
Information retrieval
cross-modal retrieval |
0.9 | 1 | 2025 | RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm · ACM Multimedia 2025 |
Information retrieval › cross-modal retrieval
image-text retrieval |
0.9 | 1 | 2025 | RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm · ACM Multimedia 2025 |
Image and video processing › image restoration › image deblurring
defocus deblurring |
0.9 | 1 | 2025 | Quad-Pixel Image Defocus Deblurring: A New Benchmark and Model · CVPR 2025 |
Collaborative and social computing
remote collaboration |
0.3 | 1 | 2017 | Codeon: On-Demand Software Development Assistance · CHI 2017 |
Requirements engineering and software design
developer support tools |
0.3 | 1 | 2017 | Codeon: On-Demand Software Development Assistance · CHI 2017 |
Interaction techniques and input › voice interaction
speech input |
0.1 | 1 | 2017 | Codeon: On-Demand Software Development Assistance · CHI 2017 |
Methods — techniques the papers use, named apart from their topics
image semantic augmented generation · 1.7hierarchical retrieval · 1.7token reconstruction loss · 1.0hungarian matching · 1.0semantic balance sampling · 0.9rotary position embedding · 0.9region transformer · 0.9mamba · 0.9local-gate attention · 0.9distributed training · 0.9contrastive learning · 0.9clustering · 0.9cluster discrimination · 0.9speech recognition · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMsabstractLarge Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we call ViCToR (Visual Comprehension via Token Reconstruction), a novel pretraining framework for LMMs. ViCToR employs a learnable visual token pool and utilizes the Hungarian matching algorithm to select semantically relevant tokens from this pool for visual token replacement. Furthermore, by integrating a visual token reconstruction loss with dense semantic supervision, ViCToR can learn tokens which retain high visual detail, thereby enhancing the large language model's (LLM's) understanding of visual information. After pretraining on 3 million publicly accessible images and captions, ViCToR achieves state-of-the-art results, improving over LLaVA-NeXT-8B by 10.4%, 3.2%, and 7.2% on the MMStar, SEEDI, and RealWorldQA benchmarks, respectively. Yin Xie, Kaicheng Yang 0002, Peirou Liang, Xiang An, Yongle Zhao, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng |
AAAI | 1 |
| 2025 | Quad-Pixel Image Defocus Deblurring: A New Benchmark and ModelabstractDefocus deblurring is a challenging task due to the spatially varying blur. Recent works have shown impressive results in data-driven approaches using dual-pixel (DP) sensors. Quad-pixel (QP) sensors represent an advanced evolution of DP sensors, providing four distinct sub-aperture views in contrast to only two views offered by DP sensors. However, research on QP-based defocus deblurring is scarce. In this paper, we propose a novel end-to-end learning-based approach for defocus deblurring that leverages QP data. To achieve this, we design a QP defocus and all-in-focus image pair acquisition method and provide a QP Defocus Deblurring (QPDD) dataset containing 4,935 image pairs. We then introduce a Local-gate assisted Mamba Network (LMNet), which includes a two-branch encoder and a Simple Fusion Module (SFM) to fully utilize features of sub-aperture views. In particular, our LMNet incorporates a Local-gate assisted Mamba Block (LAMB) that mitigates local pixel forgetting and channel redundancy within Mamba, and effectively captures global and local dependencies. By extending the defocus deblurring task from a DP-based to a QP-based approach, we demonstrate significant improvements in restoring sharp images. Comprehensive experimental evaluations further indicate that our approach outperforms state-of-the-art methods. Yin Xie, Xiaoxiu Peng, Lihu Sun, Wenkai Su, Chengming Liu |
CVPR | 2 |
| 2025 | Region-based Cluster Discrimination for Visual Representation LearningabstractLearning visual representations is foundational for a broad spectrum of downstream tasks. Although recent vision-language contrastive models, such as CLIP and SigLIP, have achieved impressive zero-shot performance via large-scale vision-language alignment, their reliance on global representations constrains their effectiveness for dense prediction tasks, such as grounding, OCR, and segmentation. To address this gap, we introduce Region-Aware Cluster Discrimination (RICE), a novel method that enhances region-level visual and OCR capabilities. We first construct a billion-scale candidate region dataset and propose a Region Transformer layer to extract rich regional semantics. We further design a unified region cluster discrimination loss that jointly supports object and OCR learning within a single classification framework, enabling efficient and scalable distributed training on large-scale data. Extensive experiments show that RICE consistently outperforms previous methods on tasks, including segmentation, dense detection, and visual perception for Multimodal Large Language Models (MLLMs). The pre-trained models have been released at https://github.com/deepglint/MVT. Yin Xie, Kaicheng Yang 0002, Xiang An, Yongle Zhao, Weimo Deng, Zimin Ran, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng |
ICCV | 1 |
| 2025 | RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation ParadigmabstractAfter pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of multimodal interleaved documents remains underutilized for contrastive vision-language representation learning. To fully leverage these unpaired documents, we initially establish a Real-World Data Extraction pipeline to extract high-quality images and texts. Then we design a hierarchical retrieval method to efficiently associate each image with multiple semantically relevant realistic texts. To further enhance fine-grained visual information, we propose an image semantic augmented generation module for synthetic text production. Furthermore, we employ a semantic balance sampling strategy to improve dataset diversity, enabling better learning of long-tail concepts. Based on these innovations, we construct RealSyn, a dataset combining realistic and synthetic texts, available in three scales: 15M, 30M, and 100M. We compare our dataset with other widely used datasets of equivalent scale for CLIP training. Models pre-trained on RealSyn consistently achieve state-of-the-art performance across various downstream tasks, including linear probe, zero-shot transfer, zero-shot robustness, and zero-shot retrieval. Furthermore, extensive experiments confirm that RealSyn significantly enhances contrastive vision-language representation learning and demonstrates robust scalability. The code will be released in https://garygutc.github.io/RealSyn. Tiancheng Gu, Kaicheng Yang 0002, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai, Jiankang Deng |
ACM Multimedia | 4 |
| 2025 | Grounding Deliberate Reasoning in Multimodal Large Language Models
Yuxuan Liu 0011, Dehu Li, Xiang An, Weimo Deng, Ziyong Feng, Yongle Zhao, Yin Xie |
MMM (2) | 8 |
| 2025 | UniViT: Unifying Image and Video Understanding in One Vision EncoderabstractDespite the impressive progress of recent pretraining methods on multimodal tasks, existing methods are inherently biased towards either spatial modeling (e.g., CLIP) or temporal modeling (e.g., V-JEPA), limiting their joint capture of spatial details and temporal dynamics. To this end, we propose UniViT, a cluster-driven unified self-supervised learning framework that effectively captures the structured semantics of both image spatial content and video temporal dynamics through event-level and object-level clustering and discrimination. Specifically, we leverage offline clustering to generate semantic clusters across both modalities. For videos, multi-granularity event-level clustering progressively expands from single-event to structured multi-event segments, capturing coarse-to-fine temporal semantics; for images, object-level clustering captures fine-grained spatial semantics. However, while global clustering provides semantically consistent clusters, it lacks modeling of structured semantic relations (e.g., temporal event structures). To address this, we introduce a contrastive objective that leverages these semantic clusters as pseudo-label supervision to explicitly enforce structural constraints, including temporal event relations and spatial object co-occurrences, capturing structured semantics beyond categories. Meanwhile, UniViT jointly embeds structured object-level and event-level semantics into a unified representation space. Furthermore, UniViT introduces two key components: (i) Unified Rotary Position Embedding integrates relative positional embedding with frequency-aware dimension allocation to support position-invariant semantic learning and enhance the stability of structured semantics in the discrimination stage; and (ii) Variable Spatiotemporal Streams adapt to inputs of varying frame lengths, addressing the rigidity of conventional fixed-input approaches. Extensive experiments across varying model scales demonstrate that UniViT achieves state-of-the-art performance on linear probing, attentive probing, question answering, and spatial understanding tasks. Xiang An, Yin Xie, Kaicheng Yang 0002, Zimin Ran, Muhammad Imran Razzak, Ziyong Feng, Behzad Bozorgtabar, Jiankang Deng, ZongYuan Ge |
NeurIPS | 4 |
| 2023 | Computation of Mobile Phone Collaborative Embedded Devices for Object Detection TaskabstractIn the past decade, computer vision has developed rapidly, and its application scenarios are increasing. But in the process of its application, the limited embedded compute capability is still one of the most important reasons hindering its development. In contrast, with the continuous improvement of mobile computing capability in recent years, the reasoning of neural network models on mobile phones has become a closer and closer fact. The most of tasks of computer vision are continuous and fixed order of the calculation processes. According to the characteristic, we propose a method for collaborative embedded inference on mobile phones. This method divides computer vision tasks, moves part of the calculation to the mobile phone, and runs in a pipeline scheme to achieve the effect of accelerating inference. This method can realize the running acceleration of such tasks and reducing the computational burden of the embedded platform. Codes are available at https://github.com/yiyexy/pipeline. Yin Xie, Yigui Luo, Haihong She, Zhaohong Xiang |
CSCWD | 1 |
| 2023 | Neural Network Model Pruning without Additional Computation and Structure RequirementsabstractIn past work, deep learning researchers always designed hyperparameters such as model structure and learning rate first and then used the training set to train the weights in this model. While unrestricted model structure design leads to massive neuron redundancy in neural network models. By pruning these redundant neurons, not only can the storage be compressed effectively, but also the operation can be accelerated. In this paper, we propose a method to utilize the training set to prune the model structure during training: 1) train the initialized model and bring it to basic convergence; 2) feed the entire training set into the model and calculate the activations of neurons in each layer; 3) calculate the threshold for neuron pruning in each layer according to the pruning ratio, delete neurons whose activation value is lower than the threshold, and correspondingly delete the weights of the upper and lower layers; 4) further train the pruned model so that it eventually converges. This method of deleting redundant neurons not only greatly deletes the parameters in the model but also achieves model acceleration. We applied this method to some mainstream neural network models: VGGNet and ResNet, and achieved good results. Yin Xie, Yigui Luo, Haihong She, Zhaohong Xiang |
CSCWD | 1 |
| 2023 | Accurate Latency Prediction of Deep Learning Model Inference Under Dynamic Runtime Resource
Haihong She, Yigui Luo, Zhaohong Xiang, Weiming Liang, Yin Xie |
ICONIP (7) | 5 |
| 2023 | CDCP: A Framework for Accelerating Dynamic Channel Pruning MethodsabstractDynamic channel pruning is a technique aimed at reducing the theoretical computational complexity and inference latency of convolutional neural networks. Dynamic channel pruning methods introduce complex additional modules for dynamically selecting channels for images. Due to the additional modules, dynamic channel pruning methods never achieve optimal acceleration effect in real world. To address this problem, we propose Consecutive Dynamic Channel Pruning (CDCP), a novel dynamic channel pruning framework unified for almost all dynamic pruning methods designed for continuous image processing. The core idea of CDCP stems from our observation that adjusting the network for all frames in semantically continuous scenes is unnecessary since adjacent frames often share similar network structures in dynamic channel pruning. CDCP introduces a simple binary classifier to determine whether the network structure needs to be adjusted for a new frame. Our method can also be used for semantically non-continuous image processing tasks with a slightly lower probability of model reuse. We validate the effectiveness of CDCP on three dynamic channel pruning methods and better acceleration effects are achieved when applied them with CDCP to the semantically continuous Waymo dataset, the nuScenes dataset, and the semantically discontinuous COCO dataset. Zhaohong Xiang, Yigui Luo, Yin Xie, Haihong She, Weiming Liang, Laigang Zhao |
ICPADS | 3 |
| 2017 | Codeon: On-Demand Software Development AssistanceabstractSoftware developers rely on support from a variety of resources---including other developers---but the coordination cost of finding another developer with relevant experience, explaining the context of the problem, composing a specific help request, and providing access to relevant code is prohibitively high for all but the largest of tasks. Existing technologies for synchronous communication (e.g. voice chat) have high scheduling costs, and asynchronous communication tools (e.g. forums) require developers to carefully describe their code context to yield useful responses. This paper introduces Codeon, a system that enables more effective task hand-off between end-user developers and remote helpers by allowing asynchronous responses to on-demand requests. With Codeon, developers can request help by speaking their requests aloud within the context of their IDE. Codeon automatically captures the relevant code context and allows remote helpers to respond with high-level descriptions, code annotations, code snippets, and natural language explanations. Developers can then immediately view and integrate these responses into their code. In this paper, we describe Codeon, the studies that guided its design, and our evaluation that its effectiveness as a support tool. In our evaluation, developers using Codeon completed nearly twice as many tasks as those who used state-of-the-art synchronous video and code sharing tools, by reducing the coordination costs of seeking assistance from other developers. Yan Chen 0033, Sang Won Lee 0002, Yin Xie, Yiwei Yang 0004, Walter S. Lasecki, Steve Oney |
CHI | 3 |
| 2012 | Outage Performance of Cognitive Relay Networks with Primary User's ISR ConstraintabstractIn the underlay spectrum sharing systems, secondary users (SUs) are allowed to transmit their data in the licensed spectrum band when primary users(PUs) are also transmitting, as long as the transmission of SUs do not interfere PUs' communications. In cognitive relay networks, the source and relay nodes both need to tune their transmit power to mitigate the interference to PU. In this paper, we investigate the outage performance of cognitive relay networks with PU's interference to signal ratio (ISR) constraint, where both average and peak ISR constraint are considered. Finally, We derive the exact outage probability in the scenario without cooperation and the upper bound of outage probability in the scenario with cooperation. Zhiqing Wei, Yin Xie, Qixun Zhang |
VTC Fall | 2 |
| 2012 | Outage Probability Analysis of Cognitive Relay Networks in Nakagami-m Fading ChannelsabstractIn spectrum sharing systems, a secondary user (SU) is permitted to share frequency bands with a primary user (PU) as long as its transmission does not interfere with the PU's communication. In this paper, the outage probability is investigated for the cognitive relay system over Nakagami-m fading channel. By applying the interference temperature constraints at the source nodes and relay nodes in secondary systems, we analyze the outage performance in two-hop underlay spectrum sharing with the best relay selection criterion. The probability density function (PDF) and cumulative distribution function (CDF) of the signal to noise ratio (SNR) at the SU's receiver are derived to obtain the closed-form upper bound of the outage probability of the secondary relay system. Simulations results demonstrate the validity and accuracy of the theoretical analysis. Yifan Zhang 0003, Yin Xie, Yang Liu 0024, Zhiyong Feng 0001, Ping Zhang 0003, Zhiqing Wei |
VTC Fall | 2 |