Weichao Chen 0001

dblp:98/120-1 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-7226-7885ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scaling Multimodal Retrieval and Generation for Long Documents through Visual Tiling and Context Compression
abstract
Multimodal retrieval-augmented generation (MRAG) provides a powerful paradigm for long-document reasoning by integrating retrieval with generative modeling. However, scaling MRAG to multi-page documents remains challenging due to noisy cross-modal retrieval and the quadratic computational cost of long-context inference. Existing multimodal large language models (MLLMs) lack mechanisms to structure visual representations and efficiently manage contextual memory, limiting their scalability and generalization. We propose LMDocRag, a principled framework for efficient multimodal retrieval and generation over long documents. Our approach is based on the insight that document understanding benefits from structured visual decomposition and representation-level compression. Specifically, we introduce a document tiling strategy that transforms long document images into semantically localized visual units, enabling unified hybrid retrieval in a structured embedding space. Building on this decomposition, we develop a layout-guided visual token sparsification method that learns to preserve structurally salient regions while suppressing redundant visual representations, and a KV-cache compression scheme that reduces autoregressive memory growth by selectively retaining informative contextual states. Extensive experiments on multimodal long-document benchmarks demonstrate that LMDocRag improves question-answering accuracy while substantially reducing visual token count and inference complexity.
Weichao Chen 0001, Shengjie Zhao 0001
ICMR2
2025 SAVE-GSL: Scalable and Expressive Graph Structure Learning for Large Graphs
abstract
Graph structure learning (GSL) has emerged as a promising approach for optimizing graph structures to enhance downstream task performance. However, the quadratic complexity of GSL renders it impractical for large graphs, such as social networks. While several attempts have been made to mitigate the scalability issue of GSL, they struggle to ensure efficiency and expressiveness simultaneously. In this paper, we propose a novel scalable and expressive graph structure learning framework (SAVE-GSL), that models all-pair interactions with linear complexity. Specifically, we design cluster and expander sparse patterns for efficient expressiveness enhancement, and adaptively fuse them to generate the optimized graph. Moreover, leveraging these sparse patterns, we theoretically prove that SAVE-GSL is an efficient universal approximator of permutation-equivariant functions, providing a formal justification for its superior expressiveness. Extensive experiments demonstrate that SAVE-GSL outperforms state-of-the-art schemes in both efficiency and accuracy on social network datasets and other graph datasets of various scales.
Manxin Xu, Shengjie Zhao 0001, Jin Zeng 0004, Weichao Chen 0001, Shilong Dong
ICME4
2025 Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers
abstract
Contrastive Language-Image Pre-training (CLIP) has attracted a surge of attention for its superior zero-shot performance and excellent transferability to downstream tasks. However, training such large-scale models usually requires substantial computation and storage, which poses barriers for users with consumer-level computers. Motivated by this observation, in this paper we investigate how to achieve competitive performance on a single Nvidia RTX3090 GPU and with one terabyte of storage for the dataset. On one hand, we simplify the transformer block structure and combine Weight Inheritance with multi-stage Knowledge Distillation (WIKD), thereby reducing the number of parameters and improving the inference speed during training as well as deployment. On the other hand, confronted with the convergence challenge posed by a limited dataset, we generate synthetic captions for each sample as data augmentation, and devise a novel Pair Matching (PM) loss to fully exploit the distinction among positive and negative image-text pairs. Extensive experiments demonstrate that our model can achieve a new state-of-the-art datascale-parameter-accuracy tradeoff, which could further popularize the CLIP model and enable its deployment in consumer devices.
Shengjie Zhao 0001, Weichao Chen 0001, Deniz Gündüz
IJCNN3
2025 ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
abstract
Recent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicated benchmark for assessing VLMs’ understanding of cinematic language. ShotBench includes 3,049 still images and 500 video clips drawn from more than 200 films, with each sample annotated by trained annotators or curated from professional cinematography resources, resulting in 3,608 high-quality question-answer pairs. We conduct a comprehensive evaluation of over 20 state-of-the-art VLMs across eight core cinematography dimensions. Our analysis reveals clear limitations in fine-grained perception and cinematic reasoning of current VLMs. To improve VLMs capability in cinematography understanding, we construct a large-scale multimodal dataset, named ShotQA, which contains about 70k Question-Answer pairs derived from movie shots. Besides, we propose ShotVL and train this VLM model with a two-stage training strategy, integrating both supervised fine-tuning and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that our model achieves substantial improvements, surpassing all existing strongest open-source and proprietary models evaluated on ShotBench, establishing a new state-of-the-art performance.
Jingwen He, Dian Zheng, Yuhao Dong, Fan Zhang 0045, Yinan He, Weichao Chen 0001, Yu Qiao 0001, Wanli Ouyang, Shengjie Zhao 0001, Ziwei Liu 0002
NeurIPS9
2023 Computation of Rate-Distortion-Perception Functions With Wasserstein Barycenter
abstract
The nascent field of Rate-Distortion-Perception (RDP) theory is seeing a surge of research interest due to the application of machine learning techniques in the area of lossy compression. The information RDP function characterizes the three-way trade-off between description rate, average distortion, and perceptual quality measured by discrepancy between probability distributions. However, computing RDP functions has been a challenge due to the introduction of the perceptual constraint, and existing research often resorts to data-driven methods. In this paper, we show that the information RDP function can be transformed into a Wasserstein Barycenter problem. The non-strictly convexity brought by the perceptual constraint can be regularized by an entropy regularization term. We prove that the entropy regularized model converges to the original problem. Furthermore, we propose an alternating iteration method based on the Sinkhorn algorithm to numerically solve the regularized optimization problem. Experimental results demonstrate the efficiency and accuracy of the proposed algorithm.
Chunhui Chen 0005, Xueyan Niu 0001, Wenhao Ye, Shitong Wu, Bo Bai 0001, Weichao Chen 0001, Sian-Jheng Lin
ISIT6
2023 Real-Time Super-Resolution: A New Mechanism for XR over 5G-Advanced
abstract
Extended Reality (XR) has attracted great attention from both academic and industry, for providing users with an immersive experience anywhere. Nowadays, XR video streaming service is evolving to high definition (HD), which results in massive data traffic with more stringent latency requirement. Due to the two characteristics, it is challenging to support the commercial use for XR service in the current New Ratio (NR) network. In this paper, we propose a real-time super-resolution (RTSR) framework for XR HD video transmission. The basic idea is to utilize the overfitting feature of Deep Neural Network (DNN) to learn the non-linear mapping between low-definition (LD) video frames and HD video frames. The cloud XR server can transmit the LD frames together with the dedicated super resolution (SR) models instead of sending HD frames directly. The receiver can recover the HD frames locally with the inferencing ability of SR model. In addition, by introducing the online training and layered transmission strategy, the SR model update period can be adaptively adjusted according to the scenario changes, which also reduces the transmission overhead. Simulation results demonstrate the superiority of our proposed RTSR, which can save up to 50% traffic and increase the XR capacity about 40% compared with the conventional SR scheme. In terms of the system capacity, our results show that the average number of UEs can reach about 23 per cell under the common settings of Dense Urban.
Weichao Chen 0001, Youlong Cao, Erkai Chen, Guohua Zhou, Weichao Li 0001
WCNC1
2021 XR Quality Index: Evaluating RAN Transmission Quality for XR services over 5G and Beyond
abstract
Recently, eXtended Reality (XR) has gained significant attention for providing immersive user experience in various applications with the help of 5G new radio (NR) network. To ensure the quality of experience (QoE) in XR services, wireless transmission plays an important role. In this paper, we propose a novel performance metric that can reflect the impact of network transmission on XR services, naming it as XR Quality Index (XQI). Specifically, fine-grained and coarse-grained XQI models are provided with detailed calculation procedures. The aim of XQI is to produce a final score that can reflect the impact of network transmission on QoE, and then could be used for network optimization. Simulation results verify the effects of XQI in evaluating the quality of XR videos.
Shengyue Dou, Shuri Liao, Kedi Wu, Erkai Chen, Weichao Chen 0001, Nijun Li
PIMRC6
2021 UAV-Assisted Data Collection With Nonorthogonal Multiple Access
abstract
Unmanned aerial vehicles (UAVs) facilitate information collection greatly in the Internet-of-Things (IoT) systems due to their superior flexibility and mobility. On the other hand, nonorthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in fifth-generation networks. The integration of NOMA into UAV-assisted wireless networks shows great potential, but how to determine the user grouping and power allocation in NOMA according to the high mobility of UAV is challenging. In this article, we propose a general NOMA-enabled UAV-assisted data collection (NUDC) protocol to maximize the sum rate of a wireless sensor network (WSN), where the location of UAV, sensor grouping, and power control are jointly considered. Moreover, a joint signal-to-interference ratio (SIR) hypergraph-based grouping and power control (SHG-PC) NOMA scheme is provided to obtain the appropriate sensor grouping and the optimal power control solutions efficiently, in which the hypergraph and the greedy coloring algorithm are exploited to find out the optimized group relationships. Extensive simulation results demonstrate the efficiency of our proposed protocol.
Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Yi Chen 0013, Liuqing Yang 0001
IEEE Internet Things J.1
2021 Generalized User Grouping in NOMA Based on Overlapping Coalition Formation Game
abstract
Non-orthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in 5G systems. In most existing NOMA user grouping approaches, users are grouped into disjoint groups, which may lead to a waste of power resources within each NOMA group. Motivated by this, in this paper we propose a novel generalized user grouping (GuG) concept for NOMA from an overlapping perspective, which allows each user to participate in multiple groups but subject to individual maximum power constraint. In order to achieve effective GuG and maximize the system sum rate, we formulate a joint power control and GuG optimization problem. Then, we address this problem by exploiting the overlapping coalition formation (OCF) game framework, and we further propose an OCF-based algorithm in which each user can be self-organized into a desirable overlapping coalition structure. Simulation results verify the efficiency of GuG in NOMA systems and indicate that compared with traditional NOMA user grouping schemes, our proposed OCF-based GuG NOMA scheme achieves significant performance gains in terms of system sum rate.
Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Liuqing Yang 0001
IEEE J. Sel. Areas Commun.1
2021 Generalized User Grouping in NOMA: An Overlapping Perspective
abstract
Non-orthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in 5G systems. Traditionally, NOMA user grouping is non-overlapping, leading to a waste of power resources within each NOMA group. Motivated by this, in this paper we propose a novel generalized user grouping (GuG) concept for NOMA from an overlapping perspective, which allows each user to participate in multiple user groups but subject to individual maximum power constraint. In order to achieve effective GuG and maximize the system sum rate, we formulate a joint power control and GuG optimization problem. Then we further provide a machine learning-based GuG scheme to obtain the optimized feasible GuG and the optimal power control solutions efficiently, in which the established machine learning-based model is exploited to explore the relative relationships of channel gains of users and obtain several fixed grouping patterns via Merge operation. Simulation results verify the efficiency of GuG in NOMA systems and indicate that compared with traditional NOMA user grouping schemes, our proposed GuG scheme achieves significant performance gains in terms of system sum rate.
Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Hong Chen 0003, Liuqing Yang 0001
IEEE Trans. Wirel. Commun.1
2020 Generalized User Grouping in NOMA Based on Overlapping Coalition Formation Game
abstract
Non-orthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in 5G systems. In most existing NOMA user grouping approaches, users are grouped into disjoint groups, which may lead to a waste of power resources within each NOMA group. Motivated by this, in this paper we propose a novel generalized user grouping (GuG) concept for NOMA from an overlapping perspective, which allows each user to participate in multiple groups but subject to individual maximum power constraint. In order to achieve effective GuG and maximize the system sum rate, we formulate a joint power control and GuG optimization problem. Then, we address this problem by exploiting the overlapping coalition formation (OCF) game framework, and we further propose an OCF-based algorithm in which each user can be self-organized into a desirable overlapping coalition structure. Simulation results verify the efficiency of GuG in NOMA systems and show that our proposed OCF-based GuG NOMA scheme achieves significant performance gains in terms of system sum rate.
Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Yi Chen 0013, Liuqing Yang 0001
GLOBECOM1
2020 Machine Learning-Based Generalized User Grouping in NOMA
abstract
Non-orthogonal multiple access (NOMA) provides high spectral efficiency and supports massive connectivity in 5G systems. Traditionally, NOMA user grouping is non-overlapping, leading to a waste of power resources within each NOMA group. Motivated by this, we propose a novel generalized user grouping (GuG) concept for NOMA from an overlapping perspective, which allows each user to participate in multiple user groups but subject to individual maximum power constraint. We formulate a joint power control and GuG optimization problem, and then provide a machine learning-based GuG scheme to obtain the optimized feasible GuG and the optimal power control solutions efficiently. Simulation results show significant performance gains in terms of system sum rate.
Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Yi Chen 0013, Liuqing Yang 0001
GLOBECOM1
2020 UAV-Assisted Data Collection with Non-Orthogonal Multiple Access
abstract
Unmanned aerial vehicles (UAVs) facilitate information collection greatly in Internet of Things (IoT) systems. On the other hand, non-orthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in 5G networks. The integration of NOMA into UAV-assisted wireless networks shows great potential, but how to determine the user grouping and power allocation in NOMA according to the different locations of UAV is challenging. In this paper, we propose a general NOMA-enabled UAV-assisted data collection (NUDC) protocol to solve the formulated sum rate maximization problem such that the location of UAV, sensor grouping, and power control are jointly considered. Moreover, a joint signal-to-interference-ratio (SIR) hypergraph-based grouping and power control (SHG-PC) NOMA scheme is provided to obtain the appropriate sensor grouping and the optimal power control solutions efficiently. Extensive simulation results demonstrate the effectiveness of our proposed protocol.
Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Liuqing Yang 0001
WCNC1