Rui Zhong 0005

dblp:95/305-5 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-0126-7543ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From BOPPPS Stages to Cognitive-Adaptive Prompts: Controlling Instructional Drift in LLM-Based Tutoring Dialogues
Longxiang Du, Rui Zhong 0005, Zhenwan Zhu, Shi Dong 0004, Shuyuan Sun
AIED2
2026 A Synergistic Framework for Cognitive Diagnosis with LLM-Empowered Data Augmentation
abstract
Accurate cognitive diagnosis is crucial for enabling personalized learning. However, its effectiveness is severely hampered by data sparsity and class imbalance in student-exercise interactions. To address these challenges, this paper proposes a synergistic framework that integrates Artificial Intelligence and Learning Analytics. First, to mitigate data sparsity, we leverage a Large Language Model (LLM) with chain-of-thought prompting to enrich the Q-matrix, thereby expanding knowledge concept coverage. Second, to address class imbalance, we introduce a Generative Adversarial Network (GAN) constrained by this refined knowledge structure to synthesize pedagogically meaningful negative samples. Finally, a dual-level graph attention network is employed to infer precise knowledge states from both the LLM-enriched Q-matrix and the balanced response data. Extensive experiments on the Junyi dataset demonstrate that our framework significantly outperforms strong baselines, achieving a 6.21% improvement in AUC and increasing specificity from 54.24% to 85.83%. These results underscore its potential for delivering reliable and actionable learning diagnostics.
Zhenwan Zhu, Rui Zhong 0005, Longxiang Du
LAK2
2025 Global-to-Local Color Correction with Full-Region Coverage for Multi-view Light Field Images
abstract
Color correction methods for multi-view images are typically divided into global-based and local-based approaches. Global methods perform global color mapping but fail to address local differences, leading to local color inconsistencies. Local methods focus on regional color mapping based on distributions or semantics but struggle with sparse semantic correspondences and are sensitive to lighting and noise. To address these, we propose a hybrid color correction method for light field images. First, a global color correction is applied to ensure overall correction. Moreover, an object matching algorithm is designed to match regions and calculate both global and local similarities between images to refine the corrections. Next, a local optimization module is introduced to optimize the adjustment of specific regions. Finally, gradient preservation is incorporated to maintain structural consistency. Experiments on our proposed dataset using a 3×3 light field camera array demonstrate that our method outperforms existing approaches.
Yixu Huang, Rui Zhong 0005, Ségolène Rogge, Adrian Munteanu 0001
ICME2
2025 Leveraging label semantics and meta-label refinement for multi-label question classification
Shi Dong 0004, Xiaobei Niu, Rui Zhong 0005, Zhifeng Wang 0001, Mingzhang Zuo
Knowl. Based Syst.3
2025 LF-GS: 3D Gaussian Splatting for View Synthesis of Multi-View Light Field Images
abstract
3D Gaussian Splatting (3D-GS) has emerged as a groundbreaking approach for view synthesis. However, when applied to light field image synthesis, the issue of a too narrow field of view (FOV) that leaves some areas uncovered, compounded by the problem of data sparsity, significantly compromises the quality of synthesized views using 3D-GS. To overcome these limitations, we present LF-GS, a specialized 3D-GS variant optimized for light field image synthesis. Our methodology incorporates two key innovations. First, by harnessing the unique advantage of light field sub-aperture images that provide dense geometric cues, our method enables the effective incorporation of enhanced depth and normal priors derived from light field images. This allows for more accurate depth than monocular depth estimation. Second, unlike other methods that struggle to control the generation of unreasonable Gaussians, we introduce adaptive regularization mechanisms. These mechanisms strategically regulate Gaussian opacity and spatial scale during optimization, thereby preventing model overfitting and preserving essential scene details. Comprehensive experiments on our newly constructed light field dataset demonstrate that LF-GS achieves significant quality improvements over 3D-GS.
Yixu Huang, Rui Zhong 0005, Ségolène Rogge, Adrian Munteanu 0001
IEEE Signal Process. Lett.2
2023 Semantic Representation and Attention Alignment for Graph Information Bottleneck in Video Summarization
abstract
End-to-end Long Short-Term Memory (LSTM) has been successfully applied to video summarization. However, the weakness of the LSTM model, poor generalization with inefficient representation learning for inputted nodes, limits its capability to efficiently carry out node classification within user-created videos. Given the power of Graph Neural Networks (GNNs) in representation learning, we adopted the Graph Information Bottle (GIB) to develop a Contextual Feature Transformation (CFT) mechanism that refines the temporal dual-feature, yielding a semantic representation with attention alignment. Furthermore, a novel Salient-Area-Size-based spatial attention model is presented to extract frame-wise visual features based on the observation that humans tend to focus on sizable and moving objects. Lastly, semantic representation is embedded within attention alignment under the end-to-end LSTM framework to differentiate indistinguishable images. Extensive experiments demonstrate that the proposed method outperforms State-Of-The-Art (SOTA) methods.
Rui Zhong 0005, Rui Wang 0107, Wenjin Yao, Shi Dong 0004, Adrian Munteanu 0001
IEEE Trans. Image Process.1
2021 Visual and Semantic Feature Coordinated Bi-Lstm Model for Unsupervised Video Summarization
abstract
While dealing with user-created video, the prior methods suffer from the problem of high redundancy among keyframes. To address the critical issue, we present a Visual and Semantic Feature coordinated Bi-LSTM (VSFB) model for unsupervised video summarization. First, a novel Salient-Area-Size-based spatial attention model is presented to extract frame-wise visual features on the observation that humans tend to focus on sizable and moving objects. Second, the visual features are integrated with semantic features processed by Bi-LSTM to refine the frame-wise probability of being selected as keyframes. Finally, an index adjusted diversity and representativeness reward is utilized to reinforce the learning operation of the VSFB model in the video summarization. Extensive experiments demonstrate that our method outperforms state-of-the-art methods in terms of the F-score.
Zhiqiang Hong, Rui Zhong 0005
ICME2
2021 Unsupervised Learning of Visual and Semantic Features for Video Summarization
abstract
The high redundancy among keyframes is a critical issue for the existing summarizing methods in dealing with user-created videos. To address the critical issue, we present an unsupervised learning method, Spatial Attention Model guided Bi-directional Long Short-term Memory network (Bi-LSTM), on the combination of visual and semantic features. As for the visual feature, we design a Salient-Area- Size-based spatial attention model on the observation that humans tend to focus on sizable and moving objects in videos. Moreover, the Bi-LSTM network is leveraged to exploit the semantic feature. Afterward, the Soft Selected Probability generated from the spatial attention and semantic feature is fused to obtain the final probability for keyframe selection. The reinforcement learning framework, trained by the Deep Deterministic Policy Gradient algorithm, is adopted to do unsupervised training. Extensive experiments on the SumMe and TVSum datasets demonstrate that our method outperforms the state-of-the-art methods in terms of F-score.
Yansen Huang, Rui Zhong 0005, Wenjin Yao, Rui Wang 0107
ISCAS2
2021 Graph Attention Networks Adjusted Bi-LSTM for Video Summarization
abstract
The high redundancy among keyframes is a critical issue for the prior summarizing methods in dealing with user-created videos. To address the critical issue, we present a Graph Attention Networks (GAT) adjusted Bi-directional Long Short-term Memory (Bi-LSTM) model for unsupervised video summarization. First, the GAT is adopted to transform an image's visual features into higher-level features by the Contextual Features based Transformation (CFT) mechanism. Specifically, a novel Salient-Area-Size-based spatial attention model is presented to extract frame-wise visual features on the observation that humans tend to focus on sizable and moving objects. Second, the higher-level visual features are integrated with semantic features processed by Bi-LSTM to refine the frame-wise probability of being selected as keyframes. Extensive experiments demonstrate that our method outperforms state-of-the-art methods.
Rui Zhong 0005, Rui Wang 0107, Zhiqiang Hong
IEEE Signal Process. Lett.1
2019 Dictionary Learning-Based, Directional, and Optimized Prediction for Lenslet Image Coding
abstract
In this paper, a novel approach to encode lenslet (LL) images is proposed. The method departs from traditional block-based coding structures and employs a hexagonal-shaped pixel cluster, called macro-pixel, as an elementary coding unit. A novel prediction mode based on dictionary learning is proposed, whereby macro-pixels are represented by a sparse linear combination of atoms from a generic dictionary. Additionally, an optimized linear prediction mode and a directional prediction mode specifically designed for macro-pixels are proposed. Rate-distortion optimization is utilized to select the best intra prediction mode for each macro-pixel. Experimental results on the light field image data set show that the proposed coding system outperforms HEVC and the state-of-the-art in LL image coding with an average peak signal to noise ratio gain of 3.33 and 1.41 dB, respectively, and with rate savings of 67.13% and 34.30%, respectively.
Rui Zhong 0005, Ionut Schiopu, Bruno Cornelis, Shao-Ping Lu, Junsong Yuan 0001, Adrian Munteanu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Efficient directional and L1-optimized intra-prediction for light field image compression
abstract
Light field images can be conveniently captured by consumer-level plenoptic cameras. However, as the resulting data rates are very high, providing efficient compression for this type of data is of critical importance. This remains an open problem which has recently attracted a lot of attention from the coding community. State-of-the-art compression systems prove to be inefficient when directly applied on this type of data due to the inherent spatial discontinuities in light field images. In this paper, a novel intra-prediction method for disk-shaped pixel clusters is proposed. An L1 minimization of the prediction residuals is performed followed by clustering of the predictors, leading to an optimized set of predictors for the macro-pixels. Furthermore, directional intra-prediction modes based on HEVC are devised for the macro-pixels. Experimental results obtained on the EPFL light field image dataset demonstrate that the proposed coding scheme yields an average of 3.22 dB and 1.45 dB gain in PSNR, and 59.6% and 30.88% average rate savings compared to HEVC and the state-of-the-art in light field image coding respectively.
Rui Zhong 0005, Shizheng Wang, Bruno Cornelis, Yuanjin Zheng, Junsong Yuan 0001, Adrian Munteanu 0001
ICIP1
2016 L1-optimized linear prediction for light field image compression
abstract
The advent of consumer-level plenoptic cameras has sparkled the interest towards the design of efficient compression techniques for light field images. State-of-the-art compression systems such as HEVC prove to be inefficient when directly applied on this type of data due to the inherent spatial discontinuities among neighboring microlens images. In this paper, a novel light field image compression system is proposed. The disk-shaped pixel clusters corresponding to each microlens in the light field image are efficiently predicted based on the neighboring disks. In this context, an optimized linear prediction design based on L1 minimization of the residuals is proposed. K-means clustering is employed on training data in order to determine the optimized set of predictors. The experimental results on an extensive set of light field images demonstrate that the proposed coding scheme yields an average of 2.93 dB and 3.22 dB gain in PSNR, and 52.67% and 57.27% average rate savings compared to HEVC and JPEG2000 respectively.
Rui Zhong 0005, Shizheng Wang, Bruno Cornelis, Yuanjin Zheng, Junsong Yuan 0001, Adrian Munteanu 0001
PCS1