VLDB 2026 Research / reviewers in the wild / expert
Hengrui Li
dblp:332/5994
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reconstructing Temporal Heterogeneity: A Multidomain Collaborative Analysis Framework for Robust Time-Series Forecasting
Hengrui Li, Wenxue Cui, Yifeng Wang 0001, Chunshan Dong, Wenju Li, Jiangpeng Shi, Yongbing Zhang 0002, Shaohui Liu |
IEEE Internet Things J. | 1 |
| 2025 | Multi-scale Context Intertwining for Panoramic Renal Pathology SegmentationabstractPanoramic segmentation of renal pathological tissues plays a crucial role in diagnosing renal carcinoma and other kidney-related diseases. The multi-scale nature of kidney tissues, which requires different magnification levels for accurate analysis, presents a significant challenge for segmentation models. In this work, we propose a Multi-scale Context Intertwining Network (MCINet) to address this issue. Our approach utilizes an auxiliary interaction network to enhance feature interaction between different scales and generate pseudo-labels for unannotated structures. By incorporating exponential moving average strategies, we ensure seamless feature integration across scales. Extensive experiments demonstrate that MCINet outperforms state-of-the-art models in key metrics such as Dice and Hausdorff Distance, proving its efficacy in renal tissue segmentation tasks. Ye Zhang 0043, Xianchao Guan, Hengrui Li, Xiangming Yan, Ziyue Wang 0005, Yongbing Zhang 0002 |
ICASSP | 3 |
| 2025 | MAMF-Net: Modality-Adaptive Masked Fusion Network for Speech Emotion RecognitionabstractThis paper introduces a novel multimodal emotion recognition model, the Modality-Adaptive Masked Fusion Network (MAMF-Net), designed to mitigate information loss and improve cross-modal alignment during the fusion of speech and text modalities. MAMF-Net employs an audio-guided text encoder to enhance the semantic representation of text by leveraging the temporal resolution and contextual information inherent in speech, thereby ensuring accurate alignment of modal features. Additionally, the model utilizes a modality transfer-based MAE masking strategy, which effectively captures complementary information between modalities by partially masking transferred information, thus improving fusion effectiveness and system stability. The experimental results show that MAMF-Net outperforms existing methods on datasets such as CMU-MOSI and CMU-MOSEI, highlighting its significant potential for multimodal emotion analysis. Hengrui Li, Xiaopei Chen, Shaohui Liu |
ICME | 1 |
| 2025 | AMH-Net: Adaptive Multi-Band Hybrid-Aware Network for Emotion Recognition in SpeechabstractSpeech emotion recognition (SER) technology analyzes speech signals to automatically identify the speaker's emotional state. However, existing methods overlook feature extraction based on human acoustic characteristics. In this paper, we propose AMH-Net, an Adaptive Multi-band Hybridaware Network designed for SER. The model leverages formant (F1, F2, F3) from speech science, which describe the human vocal tract, to partition speech signals into multiple frequency bands. A variable-depth residual network structure is employed for more precise extraction of emotional characteristics. In addition, a hybrid attention mechanism is integrated to combine information, resulting in a more comprehensive emotional representation. Experimental evaluations of six diverse datasets show that AMH-Net outperforms state-of-the-art methods, achieving improvements of 2.11% and 2.64% in average UAR and WAR, respectively, on each corpus. The code is publicly available at https://github.com/hengruili1997/AMH-net Hengrui Li, Yongbing Zhang 0002, Shaohui Liu |
IEEE Signal Process. Lett. | 1 |
| 2025 | FDDCC-VSR: a lightweight video super-resolution network based on deformable 3D convolution and cheap convolution
Hengrui Li |
Vis. Comput. | 3 |
| 2024 | Incremental Model Predictive Control for Velocity-Controlled Robot ManipulatorsabstractIn this article, an incremental model predictive controller is proposed for the velocity-controlled robot manipulator. First, the time-delay estimation (TDE) technique is used to approximate the unknown discrepancy between the command and the real joint velocity due to velocity dynamics, and the equation of motion in the incremental form is obtained. Then, on the basis of the resulting equation of motion and taking into account joint position and velocity constraints, the incremental model predictive controller is developed by formulating a constrained optimal control problem (OCP). The constrained OCP is cast to a quadratic programming (QP) problem, making it possible to employ computationally efficient QP solvers to solve the constrained OCP. As a result, real-time control is guaranteed. Finally, experiments on a robot manipulator are implemented to verify the effectiveness of the developed incremental model predictive controller. Hengrui Li, Cong Li 0015, Fangzhou Liu 0001 |
IECON | 2 |
| 2023 | SCCADC-SR: a real image super-resolution based on self-calibration convolution and adaptive dense connection
Xin Yang 0002, Hengrui Li, Chenhuan Wu, Tao Li 0011 |
Multim. Tools Appl. | 2 |
| 2023 | HIFGAN: A High-Frequency Information-Based Generative Adversarial Network for Image Super-ResolutionabstractSince the neural network was introduced into the super-resolution (SR) field, many SR deep models have been proposed and have achieved excellent results. However, there are two main drawbacks: one is that the methods based on the best peak-signal-to-noise ratio (PSNR) do not have enough comfortable visual quality; the other is that although the SR models based on generative adversarial network (GAN) have satisfactory visual quality, the structure of the reconstructed image has apparent defects. Therefore, according to the characteristics that human eyes are sensitive to high-frequency components in images, this article proposes an improved image SRGAN model based on high-frequency information fusion (HIFGAN). It builds a feature extraction network for high-frequency information fusion by designing a lightweight spatial attention module and improving the network architecture of enhanced super-resolution GAN (ESRGAN). It makes the generator in the GAN network have better feature recovery ability, reduces the dependence of the later training on the decider and loss function, and makes the generated image structure more consistent with the real situation. In addition, we build a high-frequency loss function to optimize the training of the generator network. Detailed experimental results show that HIFGAN performs excellently in both objective criterion evaluation and subjective visual effect. Compared with the state-of-the-art GAN-based SR networks, the reconstructed image by our model is more precise and complete in texture details. Xin Yang 0002, Hengrui Li, Tao Li 0011 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |