Rumin Zhang

dblp:158/1364 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Mental-Perceiver: Audio-Textual Multi-Modal Learning for Estimating Mental Disorders
abstract
Mental disorders, such as anxiety and depression, have become a global concern that affects people of all ages. Early detection and treatment are crucial to mitigate the negative effects these disorders can have on daily life. Although AI-based detection methods show promise, progress is hindered by the lack of publicly available large-scale datasets. To address this, we introduce the Multi-Modal Psychological assessment corpus (MMPsy), a large-scale dataset containing audio recordings and transcripts from Mandarin-speaking adolescents undergoing automated anxiety/depression assessment interviews. MMPsy also includes self-reported anxiety/depression evaluations using standardized psychological questionnaires. Leveraging this dataset, we propose Mental-Perceiver, a deep learning model for estimating mental disorders from audio and textual data. Extensive experiments on MMPsy and the DAIC-WOZ dataset demonstrate the effectiveness of Mental-Perceiver in anxiety and depression detection.
Jinghui Qin, Changsong Liu, Tianchi Tang, Dahuang Liu, Qianying Huang, Rumin Zhang
AAAI7
2025 MCoreOPU: An FPGA-based Multi-Core Overlay Processor for Transformer-based Models
abstract
Transformer-based models have achieved extensive success with increasingly large numbers of parameters and computations, for which many multi-core accelerators have been developed. Nevertheless, they suffer from limited throughput due to either low operating frequency or high communication overhead between cores. This article proposes an FPGA-based multi-core overlay processor, named MCoreOPU, to optimize intra-core computation and inter-core communication. First, we boost the operating frequency of the processing element (PE) array to double the rest of the processor to improve the intra-core throughput. Second, we develop on-chip synchronization routers to reduce off-chip memory traffic, where only the partial sum and maximum are communicated between cores rather than entire vectors for layer normalization and softmax. Moreover, we pipeline synchronization to reduce synchronization latency and develop a bypass of the interconnect bus to reduce the off-chip memory access latency. Finally, we optimize the multi-core model allocation and scheduling to minimize the inter-core communications and maximize the intra-core computation efficiency. The MCoreOPU is implemented in 8-bit fixed-point precision with four cores and four DDRs on the Xilinx U200 FPGA, where the PE array runs at 600 MHz while the rest runs at 300 MHz. Experimental results show that the throughput per MAC of MCoreOPU for BERT, ViT, GPT-2, and LLaMA inference is 1.31 \(\times\) –7.18 \(\times\) higher than other FPGA-based accelerators. Compared with the A100 GPU, the throughput per equivalent MAC efficiency is improved by 22.52 \(\times\) –27.12 \(\times\) .
Shaoqiang Lu, Tiandong Zhao, Ting-Jung Lin, Rumin Zhang, Lei He 0001
ACM Trans. Reconfigurable Technol. Syst.4
2024 CAT: Continual Adapter Tuning for aspect sentiment classification
Qiangpu Chen, Jiahua Huang, Wushao Wen, Qingling Li, Rumin Zhang, Jinghui Qin
Neurocomputing5
2023 RankSearch: An Automatic Rank Search Towards Optimal Tensor Compression for Video LSTM Networks on Edge
abstract
Various industrial and domestic applications call for optimized lightweight video LSTM network models on edge. The recent tensor-train method can transform space-time features into tensors, which can be further decomposed into low-rank network models for lightweight video analysis on edge. The rank selection of tensor is however manually performed with no optimization. This paper formulates a rank search algorithm to automatically decide tensor ranks with consideration of the trade-off between network accuracy and complexity. A fast rank search method, called RankSearch, is developed to find optimized low-rank video LSTM network models on edge. Results from experiments show that RankSearch achieves a$4.84 >$reduction in model complexity, and$1.96\times$speed-up in run time while delivering a 3.86% accuracy improvement compared with the manual-ranked models.
Changhai Man, Chenchen Ding, Shaobo Luo, Rumin Zhang, Ngai Wong 0001, Hao Yu 0001
DATE9
2022 Visual aggregation of large multivariate networks with attribute-enhanced representation learning
Yuhua Liu, Miaoxin Hu, Rumin Zhang, Ting Xu 0002, Yigang Wang, Zhiguang Zhou
Neurocomputing3
2022 Lightweight single image super-resolution with attentive residual refinement network
Jinghui Qin, Rumin Zhang
Neurocomputing2
2022 A Fall Detection Network by 2D/3D Spatio-temporal Joint Models with Tensor Compression on Edge
abstract
Falling is ranked highly among the threats in elderly healthcare, which promotes the development of automatic fall detection systems with extensive concern. With the fast development of the Internet of Things (IoT) and Artificial Intelligence (AI), camera vision-based solutions have drawn much attention for single-frame prediction and video understanding on fall detection in the elderly by using Convolutional Neural Network (CNN) and 3D-CNN, respectively. However, these methods hardly supervise the intermediate features with good accurate and efficient performance on edge devices, which makes the system difficult to be applied in practice. This work introduces a fast and lightweight video fall detection network based on a spatio-temporal joint-point model to overcome these hurdles. Instead of detecting fall motion by the traditional CNNs, we propose a Long Short-Term Memory (LSTM) model based on time-series joint-point features extracted from a pose extractor . We also introduce the increasingly mature RGB-D camera and propose 3D pose estimation network to further improve the accuracy of the system. We propose to apply tensor train decomposition on the model to reduce storage and computational consumption so the deployment on edge devices can to realized. Experiments are conducted to verify the proposed framework. For fall detection task, the proposed video fall detection framework achieves a high sensitivity of 98.46% on Multiple Cameras Fall, 100% on UR Fall, and 98.01% on NTU RGB-D 120. For pose estimation task, our 2D model attains 73.3 mAP in the COCO keypoint challenge, which outperforms the OpenPose by 8%. Our 3D model attains 78.6% mAP on NTU RGB-D dataset with 3.6× faster speed than OpenPose.
Shuwei Li, Changhai Man, Wei Mao 0002, Shaobo Luo, Rumin Zhang, Hao Yu 0001
ACM Trans. Embed. Comput. Syst.7
2020 Semantically-Aligned Universal Tree-Structured Solver for Math Word Problems
abstract
A practical automatic textual math word problems (MWPs) solver should be able to solve various textual MWPs while most existing works only focused on one-unknown linear MWPs.Herein, we propose a simple but efficient method called Universal Expression Tree (UET) to make the first attempt to represent the equations of various MWPs uniformly.Then a semantically-aligned universal tree-structured solver (SAU-Solver) based on an encoder-decoder framework is proposed to resolve multiple types of MWPs in a unified model, benefiting from our UET representation.Our SAU-Solver generates a universal expression tree explicitly by deciding which symbol to generate according to the generated symbols' semantic meanings like human solving MWPs.Besides, our SAU-Solver also includes a novel subtree-level semanticallyaligned regularization to further enforce the semantic constraints and rationality of the generated expression tree by aligning with the contextual information.Finally, to validate the universality of our solver and extend the research boundary of MWPs, we introduce a new challenging Hybrid Math Word Problems dataset (HMWP), consisting of three types of MWPs.Experimental results on several MWPs datasets show that our model can solve universal types of MWPs and outperforms several state-of-the-art models 1 .
Jinghui Qin, Lihui Lin, Xiaodan Liang, Rumin Zhang, Liang Lin 0004
EMNLP (1)4
2019 Adaptive illumination normalization via adaptive illumination preprocessing and modified weber-face
Rumin Zhang, Wenyi Wang 0005
Appl. Intell.3
2018 High Efficient VR Video Coding Based on Auto Projection Selection Using Transferable Features
abstract
Given multiple texture projection methods from the sphere surface to the planar surface, this paper proposes an adaptive selection mode that automatically chooses the appropriate projection method to obtain high compression efficiency of the VR video. The video compression efficiency is inherently affected by the video content, which is closely related to the projection method in the case of VR video encoding. In order to represent the VR video content in a compact manner, a feature vector (transferable feature) for each frame is extracted by a Res-CNN which is pre-trained by a large scale data set for general classification. Afterwards, the relation between the feature and the optimal projection method is investigated by using PCA-KNN, which can project the initial feature vector to a subspace where the VR videos can be efficiently classified with low ambiguity. The experimental results show that the proposed method can select the appropriate projection method that generates the best BD rate.
Lili Zhao 0001, Wenyi Wang 0005, Rumin Zhang, Liaoyuan Zeng
VCIP4