Zhiguo Hu

dblp:117/7164 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorComputer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Unsupervised Multi-View Outlier Detection via Optimal Graph Filtering
abstract
Unsupervised multi-view outlier detection has garnered increasing attention in recent years, yet existing methods face persistent challenges. Many approaches rely predominantly on first-order neighborhood information, overlooking the richer insights offered by higher-order structures, which can degrade detection accuracy. Additionally, some methods suffer from outlier domination in their objective functions, leading to suboptimal performance. Integrating information effectively across multiple views also remains a significant hurdle. To address these challenges, we propose a novel Multi-View Outlier Detection method based on Optimal Graph Filtering (MODGF). Our approach detects outliers using a high-order graph filtering mechanism, ensuring consistency between feature and neighborhood spaces by sharing filtering parameters. Furthermore, we incorporate the Corr-entropy Induced Metric (CIM) to refine the objective function and introduce an efficient scoring strategy for enhanced detection reliability. Extensive experimental results demonstrate that our method is both stable and efficient across scenarios. The code is available at https://github.com/criticcc/MODGF.
Zhiguo Hu, Peng Zhou 0006, Liang Du 0003
ICASSP1
2024 NetDS: Distributed Search Framework with Hybrid Acceleration Methods
abstract
Approximate neighbour nearest search has achieved great success for indexing similar high-dimensional data in distributed search systems. As the scale of data vectors grows, distributed search require large storage, low latency, and high throughput on processing vectors. To achieve this, researchers tend to load balance data with more machines and implement efficient distributed frameworks, but they need to pay huge storage overhead, which leads to inefficient network transmission.To address this gap, we propose NetDS, which exploits the computational capacity of in-network computation and the storage capacity of solid-state drives. NetDS utilizes a multi-level constrained balanced tree to process data vectors and construct multi-level tables. Then, NetDS proposes a heuristic neighbour graph to solve the boundary data problem. NetDS also offloads central tables and graphs into switches to accelerate vector classification. Finally, NetDS designs hybrid storage and query pre-match methods to accelerate the ANNS distributed system. We deploy NetDS on a programmable switch and evaluate it. NetDS completes the data and query processing in a shorter time than other typical distributed frameworks.
Penghao Zhang, Zhiguo Hu
ISPA2
2024 RS-DETR: An Improved DETR for High-Resolution Remote Sensing Image Object Detection
abstract
High-resolution remote sensing image object detection is an important research area in remote sensing information processing and has substantial practical applications. This domain presents unique challenges, including variable object scales, complex backgrounds, prevalent small objects, and densely arranged items, distinguishing it from traditional object detection in natural images. This paper proposes a novel object detection algorithm(RS-DETR), which builds upon the DETR framework and integrates the Swin Transformer. The algorithm features a dual-branch structure in its feature extraction module, markedly improving detection accuracy, especially for objects of varying scales. The addition of the GAM convolutional attention mechanism allows the model to concentrate more effectively on relevant regions, minimizing background complexities. Moreover, we have included the scale-invariant intersection over union (SIoU) loss function to enhance the precise localization of closely packed objects. To demonstrate the efficacy of the algorithm, RS-DETR was applied to the HRSC2016 and NWPU VHR-10 datasets. The results show average detection accuracies of 86.1% and 57.9% on these datasets, respectively, outperforming the baseline models by 1.1% and 0.9%, respectively.
Zhiguo Hu
SMC4
2024 A Lightweight_PAEKS-based energy scheduling model considering priority in MicroGrid
Xialei Zhang, Yaoyang Wang, Tianjun Ma, Lifeng Guo, Zhiguo Hu
Ad Hoc Networks5
2024 Understanding Flow Experience in Video Learning by Multimodal Data
abstract
Video-based learning has successfully become an effective alternative to face-to-face instruction. In such situations, modeling or predicting learners’ flow experience during video learning is critical for enhancing the learning experience and advancing learning technologies. In this study, we set up an instructional scenario for video learning according to flow theory. Different learning states, i.e., boredom, fit (flow), and anxiety, were successfully induced by varying the difficulty levels of the learning task. We collected learners’ electrocardiogram (ECG) signals as well as facial video, upper body posture and speech data during the learning process. We proposed classification models of the learning state and regression models to predict flow experience by utilizing different combinations of the data from the four modalities. The results showed that the model performance of learning state recognition was significantly improved by the decision-level fusion of multimodal data. By using the selected important features from all data sources, such as the standard deviation of normal to normal R-R intervals (SDNN), high-frequency (HF) heart rate variability and mel-frequency cepstral coefficients (MFCC), the multilayer perceptron (MLP) classifier gave the best recognition result of learning states (i.e., mean AUC of 0.780). The recognition accuracy of boredom, fit (flow) and anxiety reached 47.48%, 80.89% and 47.41%, respectively. For flow experience prediction, the MLP regressor based on the fusion of two modalities (i.e., ECG and posture) achieved the optimal prediction (i.e., mean RMSE of 0.717). This study demonstrates the feasibility of modeling and predicting the flow experience in video learning by combining multimodal data.
Zhiguo Hu
Int. J. Hum. Comput. Interact.4
2023 Is This Science Video Popular? Let Us See How the Audience Reacts!
abstract
The popularity of science videos is critical for the dissemination of knowledge, and predicting the popularity of a video is a hot topic among researchers. The existing research is mainly based on the videos’ content (e.g., theme) or video-related external information (e.g., comments). However, videos with different popularity can bring different learning states and viewing experiences (emotional awakening, flow experiences). In this paper, we utilize participants’ learning states when they watch videos to predict the popularity of the videos, depending on two modal indicators, i.e., contact physiological indicators recorded by a Biopac MP150 polygraph and noncontact gesture indices recorded by a Kinect V2 body camera which can obtain data of head position. We propose two classification prediction models with each modality of indicator and filter out the indicators that make sense for modeling. Results show that the meaningful indicators in the physiological modality through stepwise logistic regression are the standard deviation of normal to normal R-R intervals (SDNN) and high-frequency heart rate variability (HF). We find that participants had higher SDNN and lower HF when watching science video with high popularity (compared with low popularity one), and the accuracy rate of classification model is 81.6%. In a same way, the selected meaningful indicators in the gesture aspect are the maximum, the standard deviation (SD) of distance between the participants’ head and Kinect. We find that the maximum and SD of head distance are smaller when participants studied highly popular science videos (contrary to the less popular videos) and the accuracy rate is 73.7%. Combining the four indicators for modeling by the direct input method, the accuracy rate is 78.9%, and the SD of head distance is probably the most important indicator for predicting popularity of videos. These results indicate that it is feasible to predict video popularity by learning states.
Guangliang Hu, Zhen Yang 0033, Zhiguo Hu
Int. J. Hum. Comput. Interact.4
2020 Evaluating QoE in VoIP networks with QoS mapping and machine learning algorithms
Zhiguo Hu, Hongren Yan, Haijun Geng, Guoqing Liu 0001
Neurocomputing1
2020 Multiscale fusion and aggregation PCNN for 3D shape recovery
Zhiguo Hu, Fengxian Liu
Inf. Sci.4
2020 3D shape reconstruction from multifocus image fusion using a multidirectional modified Laplacian operator
Zhiguo Hu, Zhiwei Qiao, Linyuan Zhang
Pattern Recognit.2
2018 A new approach for packet loss measurement of video streaming and its application
Zhiguo Hu, Qiqiang Zhang
Multim. Tools Appl.1
2012 SLDRT: A measurement technique for available bandwidth on multi-hop path with bursty cross traffic
Zhiguo Hu, Dalu Zhang, Hualei Zhou
Comput. Networks1