Yong Gu

dblp:32/7152 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 8 · 5 first-authorTheory of computation · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Lightweight Multi-Variable Spatio-Temporal Convolutional Framework for Dynamic Gesture Recognition
abstract
Transformer-based hybrid architectures have achieved remarkable performance in dynamic hand gesture recognition. However, their high computational overhead and model size limit deployment in resource-limited environments. Motivated by this limitation, we propose the Decoupled Spatio-Temporal Convolutional Network (DSTCNet), a lightweight, pure convolutional framework trained end-toend delivering high accuracy with a fraction of the complexity. DSTCNet integrates two components: (1) an efficient pseudo-3D spatial backbone, the Pseudo-3D Gated Attentional Fusion Network (P3D-GAFNet), enhancing spatial feature extraction via positional prior injection, and (2) a temporal modeling network, the Multi-Variable Decomposition Temporal Convolutional Network (MVD-TCN), leveraging multi-variable feature decomposition with modern convolutional blocks to capture long-range temporal dependencies without the cost of self-attention. With only 9.6M parameters, DSTCNet matches or surpasses the accuracy of substantially larger models on several challenging benchmarks, while offering high computational efficiency, lower memory usage, and reduced energy consumption—making it a practical solution for deployment on edge devices. Our results demonstrate that modernized pure convolutional architectures can serve as a robust and efficient alternative to hybrid designs, offering valuable insights for the broader field of video understanding.
Guoqiong Liao, Longjie Huang, Yong Gu
3DV4
2026 MDRA: A Motion-guided Dual-stream Recurrent Attention Framework for Dynamic Hand Gesture Recognition
abstract
Dynamic Hand Gesture Recognition (DHGR) aims to detect dynamic hand movements by leveraging the features and continuity of video frames. Existing methods mainly utilize backbone networks to extract latent features from individual video frames and sequence modeling through a Transformer. However, the hand usually occupies a relatively small proportion in the video, resulting in a large amount of invalid information in the extracted features, which affects the model’s robustness and subsequent temporal modeling performance. Moreover, the traditional Transformer structure has a high time complexity, which affects the model’s operational efficiency. To address these issues, we propose a novel data preprocessing and data fusion approach. It filters the hand contour using the motion vector of video coding and extracts features from RGB images and contour images through a dual-stream network. Additionally, a Gated-MLP GCN (GM-GCN) fusion module is proposed to fully fuse the dual-stream features. Meanwhile, we developed an Efficient Multi-scale Recurrent Attention (EMRA) module as our temporal modeling network, which adopts a recurrent structure similar to an RNN for attention calculation, enabling efficient parallel training of the model. Moreover, it better captures the details and dynamic changes of gestures through wavelet transform and multi-scale pooling strategies. Extensive experiments demonstrate that our proposed framework achieves highly competitive results on key benchmarks (e.g., 83.87% accuracy on NVGesture) while ensuring computational efficiency, reducing MACs by 28% compared to the standard Transformer.
Guoqiong Liao, Longjie Huang, Yong Gu, Tao Zhu 0006
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Time-channel Adaptive Fusion and Hierarchical Attention Mechanism for Dynamic Hand Gesture Recognition
Longjie Huang, Jianhai Liu, Yong Gu
ICMI3
2025 Diffusion in Time and Frequency Domains for Efficient 3D Human Pose Estimation in Videos
Yong Gu, Longjie Huang
PRCV (10)1
2024 Generalizing Soft Actor-Critic Algorithms to Discrete Action Spaces
Le Zhang 0015, Yong Gu, Yanshuo Zhang, Yifei Jin, Xinxin Wu
PRCV (1)2
2021 Constructing a Distance Sensitivity Oracle in O(n^2.5794 M) Time
abstract
We continue the study of distance sensitivity oracles (DSOs). Given a directed graph $G$ with $n$ vertices and edge weights in $\{1, 2, \dots, M\}$, we want to build a data structure such that given any source vertex $u$, any target vertex $v$, and any failure $f$ (which is either a vertex or an edge), it outputs the length of the shortest path from $u$ to $v$ not going through $f$. Our main result is a DSO with preprocessing time $O(n^{2.5794}M)$ and constant query time. Previously, the best preprocessing time of DSOs for directed graphs is $O(n^{2.7233}M)$, and even in the easier case of undirected graphs, the best preprocessing time is $O(n^{2.6865}M)$ [Ren, ESA 2020]. One drawback of our DSOs, though, is that it only supports distance queries but not path queries. Our main technical ingredient is an algorithm that computes the inverse of a degree-$d$ polynomial matrix (i.e. a matrix whose entries are degree-$d$ univariate polynomials) modulo $x^r$. The algorithm is adapted from [Zhou, Labahn, and Storjohann, Journal of Complexity, 2015], and we replace some of its intermediate steps with faster rectangular matrix multiplication algorithms. We also show how to compute unique shortest paths in a directed graph with edge weights in $\{1, 2, \dots, M\}$, in $O(n^{2.5286}M)$ time. This algorithm is crucial in the preprocessing algorithm of our DSO. Our solution improves the $O(n^{2.6865}M)$ time bound in [Ren, ESA 2020], and matches the current best time bound for computing all-pairs shortest paths.
Yong Gu, Hanlin Ren
ICALP1
2021 Approximate Distance Oracles Subject to Multiple Vertex Failures
abstract
Given an undirected graph G = (V, E) of n vertices and m edges with weights in [1, W], we construct vertex sensitive distance oracles (VSDO), which are data structures that preprocess the graph, and answer the following kind of queries: Given a source vertex u, a target vertex v, and a batch of d failed vertices D, output (an approximation of) the distance between u and v in G – D (that is, the graph G with vertices in D removed). An oracle has stretch α if it always holds that , where δG–D(u, v) is the actual distance between u and v in G – D, and is the distance reported by the oracle. In this paper we construct efficient VSDOs for any number d of failures. For any constant c ≥ 1, we propose two oracles: The first oracle has size n2+1/c(log n/∊)O(d) · log W, answers a query in poly(log n, dc, log log W, ∊–1) time, and has stretch 1 + ∊, for any constant ∊ > 0. The second oracle has size n2+1/cpoly (log(nW), d), answers a query in poly (log n, dc, log log W) time, and has stretch poly (log n, d). Both of these oracles can be preprocessed in time polynomial in their space complexity. These results are the first approximate distance oracles of poly-logarithmic query time for any constant number of vertex failures in general undirected graphs. Previously there are (1 + ∊)-approximate d-edge sensitive distance oracles [Chechik et al. 2017] answering distance queries when d edges fail, which have size O(n2(log n/∊)d · d log W) and query time poly (log n, d, log log W).
Yong Gu, Hanlin Ren
SODA2
2020 Roundtrip Spanners with (2k-1) Stretch
abstract
A roundtrip spanner of a directed graph G is a subgraph of G preserving roundtrip distances approximately for all pairs of vertices. Despite extensive research, there is still a small stretch gap between roundtrip spanners in directed graphs and undirected graphs. For a directed graph with real edge weights in [1,W], we first propose a new deterministic algorithm that constructs a roundtrip spanner with (2k-1) stretch and O(k n^(1+1/k) log (nW)) edges for every integer k > 1, then remove the dependence of size on W to give a roundtrip spanner with (2k-1) stretch and O(k n^(1+1/k) log n) edges. While keeping the edge size small, our result improves the previous 2k+ε stretch roundtrip spanners in directed graphs [Roditty, Thorup, Zwick'02; Zhu, Lam'18], and almost matches the undirected (2k-1)-spanner with O(n^(1+1/k)) edges [Althöfer et al. '93] when k is a constant, which is optimal under Erdös conjecture.
Ruoxu Cen, Yong Gu
ICALP3
2020 Deep Reinforcement Learning for Solving AGVs Routing Problem
Chengxuan Lu, Jinjun Long, Zichao Xing, Weimin Wu 0001, Yong Gu, Jiliang Luo, Yisheng Huang
VECoS5
2019 Multi-Pair Active Shielding for Security IC Protection
abstract
Active shielding techniques are effective methods to prevent attacks of probing or FIB. Many research groups proposed methods for automatically metallic shield generation but few considered the multi-pair and shield aligning requirements. In this paper, the active shielding problem is redefined considering both requirements and a novel two-phase method is proposed to solve it. A coverage-oriented initial routing algorithm is performed to generate initial random paths for each routing-pair and an ILP-based tile routing algorithm is presented to generate complex shield lines. Experimental results show that the proposed method is very effective and the generated shield can satisfy all design principles of connectivity, full-coverage, randomness, and complexity. When compared to traditional methods, it can solve all test cases while improving routability by 67.7% and reducing solving time by 99.1%.
Kan Wang 0007, Yong Gu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 Improved Time Bounds for All Pairs Non-decreasing Paths in General Digraphs
abstract
We present improved algorithms for solving the All Pairs Non-decreasing Paths (APNP) problem on weighted digraphs. Currently, the best upper bound on APNP is O~(n^{(9+omega)/4})=O(n^{2.844}), obtained by Vassilevska Williams [TALG 2010 and SODA'08], where omega<2.373 is the usual exponent of matrix multiplication. Our first algorithm improves the time bound to O~(n^{2+omega/3})=O(n^{2.791}). The algorithm determines, for every pair of vertices s, t, the minimum last edge weight on a non-decreasing path from s to t, where a non-decreasing path is a path on which the edge weights form a non-decreasing sequence. The algorithm proposed uses the combinatorial properties of non-decreasing paths. Also a slightly improved algorithm with running time O(n^{2.78}) is presented.
Yong Gu, Le Zhang 0015
ICALP2
2016 Asymmetric multiprocessing for motion control based on Zynq SoC
abstract
The Zynq™-7000 family is Xilinx's first extensible processing platform (EPP). This product combines an ARM® dual-core Cortex™-A9 MPCore™ processing system with a Xilinx 28nm field programmable gate array (FPGA) which offers the flexibility and scalability. Nowadays, real-time control systems and easy-to-use operating systems are both required for many industry applications. Combining the high-performance Bare-Metal system and flexible Linux operating system, the Asymmetric Multiprocessing (AMP) based on Zynq SoC is a powerful solution to meet various demands in motion control. This paper proposes a method to implement the AMP on Zynq SoC which runs Linux and Bare-Metal system on two separate cores, and describes two motion control demonstrations.
Yong Gu, Xuguang Guan
FPT2
2001 A text-independent speaker verification system using support vector machines classifier
abstract
Abstract In the recent years the technology for speaker verification orcall authentication has received an increasing amount ofattention in IVR industry. However due to the complexity ofspeaker information embedded in the speech signals thecurrent technology still can not produce the verificationaccuracy to meet the requirement for some applications. In thispaper we introduce a new pattern classification approach,support vector machines (SVM) for the text-independentspeaker verification. The SVM is a new way of statisticallearning based on a principle of structural risk minimisation.In the paper various evaluation results for the SVMverification system are presented and a comparison with abaseline GMM approach is also given. The results demonstratethat the SVM approach perform much better than the GMMapproach. On the same training and testing data set the SVMapproach gives an EER 1.2% versus 3.9% EER from theGMM approach. 1. Introduction In the recent years the technology for speaker verification(SV) or call authentication has received an increasing amountof attention in IVR industry. However due to the complexityof speaker information embedded in the speech signals thecurrent technology such as HMM, GMM, ANN etc. still cannot produce the verification accuracy to meet the requirementfor some applications. In this paper we introduce a newapproach support vector machines (SVM) for this problem.The SVM is a learning technique introduced by V. Vapnik [1].It can be seen as a new way to statistical learning based on aprinciple of structural risk minimisation. An explicit noisedescription in the approach and the possibility of using non-linear kernel in the dual representation makes this method veryattractive in many pattern recognition areas. The technique hasbeen applied in the area of computer vision and others.Recently some works have shown that the algorithm canachieve better phoneme classification accuracy than someconventional methods for speech processing [2][3]. This paperpresents a text-independent SV system using the SVMapproach. In the paper some alternatives in the kernelfunctions and decision functions are discussed and evaluationresults are presented. The Gaussian mixture model (GMM)technique, one of most popular approach for the text-independent SV, is used as a baseline in our evaluations. Acomparison between the SVM and the GMM approach isgiven in the paper. Results demonstrate that the SVMalgorithm perform much better than the baseline GMM. Onthe same training and testing data set the SVM approach givesan equal error rate (EER) 1.2% versus 3.9% from the GMM.
Yong Gu, Trevor Thomas
INTERSPEECH1
2001 Word level confidence measures using n-best sub-hypotheses likelihood ratio
abstract
This paper proposes an efficient confidence measure applied at the word level by combining various likelihood ratio tests. The estimates are derived from the local N-best subhypotheses. This approach allows the confidence measures to take into account the effect of neighboring words and still provides the estimate localized around the word to be verified. It produces an effective confidence measure that is usable for various tasks. We compared the results with other likelihood ratio based confidence measures including garbage model, N-best homogeneity and online garbage models. The proposed method gave more than 30% relative false accept rate reduction over other methods and the rejection performance was less task-dependent.
Beng Tiong Tan, Yong Gu, Trevor Thomas
INTERSPEECH2
2000 Speaker verification in operational environments - monitoring for improved service operation
abstract
There are very few, if any, published accounts of practical expe rience with Speaker Verification as a means to provide secure access to telematics services.Yet, there is no reason to expect that Speaker Verification is very different from speech recogni tion, for which many deployed services have shown the need for close and intensive on-line monitoring during the time when the service becomes operational.In this paper we present our expe rience with a monitoring scheme for Speaker Verification dur ing the field test of a financial investment game.Many of the issues that were monitored were suggested by our experience with a semi-operational service, viz.free access to Directory Assistance for visually impaired.A newly developed enrolment procedure, that can flag potentially weak speaker models, is an essential part of the monitoring procedure.
Yong Gu, Hans Jongebloed, Dorota J. Iskra, Els den Os, Lou Boves
INTERSPEECH1
2000 Advances on HMM-based text-dependent speaker verification
Yong Gu, Trevor Thomas
INTERSPEECH1
2000 Competition-based score analysis for utterance verification in name recognition
Yong Gu, Trevor Thomas
INTERSPEECH1
2000 Utterance verification based speech recognition system
Beng Tiong Tan, Yong Gu, Trevor Thomas
INTERSPEECH2
1999 A hybrid score measurement for HMM-based speaker verification
abstract
In speaker verification the world model based approach and the cohort model based approach have been used for better HMM score measurements for verification comparison. From theoretical analysis these two approaches represent two different paradigms for verification decision-making strategy. The two techniques could be combined for a better solution. In the paper we present a hybrid score measurement which combines the world model based technique and the cohort model based technique together. The method is evaluated with the YOHO database. The results show that the combination can lead a better score measurement which improves speaker verification performance. An experimental comparison between the world model based approach and the cohort model based approach with the YOHO database can also be found in the paper.
Yong Gu, Trevor Thomas
ICASSP1
1998 An implementation and evaluation of an on-line speaker verification system for field trials
Yong Gu, Trevor Thomas
ICSLP1
1998 Evaluation and implementation of a voice-activated dialing system with utterance verification
Beng Tiong Tan, Yong Gu, Trevor Thomas
ICSLP2