EDBT 2026 Demo / reviewers in the wild / expert
Yue Ming 0001
dblp:73/6310-1
· DBLP profile ↗
44ranked-venue papers
17as first author
24since 2021 · last 2026
0000-0001-7105-4207ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 8 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlowGS: End-to-end correspondence-guided 3D Gaussian Splatting from sparse unposed images
Kai Hong, Yue Ming 0001, Chuanchen Luo, Jiangwan Zhou |
Neurocomputing | 2 |
| 2025 | DSSLNet: dual-stream self-supervised learning network for end-to-end speech recognition
Boyang Lyu, Chunxiao Fan 0001, Yue Ming 0001, Nannan Hu |
Appl. Intell. | 3 |
| 2025 | CG-MCFNet: cross-layer guidance-based multi-scale correlation fusion network for 3D face recognition
Panzi Zhao, Yue Ming 0001, Hui Yu 0001, Jiangwan Zhou |
Appl. Intell. | 2 |
| 2024 | LMGSNet: A Lightweight Multi-scale Group Shift Fusion Network for Low-quality 3D Face RecognitionabstractWith the easy availability of low-quality 3D facial data, research on low-quality 3D face recognition (FR) has gained widespread attention. However, most existing methods struggle to strike a balance between accuracy and computational complexity, with the enormous parameters being one of the primary challenges. To address this issue, we propose a novel lightweight multi-scale group shift fusion network (LMGSNet) for low-quality 3D FR. Specifically, we construct a multi-scale group attention layer-by-layer shift fusion module (GALSF) based on our proposed channel shift fusion (CSF) method, which integrates attention convolution and grouping operation to capture critical local features (such as the nose, mouth, and forehead) while significantly reducing parameters. Furthermore, we design a novel split-aggregate local feature fusion module (SALF) to enhance local features representation and capturing rich discriminative features. Extensive experiments on three challenging low-quality 3D face datasets demonstrate that our model achieves competitive recognition accuracy with the lowest parameters. Yue Ming 0001, Panzi Zhao, Boyang Lyu, Kai Hong |
ICME | 2 |
| 2024 | F2D-SIFPNet: a frequency 2D Slow-I-Fast-P network for faster compressed video action recognition
Yue Ming 0001, Jiangwan Zhou, Xia Jia, Nannan Hu |
Appl. Intell. | 1 |
| 2024 | Action recognition in compressed domains: A survey
Yue Ming 0001, Jiangwan Zhou, Nannan Hu, Panzi Zhao, Boyang Lyu, Hui Yu 0001 |
Neurocomputing | 1 |
| 2024 | CSS-Net: A Consistent Segment Selection Network for Audio-Visual Event LocalizationabstractAudio-visual event (AVE) localization aims to localize the temporal boundaries of events that contains visual and audio contents, to identify event categories in unconstrained videos. Existing work usually utilizes successive video segments for temporal modeling. However, ambient sounds or irrelevant visual targets in some segments often cause the problem of audio-visual semantics inconsistency, resulting in inaccurate global event modeling. To tackle this issue, we present a consistent segment selection network (CSS-Net) in this paper. First, we propose a novel bidirectional guided co-attention (BGCA) block, containing two distinct attention paths from audio to vision and from vision to audio, to focus on sound-related visual regions and event-related sound segments. Then, we propose a novel context-aware similarity measure (CASM) module to select semantic consistent visual and audio segments. A cross-correlation matrix is constructed using the correlation coefficients between the visual and audio feature pairs in all time steps. By extracting highly correlated segments and discarding low correlated segments, visual and audio features can learn global event semantics in videos. Finally, we propose a novel audio-visual contrastive loss to learn the similar semantics representation for visual and audio global features under the constraints of cosine and L2 similarities. Extensive experiments on public AVE dataset demonstrates the effectiveness of our proposed CSS-Net. The localization accuracies achieve the best performance of 80.5% and 76.8% in both fully- and weakly-supervised settings compared with other state-of-the-art methods. Yue Ming 0001, Nannan Hu, Hui Yu 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Frequency Enhancement Network for Efficient Compressed Video Action RecognitionabstractThe existing frequency-based action recognition methods achieve impressive performance in improving efficiency. However, they ignore the low-frequency texture and edge clues, leading to accuracy degradation. To address this problem, we propose a novel frequency enhancement (FE) block for efficient compressed video action recognition, including a temporal-channel two-heads attention (TCTHA) module and a frequency overlapping group convolution (FOGC) module. First, the TCTHA module emphasizes the inter-frame temporal context and the inner-frame informative frequency semantics by attention. Then, the FOGC module groups channels in different frequency bands with overlap, to extract low-frequency texture and edge clues, while maintaining the interaction of groups. We integrate the FE block into 2D-CNNs with frequency I-frame input, termed FENet, focusing on the pivotal low-frequency spatio-temporal semantics for action recognition. Experiments on HMDB-51, UCF-101, Kinetics-400, and Kinetics-700 verify that our FENet achieves comparable accuracy compared with RGB-based methods with high efficiency. Yue Ming 0001, Xia Jia, Jiangwan Zhou, Nannan Hu |
ICIP | 1 |
| 2023 | See, move and hear: a local-to-global multi-modal interaction network for video action recognition
Yue Ming 0001, Nannan Hu, Jiangwan Zhou |
Appl. Intell. | 2 |
| 2023 | MAENet: A novel multi-head association attention enhancement network for completing intra-modal interaction in image captioning
Nannan Hu, Chunxiao Fan 0001, Yue Ming 0001 |
Neurocomputing | 3 |
| 2023 | CAST: Context-association architecture with simulated long-utterance training for mandarin speech recognition
Yue Ming 0001, Boyang Lyu |
Speech Commun. | 1 |
| 2023 | En-HACN: Enhancing Hybrid Architecture With Fast Attention and Capsule Network for End-to-end Speech RecognitionabstractAutomatic speech recognition (ASR) is a fundamental technology in the field of artificial intelligence. End-to-end (E2E) ASR is favored for its state-of-the-art performance. However, E2E speech recognition still faces speech spatial information loss and text local information loss, which results in the increase of deletion and substitution errors during inference. To overcome this challenge, we propose a novel Enhancing Hybrid Architecture with Fast Attention and Capsule Network (termed En-HACN), which can model the position relationships between different acoustic unit features to improve the discriminability of speech features while providing the text local information during inference. Firstly, a new CNN-Capsule Network (CNN-Caps) module is proposed to capture the spatial information in the spectrogram through the capsule output and dynamic routing mechanism. Then, we design a novel hybrid structure of LocalGRU Augmented Decoder (LA-decoder) that generates text hidden representations to obtain text local information of the target sequences. Finally, we introduce fast attention instead of self-attention in En-HACN, which improves the generalization ability and efficiency of the model in long utterances. Experiments on corpora Aishell-1 and Librispeech demonstrate that our En-HACN has achieved the state-of-the-art compared with existing works. Besides, experiments on the long utterances dataset based on Aishell-1-long show that our model has a high generalization ability and efficiency. Boyang Lyu, Chunxiao Fan 0001, Yue Ming 0001, Panzi Zhao, Nannan Hu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | LMFNet: A Lightweight Multiscale Fusion Network With Hierarchical Structure for Low-Quality 3-D Face RecognitionabstractThree-dimensional (3-D) face recognition (FR) can improve the usability and user-friendliness of human–machine interaction. In general, 3-D FR can be divided into high-quality and low-quality 3-D FR according to different interaction scenarios. The low-quality data can be easily obtained, so its application prospect is more extensive. However, the challenge is how to balance the trade-offs between data accuracy and real-time performance. To solve this problem, we propose a lightweight multiscale fusion network (LMFNet) with a hierarchical structure based on single-mode data for low-quality 3-D FR. First, we design a backbone network with only five feature extraction blocks to reduce computational complexity and improve the inference speed. Second, we devise a mid-low adjacent layer with a multiscale feature fusion (ML-MSFF) module to extract the facial texture and contour information, and a mid-high adjacent layer with a multiscale feature fusion (MH-MSFF) module to obtain the discriminative information in high-level features. Then, a hierarchical multiscale feature fusion (HMSFF) module is formed by combining these two modules mentioned above to acquire the local information of different scales. Finally, we enhance the expression of features by integrating HMSFF with a global convolutional neural network for improving recognition accuracy. Experiments on Lock3DFace, KinectFaceDB, and IIIT-D datasets demonstrate that our proposed LMFNet can achieve superior performance on low-quality datasets. Furthermore, experiments on the cross-quality database based on Bosphorus and the different intensity noise low-quality datasets based on UMB-DB and Bosphorus show that our network is robust and has a high generalization ability. It satisfies the real-time requirement, which lays a foundation for a smooth and user-friendly interactive experience. Panzi Zhao, Yue Ming 0001, Xuyang Meng, Hui Yu 0001 |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2023 | TSFNet: Triple-Steam Image CaptioningabstractImage captioning is a challenging task that generates a natural language description based on the visual understanding of the given image. Significant region representation is a milestone in image captioning. Despite the great success of existing region-based works, they only focus on the salient objects and encode these objects independently, still plagued by the lack of global contextual information and visual relationships. In fact, the global contextual information and structured visual relationships are exactly the merits of traditional grid features and emerging scene graph features. In this paper, we present a Triple-Steam Feature Fusion Network (TSFNet) to leverage the complementary advantages of the grid, region, and scene graph triple-steam visual representations in image captioning. Concretely, in our TSFNet, a novel Dual-level Attention (DA) mechanism is proposed to simultaneously explore visual intrinsic properties and word-related attributes uniformly of different features. Then attention enhanced features of different modalities are mapped into a joint representation to guide the caption generation. Moreover, we design a new global-aware decoder, which leverages the concatenated representation of triple-steam features and the joint attention representation to obtain global visual guidance information, further refine the complex multimodal reasoning. To verify the effectiveness of our feature fusion model, we perform extensive experiments on the highly competitive MSCOCO dataset to evaluate the model quantitatively and qualitatively. The results illustrate that the proposed framework outperforms many state-of-the-art image captioning approaches in various evaluation metrics, and generates more accurate and abundant captions. Nannan Hu, Yue Ming 0001, Chunxiao Fan 0001, Boyang Lyu |
IEEE Trans. Multim. | 2 |
| 2022 | M-CoTransT: Adaptive spatial continuity in visual trackingabstractAbstract Visual tracking is an important area in computer vision. Based on the Siamese network, current tracking methods employ the self‐attention block in convolutional networks to extract semantic features containing the image structure information of an object. However, spatial continuity is a point of contradiction between two seemingly unrelated challenges, that is, occlusion and similar distractor, in tracking methods. At the same time, it is a spatially discontinuous task to locate a target reappearing after occlusion accurately. The prediction of bounding boxes should be constrained by spatial continuity to prevent them from jumping into similar distractors. This study proposes a novel tracking method for introducing spatial continuity in visual tracking called M‐CoTransT; the novel tracking method is developed through the confidence‐based adaptive Markov motion model (M‐model) and a novel correlation‐based feature fusion network (CoTransT). In particular, the M‐model provides confidence for the nodes of the Markov motion model to estimate the motion state continuity. It also predicts a more accurate search region for CoTransT, which then adds a cross‐correlation branch into the self‐attention tracking network to enhance the continuity of target appearance in the feature space. Extensive experiments on five challenging datasets (LaSOT, GOT‐10k, TrackingNet, OTB‐2015 and UAV123) demonstrated the effectiveness of the proposed M‐CoTransT in visual tracking. Chunxiao Fan 0001, Runqing Zhang, Yue Ming 0001 |
IET Comput. Vis. | 3 |
| 2022 | SSLNet: A network for cross-modal sound source localization in visual scenes
Yue Ming 0001, Nannan Hu |
Neurocomputing | 2 |
| 2022 | CORNet: Context-Based Ordinal Regression Network for Monocular Depth EstimationabstractMonocular depth estimation, as one of the fundamental tasks of computer vision, plays a crucial role in three-dimensional (3D) scene understanding and perception. Usually, deep learning methods recover monocular depth maps using continuous regression manners by minimizing the errors between the ground-truth depth and the predicted depth. However, fine depth features may not be fully captured through layer-by-layer coding, which is prone to low spatial resolution depth maps and insufficient details. Furthermore, it usually converges slowly and suffers from unsatisfactory results. To tackle these issues, we propose a novel model, named context-based ordinal regression network (CORNet), to reconstruct monocular depth maps in the ordinal regression manner with context information in this paper. Firstly, we put forward a novel context-based encoder with a feature transformation (FT) module to learn context information and details from inputs, and output multi-scale feature maps. Then, we design a boundary enhancement module (BEM) with a spatial attention mechanism following each operation of feature fusion, which captures boundary features in the scene to enhance the border depth. Finally, a feature optimization module (FOM) is designed to fuse and optimize the multi-scale features and boundary features to strengthen depth learning. What’s more, we introduce an ordinal weighted inference to predict depth maps from probabilities and discretization values. Experiments and results on two challenging datasets, KITTI and NYU Depth V2, demonstrate that our proposed CORNet can estimate monocular depth maps effectively and obtain superior performance in capturing geometric features over existing methods. Xuyang Meng, Chunxiao Fan 0001, Yue Ming 0001, Hui Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | MP-LN: motion state prediction and localization network for visual object tracking
Chunxiao Fan 0001, Runqing Zhang, Yue Ming 0001 |
Vis. Comput. | 3 |
| 2021 | Faster-FCoViAR: Faster Frequency-Domain Compressed Video Action Recognition
Xia Jia, Yue Ming 0001, Jiangwan Zhou, Nannan Hu |
BMVC | 3 |
| 2021 | Re-Identify Deformable Targets for Visual Tracking
Runqing Zhang, Chunxiao Fan 0001, Yue Ming 0001 |
PRCV (1) | 3 |
| 2021 | RFRN: A recurrent feature refinement network for accurate and efficient scene text detection
Guanyu Deng, Yue Ming 0001, Jing-Hao Xue |
Neurocomputing | 2 |
| 2021 | Mutual-learning sequence-level knowledge distillation for automatic speech recognition
Yue Ming 0001, Lei Yang 0034, Jing-Hao Xue |
Neurocomputing | 2 |
| 2021 | 3D-TDC: A 3D temporal dilation convolution framework for video action recognition
Yue Ming 0001, Chao Li 0026, Jing-Hao Xue |
Neurocomputing | 1 |
| 2021 | Deep learning for monocular depth estimation: A review
Yue Ming 0001, Xuyang Meng, Chunxiao Fan 0001, Hui Yu 0001 |
Neurocomputing | 1 |
| 2020 | An Effective Hierarchical Resolution Learning Method for Low-Resolution Targets TrackingabstractSuffering from the low-resolution target's visual quality, the precisions of visual object trackers are reduced. This paper proposes an effective hierarchical resolution learning method for low-resolution targets tracking, abbreviated as HRT. We adopt a hierarchical structure to exploit information from different resolution levels. (1) At the high level: the super-resolution (SR) images, determining the target's shape, contains richer image textures and clearer target contours, and transmits the search region to the low level. (2) At the low level: low-resolution (LR) images maintain the spatial structure information of the original target, providing the precise center coordinates of the target. Experimental results demonstrate the effectiveness of the proposed tracker, which HRT achieves 90.3% precision on OTB100 LR sequences and 78.5% precision on LR sequences from UAV123 datasets, gaining 2.0%, 2.4% improvement over state-of-the-art trackers respectively. Runqing Zhang, Chunxiao Fan 0001, Yue Ming 0001, Hao Fu 0013, Xuyang Meng |
ICIP | 3 |
| 2020 | Variational Bayesian Sparsification for Distillation CompressionabstractModel compression is a critical technique for cumbersome models to reduce memory consumption and accelerate inference. Here, we propose a novel method, called Variational Bayesian Sparsification, for distilling large models into small and sparse models while maintaining accuracy. Different from prior work, our approach innovatively embeds Bayesian sparsification into distillation. The core contributions are in two folds. First, the minibatch re-weighting method is proposed to dynamically balance the hard and soft knowledge, which can largely boost distillation accuracy. Then, the Bayesian deep sparse method presents to leverage the group sparseness and element sparseness simutaneously to reduce the parameter redundancy of student networks. We validate our method in MNIST, CIFAR10 and CIFAR100 datasets. Our method achieves 98.86% compression ratio with minor accuracy loss in MNIST. It is also evaluated through a compact model with only 1.99M weights on CIFAR10 and performed favorably against state-of-the-art compression methods on CFAR100, to verify the algorithm's generality. Yue Ming 0001, Hao Fu 0013 |
ICME | 1 |
| 2020 | MD-ST: Monocular Depth Estimation Based on Spatio-Temporal Correlation Features
Xuyang Meng, Chunxiao Fan 0001, Yue Ming 0001, Runqing Zhang, Panzi Zhao |
PRCV (1) | 3 |
| 2020 | Efficient scalable spatiotemporal visual tracking based on recurrent neural networks
Yue Ming 0001, Yashu Zhang |
Multim. Tools Appl. | 1 |
| 2018 | Sparse Tikhonov-Regularized Hashing for Multi-Modal LearningabstractThis paper mainly focuses on the role of regularization in Multi-Modal Learning (MML). Existing MML studies devote most of the efforts in maximizing the consensus of models from cues of different modalities. However, regularization methods are still far from fully explored. To fill in this gap, we propose a compact and efficient coding solution, termed by sparse Tikhonov-Regularized Hashing (STRH). The STRH enforces both the ℓ0-norm induced sparsity constraints and the Tikhonov regularization on the binary solution vectors which maximize cross-modal correlation. In addition, we raise the concerns on the challenging testing scenario of `Multi-modal Learning and Single-modal Prediction' (MLSP). Finally, we demonstrate that the STRH is an efficient hashing solutions by showing its superiority under the MLSP scenario. Lei Tian 0002, Xiaopeng Hong, Chunxiao Fan 0001, Yue Ming 0001, Matti Pietikäinen, Guoying Zhao 0001 |
ICIP | 4 |
| 2018 | Sparse projections matrix binary descriptors for face recognition
Chunxiao Fan 0001, Lei Tian 0002, Yue Ming 0001, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
Neurocomputing | 3 |
| 2018 | Improving deep neural network with Multiple Parametric Exponential Linear Units
Yang Li 0032, Chunxiao Fan 0001, Yong Li 0025, Yue Ming 0001 |
Neurocomputing | 5 |
| 2018 | Mesh motion scale invariant feature and collaborative learning for visual recognition
Yue Ming 0001, Jiakun Shi |
Multim. Tools Appl. | 1 |
| 2017 | Learning spherical hashing based binary codes for face recognition
Lei Tian 0002, Chunxiao Fan 0001, Yue Ming 0001 |
Multim. Tools Appl. | 3 |
| 2016 | A unified 3D face authentication framework based on robust local mesh SIFT feature
Yue Ming 0001, Xiaopeng Hong |
Neurocomputing | 1 |
| 2016 | Learning iterative quantization binary codes for face recognition
Lei Tian 0002, Chunxiao Fan 0001, Yue Ming 0001 |
Neurocomputing | 3 |
| 2015 | 3D human behavior recognition based on spatiotemporal texture featuresabstractNowadays, more and more activity recognition algorithms begin to improve recognition performance by combining the RGB and depth information. Although, the space-time volumes (STV) algorithm and the space-time local features algorithm can combine the RGB and depth information effectively, they also have their own defects. Such as they need expensive computational cost and they are not suitable for modeling nonperiodic activity. In this paper, we propose a novel algorithm for three dimensional human activity recognition that combines spatial-domain local texture features and spatio-temporal local texture features. On the one hand, in order to extract spatial local texture features, we mix the RGB and depth image sequence which have been applied with ViBe (Visual Background extractor) and binarization operator. Then we obtain the RGB-MOHBBI and depth-MOBHBI respectively and perform intersect operation on them. Afterwards, we extract LBP feature from the mixed MOHBBI to describe spatial domain feature. On the other hand, we follow the same background subtraction and binarization method to process the RGB and depth image sequences and get the spatial-temporal local texture features. And then, we project the three dimensional image volume on plane X-T and plane Y-T to get the spatio-temporal behavior volume change image to which we apply LBP operator to extract features that can represent human activity feature in spatio-temporal domain. At last, we combine the two local features that are extracted by LBP algorithm as one integrated feature of our model final output. Extensive experiments are conducted on the BUPT Arm Activity Dataset and the BUPT Arm And Finger Activity Dataset. The experimental results demonstrate the algorithm we proposed in this paper can make up for the deficiency of traditional activity recognition algorithms effectively and provide excellent experiment results on different databases of various complexities. Chunxiao Fan 0001, Lei Tian 0002, Guangchao Wang, Yue Ming 0001, Jiakun Shi, Yi Jin 0001 |
HSI | 4 |
| 2015 | Hand fine-motion recognition based on 3D Mesh MoSIFT feature descriptor
Yue Ming 0001 |
Neurocomputing | 1 |
| 2015 | Robust regional bounding spherical descriptor for 3D face recognition and emotion analysis
Yue Ming 0001 |
Image Vis. Comput. | 1 |
| 2014 | Rigid-area orthogonal spectral regression for efficient 3D face recognition
Yue Ming 0001 |
Neurocomputing | 1 |
| 2013 | A Mandarin edutainment system integrated virtual learning environments
Yue Ming 0001, Qiuqi Ruan, Guodong Gao |
Speech Commun. | 1 |
| 2012 | Activity Recognition from RGB-D Camera with 3D Local Spatio-temporal FeaturesabstractKinect, as a 3D digital capturing device, can collect the RGB and depth information of human activities rapidly. We study fusing the depth and RGB information for activity recognition. We introduce histogram color-based image thresholding to detect skin on human body, and use a GMM model to segment human hand areas. We design a new local descriptor, called a 3D Motion Scale-Invariant Feature Transform (3D MoSIFT), which can effectively detect interesting points based on both RGB and depth information, and consequently encode the visual and motion information from both to describe the interesting points. Experiments, based on a video dataset collected by a Kinect camera, show that adding depth information in the descriptor can distinctly improve the accuracy of human activity recognition. We introduce the F1-score measurement to evaluate and compare our performance with the other algorithms. Yue Ming 0001, Qiuqi Ruan, Alex Hauptmann 0001 |
ICME | 1 |
| 2012 | A remarkable standard for estimating the performance of 3D facial expression features
Xiaoli Li 0009, Qiuqi Ruan, Yue Ming 0001 |
Neurocomputing | 3 |
| 2012 | Robust sparse bounding sphere for 3D face recognition
Yue Ming 0001, Qiuqi Ruan |
Image Vis. Comput. | 1 |
| 2010 | Learning effective features for 3D face recognitionabstract3D images provide several advantages over 2D images for face recognition, especially when considering expression variations. In this paper, a novel framework is proposes for 3D-based face recognition. The key idea in the proposed algorithm is a representation of the facial surface, by what is called a Bending Invariant (BI), invariant to isometric deformations resulting from expressions and postures. In order to encode relationships in neighboring mesh nodes, Gaussian-Hermite moments are used for the obtained geometric invariant, which is a richer representation, due to their mathematical orthogonality and effectiveness in characterizing local details of the signal. The signature images are then decomposed into their principle components based on Spectral Regression Kernel Discriminate Analysis (SRKDA) resulting in a huge time saving. Our experiments are based on FRGC v2.0 face database. Experimental results show our framework provides better effectiveness and efficiency than many commonly used existing methods and handles variations in facial expression quite well. Yue Ming 0001, Qiuqi Ruan |
ICIP | 1 |