Yueliang Qian

dblp:43/7006 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 since 2021Artificial intelligence and machine learning · 6Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Learning paradigms · 77% Segmentation and scene understanding · 23%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Human-computer interaction and pervasive computing
1 paper
Interaction techniques and input · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
multiple instance learning
0.412020
Specialized Decision Surface and Disentangled Feature for Weakly-Supervised Polyphonic Sound Event Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Audio and music processing
sound event detection
0.412020
Specialized Decision Surface and Disentangled Feature for Weakly-Supervised Polyphonic Sound Event Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Interaction techniques and input
touch and gesture input
0.112011
A ring-shaped interactive device for large remote display and mobile device control · UbiComp 2011
Interaction techniques and input › input device
input device design
0.012011
A ring-shaped interactive device for large remote display and mobile device control · UbiComp 2011

Methods — techniques the papers use, named apart from their topics

neural network · 0.9disentangled feature · 0.9attention pooling · 0.9gyroscope · 0.1accelerometer · 0.1SVM · 0.1MFCC · 0.1
YearPublicationVenuePosition
2022 MAL: Multi-modal Attention Learning for Tumor Diagnosis Based on Bipartite Graph and Multiple Branches
Menglei Jiao, Hong Liu 0007, Jianfang Liu, Hanqiang Ouyang, Huishu Yuan, Yueliang Qian
MICCAI (3)8
2021 Speech Synthesis of Chinese Braille with Limited Training Data
abstract
This paper describes to our knowledge the first Chinese Braille speech synthesis system. The system consists of modules of Braille front-end processing, prosody prediction, and speech synthesis. The Braille front-end processing includes conversion from the common Braille to Pinyin, and a high-precision Chinese character prediction model. To achieve high precision prosody prediction under limited corpus conditions, we propose a prosody prediction model based on the RoBERTa pre-trained model, which achieves an accuracy of 94.42%. Finally, a real-time TTS system based on Tacotron2 and LPCNet is proposed. We modify Tacotron2, including introducing a forward attention mechanism and extending the autoregressive correlation step size to obtain more natural speech.
Jianguo Mao, Hong Liu 0007, Yueliang Qian
ICME5
2020 Multi-Branch Learning for Weakly-Labeled Sound Event Detection
abstract
There are two sub-tasks implied in the weakly-supervised SED: audio tagging and event boundary detection. Current methods which combine multi-task learning with SED requires annotations both for these two sub-tasks. Since there are only annotations for audio tagging available in weakly-supervised SED, we design multiple branches with different learning purposes instead of pursuing multiple tasks. Similar to multiple tasks, multiple different learning purposes can also prevent the common feature which the multiple branches share from overfitting to any one of the learning purposes. We design these multiple different learning purposes based on combinations of different MIL strategies and different pooling methods. Experiments on the DCASE 2018 Task 4 dataset and the URBAN-SED dataset both show that our method achieves competitive performance.
Hong Liu 0007, Yueliang Qian
ICASSP5
2020 Guided Learning for Weakly-Labeled Semi-Supervised Sound Event Detection
abstract
We propose a simple but efficient method termed Guided Learning for weakly-labeled semi-supervised sound event detection (SED). There are two sub-targets implied in weakly-labeled SED: audio tagging and boundary detection. Instead of designing a single model by considering a trade-off between the two sub-targets, we design a teacher model aiming at audio tagging to guide a student model aiming at boundary detection to learn using the unlabeled data. The guidance is guaranteed by the audio tagging performance gap of the two models. In the meantime, the student model liberated from the trade-off is able to provide more excellent boundary detection results. We propose a principle to design such two models based on the relation between the temporal compression scale and the two sub-targets. We also propose an end-to-end semi-supervised learning process for these two models to enable their abilities to rise alternately. Experiments on the DCASE2018 Task4 dataset show that our approach achieves competitive performance.
Hong Liu 0007, Yueliang Qian
ICASSP4
2020 Specialized Decision Surface and Disentangled Feature for Weakly-Supervised Polyphonic Sound Event Detection
abstract
In this article, a special decision surface for the weakly-supervised sound event detection (SED) and a disentangled feature (DF) for the multi-label problem in polyphonic SED are proposed. We approach SED as a multiple instance learning (MIL) problem and utilize a neural network framework with a pooling module to solve it. General MIL approaches include two kinds: the instance-level approaches and embedding-level approaches. We present a method of generating instance-level probabilities for the embedding level approaches which tend to perform better than the instance-level approaches in terms of bag-level classification but can not provide instance-level probabilities in current approaches. Moreover, we further propose a specialized decision surface (SDS) for the embedding-level attention pooling. We analyze and explained why an embedding-level attention module with SDS is better than other typical pooling modules from the perspective of the high-level feature space. As for the problem of the unbalanced dataset and the co-occurrence of multiple categories in the polyphonic event detection task, we propose a DF to reduce interference among categories, which optimizes the high-level feature space by disentangling it based on class-wise identifiable information and obtaining multiple different subspaces. Experiments on the dataset of DCASE 2018 Task 4 show that the proposed SDS and DF significantly improve the detection performance of the embedding-level MIL approach with an attention pooling module and outperform the first place system in the challenge by $\mathbf {6.6}$ percentage points.
Hong Liu 0007, Yueliang Qian
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 CBConv: Service for Automatic Conversion of Chinese Characters into Braille with High Accuracy
abstract
The conversion of Chinese Characters into Braille faces the challenges of Braille word segmentation, pronunciation determination, and tone marking. In this paper, we present the CBConv service for automatic conversion of Chinese characters into Braille with high accuracy. The service supports two modes: real-time conversion of Chinese strings and asynchronous conversion of documents. To address the challenges described above, an end-to-end deep learning framework is proposed, which can perform pronunciation determination, word segmentation and tone marking simultaneously using a deep neural network. Accuracy evaluation and user study are conducted, which demonstrate the usability of the proposed system.
Jinghua Zhong, Hong Liu 0007, Yueliang Qian
ASSETS5
2019 Effective Optical Braille Recognition Based on Two-Stage Learning for Double-Sided Braille Image
Renqiang Li, Hong Liu 0007, Yueliang Qian
PRICAI (3)4
2018 Estimation of Spatial-Temporal Gait Parameters based on the Fusion of Inertial and Film-Pressure Signals
Cheng Wang 0005, Zhou Long, Mingming Gao, Xiaoping Yun, Yueliang Qian
BIBM7
2018 Multiple Visual Fields Cascaded Convolutional Neural Network for Breast Cancer Detection
Haomiao Ni, Hong Liu 0007, Zichao Guo, Taijiao Jiang, Kuansong Wang, Yueliang Qian
PRICAI (1)7
2018 RGB-D joint modelling with scene geometric information for indoor semantic segmentation
Hong Liu 0007, Wenshan Wu, Yueliang Qian
Multim. Tools Appl.4
2017 Efficient Multi-scale Plane Extraction Based RGBD Video Segmentation
Hong Liu 0007, Yueliang Qian
MMM (1)4
2017 Improving speech transcription by exploiting user feedback and word repetition
Hong Liu 0007, Yueliang Qian
Multim. Tools Appl.4
2016 Action Recognition Based on Optimal Joint Selection and Discriminative Depth Descriptor
Haomiao Ni, Hong Liu 0007, Yueliang Qian
ACCV (2)4
2016 A novel obstacle detection method based on distortion of laser pattern
abstract
Most visual based obstacle detection methods usually detect obstacles on the whole image, which are influenced by complex background and various illuminations. This paper proposes a novel and simple obstacle detection method DLP based on analyzing Distortion of Laser Pattern. To reduce computation, a laser pattern with two cross lines is used to project onto the front ground. Firstly, a segmentation algorithm based on color and edge is proposed to extract laser pixels. Then, skeletons of laser pattern are quickly extracted using table scanning strategy to further reduce noise. Secondly, seed-filling is used to extract the region of skeletons, and then laser lines are detected by Hough and merged by region based rule. Finally, according to the distortion of laser pattern projected on different obstacles, we propose a hierarchical discriminant model, which can classify the front scene as free or having obstacles. And the obstacle can be further classified as raised or sunken obstacle, and ascending or descending stairs. Experimental results show the robustness and effectiveness of our DLP method in different scenes, including indoor and outdoor, day and night with 10f/s.
Zichao Guo, Hong Liu 0007, Yueliang Qian
ICME3
2014 Segment and Label Indoor Scene Based on RGB-D for the Visually Impaired
Hong Liu 0007, Yueliang Qian
MMM (1)4
2013 Related HOG Features for Human Detection Using Cascaded Adaboost and SVM Classifiers
Hong Liu 0007, Tao Xu 0029, Yueliang Qian
MMM (2)4
2012 A Fast and Robust Pedestrian Detection Framework Based on Static and Dynamic Information
abstract
With the powerful development of pedestrian detection technique based on sliding-window and machine-learning, detection-based tracking systems have become increasingly popular. Most of these systems rely on existing static pedestrian detectors only despite the obvious potential motion information for people detection. This paper proposes a novel pedestrian detection framework fusing static and dynamic features. Motion cue is firstly used to detect potential pedestrian regions. Secondly, static detector scans potential regions to get candidate pedestrian detections. Final detection results are improved by removing false detections based on their motion distribution. The proposed framework significantly raises detection speed and detection performance. Static detector of pedestrian in this paper is trained by AdaBoost with simplified HOG feature (1HOG). Additionally, we introduce a detection-window-pyramid based scanning strategy for quickly extracting 1HOG features. The experimental results on several public data sets show the effectiveness of the proposed approach.
Tao Xu 0029, Hong Liu 0007, Yueliang Qian
ICME3
2011 A ring-shaped interactive device for large remote display and mobile device control
abstract
In this demonstration, a novel human-computer interaction device is proposed to realize finger touching for large display and mobile device control, without a touchscreen or a touch pad. In this method, interaction commands are input in a same way as traditional touchscreen and touchpad, which is convenient to develop applications for long-distance operation of display and mobile devices. An embedded module is designed to collect bone-conducted friction sound, acceleration and gyroscope sensor data, corresponding to the behavior and direction of interaction commands. For algorithm, modified MFCC and SVM are applied in sound processing and probability calculation.
Yiqiang Chen 0001, Yueliang Qian
UbiComp3
2008 Fast commercial detection based on audio retrieval
abstract
Automatic detection of commercials in digital multimedia material is a challenging task with many applications. This paper presents a novel approach to fast commercial detection based on audio retrieval. It is based on the idea of segmenting energy envelope of audio into units, using only audio signal for matching on a commercial database. Fast searching and matching can be performed with high accuracy, by searching and by novel similarity function based on units. Experimental results show that 96.8% recall rate and 98.7% precision rate can be achieved under 0.125 real-time.
Yueliang Qian, Qun Liu 0001, Shouxun Lin
ICME3
2007 The PICA Framework for Performance Analysis of Pattern Recognition Systems and Its Application in Broadcast News Segmentation
Meiyin Li, Shouxun Lin, Yueliang Qian, Qun Liu 0001
IEA/AIE4
2007 HTRDP evaluations on Chinese information processing and intelligent human-machine interface
Qun Liu 0001, Hong Liu 0007, Le Sun 0001, Sheng Tang, Deyi Xiong, Hongxu Hou, Yuanhua Lv, Shouxun Lin, Yueliang Qian
Frontiers Comput. Sci. China11
2005 Parsing the Penn Chinese Treebank with Semantic Knowledge
Deyi Xiong, Shuanglong Li, Qun Liu 0001, Shouxun Lin, Yueliang Qian
IJCNLP5