Mingyu You

dblp:74/6569 · DBLP profile ↗
← Back
39ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0003-2758-167XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 A novel multi-scale domain-adaptive long-distance forecasting model for electric shovel digging resistance load
Haifeng Yue, Chengjin Qin, Mingyu You, Pengcheng Xia 0005, Chengliang Liu 0001
Adv. Eng. Informatics5
2025 Imagine: Image-Guided 3D Part Assembly with Structure Knowledge Graph
abstract
3D part assembly is a promising task in 3D computer vision and robotics, focusing on assembling 3D parts together by predicting their 6-DoF poses. Like most 3D shape understanding tasks, existing methods primarily address this task by memorizing the poses of parts during the training process, leading to inaccuracies in complex assemblies and poor generalization to novel categories. In order to essentially improve the performance, structure knowledge of the target assembly is indispensable before assembling, which abstracts the potential part composition and their structural relationships. An image of the target assembly can serve as a common source for constructing this structure knowledge. Nevertheless, the image is far from enough, as its knowledge can be incomplete and ambiguous due to part occlusion and varying views. To tackle these issues, we propose Imagine, a novel Image-guided 3D part assembly framework with structure knowledge graph. As a novel assembly prior, the structure knowledge graph originates from the image and is refined as understanding the 3D parts. It encodes robust part-aware structural and semantic information of the assembly, guides the 3D parts from a coarse super-structure to a fine assembly, and co-evolves progressively throughout the assembly process. Extensive experiments demonstrate the state-of-the-art performance of our framework, along with strong generalization to novel images and categories.
Mingyu You, Bin He 0003
AAAI3
2025 Toward Better Out-Painting: Improving the Image Composition With Initialization Policy Model
Yanhao Ge, Mingyu You
ICCV4
2025 Completing 3D Partial Assemblies with View-Consistent 2D-3D Correspondence
Mingyu You, Bin He 0003
ICCV3
2025 Contrast, Imitate, Adapt: Learning Robotic Skills From Raw Human Videos
abstract
Learning robotic skills from raw human videos remains a non-trivial challenge. Previous works tackled this problem by leveraging behavior cloning or learning reward functions from videos. Despite their remarkable performances, they may introduce several issues, such as the necessity for robot actions, requirements for consistent viewpoints and similar layouts between human and robot videos, as well as low sample efficiency. To this end, our key insight is to learn task priors by contrasting videos and to learn action priors through imitating trajectories from videos, and to utilize the task priors to guide trajectories to adapt to novel scenarios. We propose a three-stage skill learning framework denoted as Contrast-Imitate-Adapt (CIA). An interaction-aware alignment transformer is proposed to learn task priors by temporally aligning video pairs. Then a trajectory generation model is used to learn action priors. To adapt to novel scenarios different from human videos, the Inversion-Interaction method is designed to initialize coarse trajectories and refine them by limited interaction. In addition, CIA introduces an optimization method based on semantic directions of trajectories for interaction security and sample efficiency. The alignment distances computed by IAAformer are used as the rewards. We evaluate CIA in six real-world everyday tasks, and empirically demonstrate that CIA significantly outperforms previous state-of-the-art works in terms of task success rate and generalization to diverse novel scenarios layouts and object instances.Note to Practitioners—This work aims to study robot skill learning from raw human videos. Compared with teleoperation or kinesthetic teaching in the laboratory, such learning method can flexibly utilize large-scale human videos available on the Internet, thereby improving the robot’s ability to generalize to various complex scenarios. Previous works on learning from videos usually have some issues, including requirements for robot actions, consistent viewpoints, similar layouts and low sample efficiency. To alleviate these issues, we propose a three-stage skill learning framework CIA. Temporal alignment is utilized to learn task priors through our proposed transformer-based model and self-supervised loss functions. A trajectory generation model is trained to learn the action priors. To further adapt to diverse scenarios, we propose a two-stage policy improvement method by initialization and interaction. An optimization method is introduced to ensure safe interaction and sample efficiency, where the optimization objective is guided by the learned task priors. The experimental results show that our CIA outperforms other state-of-the-art methods in task success rate and generalization to novel scenarios.
Zhifeng Qian, Mingyu You, Hongjun Zhou, Xuanhui Xu, Jinzhe Xue, Bin He 0003
IEEE Trans Autom. Sci. Eng.2
2025 Movement Primitive Categorization Balancing the Learnability and Adaptability
abstract
Given the rapid advancement of robotic technologies, robots will eventually enter our daily lives, performing complex long-horizon tasks. Complex long-horizon tasks, such as furniture assembly, typically contain dozens of subtasks with various scenes. Although complex, assembly is based on several reusable movements. Movement Primitive (MP) is a promising framework for learning reusable movements from demonstrations and adapting the learned movements to the test scenes. The critical step in employing MP methods is categorizing the unlabeled demonstrations into different MPs. However, current MP methods focus on individual MP learning using manually selected demonstrations, neglecting categorization. Manual categorization of demonstrations is easy to fall into suboptimal. If the demonstrations within the same category are too similar, the learned MP cannot be adapted to task scenes with various obstacles. Conversely, a significant distance between demonstrations leads to the MP’s failure in learning. To this end, we propose the following principle for MP categorization: balance the Learnability and Adaptability. Following this principle, we introduce an optimal transportation (OT)-based theoretical framework and a practical solution utilizing an auto-encoder network. We obtain the lower threshold of learnability by OT. Then we increase the adaptability of MP until it reaches the lower threshold of learnability. For complex long-horizon task learning, we propose a balanced MPs-based learning framework that contains four modules, termed BaMPs. BaMPs achieved success rates of 100% and 80%, respectively, in the 12-step and 20-step tasks.
Xuanhui Xu, Mingyu You, Hongjun Zhou, Zhifeng Qian, Jinzhe Xue, Weisheng Xu, Bin He 0003
IEEE Trans Autom. Sci. Eng.2
2025 PhysFiT: Physical-aware 3D Shape Understanding for Finishing Incomplete Assembly
abstract
Understanding the part composition and structure of 3D shapes is crucial for a wide range of 3D applications, including 3D part assembly and 3D assembly completion. Compared to 3D part assembly, 3D assembly completion is more complicated, which involves repairing broken or incomplete furniture that miss several parts with a toolkit. Given an incomplete assembly, 3D assembly completion seeks to identify its missing parts from multiple candidates, determine their poses, and produce complete assembly that is well-connected, structurally stable, and aesthetically pleasing. This task necessitates not only specialized knowledge of part composition but, more importantly, an awareness of physical constraints, i.e., connectivity, stability, and symmetry. Neglecting these constraints often results in assemblies that, although visually plausible, are impractical. To address this challenge, we propose PhysFiT, a physical-aware 3D shape understanding framework. This framework is built upon attention-based part relation modeling and incorporates connection modeling, simulation-free stability optimization and symmetric transformation consistency. We evaluate its efficacy on 3D part assembly and 3D assembly completion, a novel assembly task presented in this work. Extensive experiments demonstrate the effectiveness of PhysFiT in constructing geometrically sound and physically compliant assemblies.
Mingyu You, Hongjun Zhou, Bin He 0003
ACM Trans. Graph.2
2024 HybridBooth: Hybrid Prompt Inversion for Efficient Subject-Driven Generation
Shanyan Guan, Yanhao Ge, Ying Tai, Jian Yang 0003, Mingyu You
ECCV (9)6
2024 Scene Diffusion: Text-driven Scene Image Synthesis Conditioning on a Single 3D Model
abstract
Scene image is one of the important windows for showcasing product design. To obtain it, the standard 3D-based pipeline requires designer to not only create the 3D model of product, but also manually construct the entire scene in software, which hindering its adaptability in situations requiring rapid evaluation. This study aims to realize a novel conditional synthesis method to create the scene image based on a single-model rendering of the desired object and the scene description. In this task, the major challenges are ensuring the strict appearance fidelity of drawn object and the overall visual harmony of synthesized image. The former's achievement relies on maintaining an appropriate condition-output constraint, while the latter necessitates a well-balanced generation process for all regions of image. In this work, we propose Scene Diffusion framework to meet these challenges. Its first progress is introducing the Shading Adaptive Condition Alignment (SACA), which functions as an intensive training objective to promote the appearance consistency between condition and output image without hindering the network's learning to the global shading coherence. Afterwards, a novel low-to-high Frequency Progression Training Schedule (FPTS) is utilized to maintain the visual harmony of entire image by moderating the growth of high-frequency signals in the object area. Extensive qualitative and quantitative results are presented to support the advantages of the proposed method. In addition, we also demonstrate the broader uses of Scene Diffusion, such as its incorporation with ControlNet.
Mingyu You
ACM Multimedia3
2024 Improving the Conditional Fine-Grained Image Generation With Part Perception
abstract
Synthesizing the images in line with the given condition is a cardinal issue of image generation. The fine-grained conditional image generation, due to its emphasis on the fidelity of details, is of profound worth to the studies in this field. To learn the conditional distribution of data, the discriminating to class semantic of generated samples is necessitated. Though, most existing methods realize it solely based on the condensed global feature, which potentially impedes the model's focus on the detailed local features and in turn causes the inaccuracy or unstable local appearances in generated images. In this context, we propose PartGAN, which features a novel part perception mechanism to strengthen the model's concentration on the nuts-and-bolts of fine-grained objects. In proposed method, the image given to the discriminator will be deconstructed and encoded into a set of embeddings that represent the semantics of parts. This scheme not only assists the model to capture the discriminative local features more accurately, but also prevents the omission of other general local features. Under the effect of the newly designed condition loss term, every part of generated image is equally encouraged to be closer to the corresponding real part, which helps to ensure that the general parts have a stable appearance that conforms to class semantic. The experiments on the popular benchmarks show that the proposed method significantly improves the effect of the generation for fine-grained images.
Mingyu You, Ping Lu 0013
IEEE Trans. Multim.2
2023 3D Assembly Completion
abstract
Automatic assembly is a promising research topic in 3D computer vision and robotics. Existing works focus on generating assembly (e.g., IKEA furniture) from scratch with a set of parts, namely 3D part assembly. In practice, there are higher demands for the robot to take over and finish an incomplete assembly (e.g., a half-assembled IKEA furniture) with an off-the-shelf toolkit, especially in human-robot and multi-agent collaborations. Compared to 3D part assembly, it is more complicated in nature and remains unexplored yet. The robot must understand the incomplete structure, infer what parts are missing, single out the correct parts from the toolkit and finally, assemble them with appropriate poses to finish the incomplete assembly. Geometrically similar parts in the toolkit can interfere, and this problem will be exacerbated with more missing parts. To tackle this issue, we propose a novel task called 3D assembly completion. Given an incomplete assembly, it aims to find its missing parts from a toolkit and predict the 6-DoF poses to make the assembly complete. To this end, we propose FiT, a framework for Finishing the incomplete 3D assembly with Transformer. We employ the encoder to model the incomplete assembly into memories. Candidate parts interact with memories in a memory-query paradigm for final candidate classification and pose prediction. Bipartite part matching and symmetric transformation consistency are embedded to refine the completion. For reasonable evaluation and further reference, we design two standard toolkits of different difficulty, containing different compositions of candidate parts. We conduct extensive comparisons with several baseline methods and ablation studies, demonstrating the effectiveness of the proposed method.
Rufeng Zhang, Mingyu You, Hongjun Zhou, Bin He 0003
AAAI3
2023 Dynamic dense CRF inference for video segmentation and semantic SLAM
Mingyu You, Chaoxian Luo, Hongjun Zhou, Shaoqing Zhu 0003
Pattern Recognit.1
2023 MBFQuant: A Multiplier-Bitwidth-Fixed, Mixed-Precision Quantization Method for Mobile CNN-Based Applications
abstract
Deploying Convolutional Neural Network (CNN)-based applications to mobile platforms can be challenging due to the conflict between the restricted computing capacity of mobile devices and the heavy computational overhead of running a CNN. Network quantization is a promising way of alleviating this problem. However, network quantization can result in accuracy degradation and this is especially the case with the compact CNN architectures that are designed for mobile applications. This paper presents a novel and efficient mixed-precision quantization pipeline, called MBFQuant. It redefines the design space for mixed-precision quantization by keeping the bitwidth of the multiplier fixed, unlike other existing methods, because we have found that the quantized model can maintain almost the same running efficiency, so long as the sum of the quantization bitwidth of the weight and the input activation of a layer is a constant. To maximize the accuracy of a quantized CNN model, we have developed a Simulated Annealing (SA)-based optimizer that can automatically explore the design space, and rapidly find the optimal bitwidth assignment. Comprehensive evaluations applying ten CNN architectures to four datasets have served to demonstrate that MBFQuant can achieve improvements in accuracy of up to 19.34% for image classification and 1.12% for object detection, with respect to a corresponding uniform bitwidth quantized model.
Mingyu You, Kai Jiang 0004, Youzao Lian, Weisheng Xu
IEEE Trans. Image Process.2
2022 Mask encoding: A general instance mask representation for object segmentation
Rufeng Zhang, Tao Kong, Mingyu You
Pattern Recognit.4
2022 Part-Guided Attention Learning for Vehicle Instance Retrieval
abstract
Vehicle instance retrieval (IR) often requires one to recognize the fine-grained visual differences between vehicles. Besides the holistic appearance of vehicles which is easily affected by the viewpoint variation and distortion, vehicle parts also provide crucial cues to differentiate near-identical vehicles. Motivated by these observations, we introduce aPart-Guided Attention Network(PGAN) to pinpoint the prominent part regions and effectively combine the global and local information for discriminative feature learning. PGAN first detects the locations of different part components and salient regions regardless of the vehicle identity, which serves as thebottom-up attentionto narrow down the possible searching regions. To estimate the importance of detected parts, we propose aPart Attention Module(PAM) to adaptively locate the most discriminative regions with high-attention weights and suppress the distraction of irrelevant parts with relatively low weights. The PAM is guided by the identification loss and therefore providestop-down attentionthat enables attention to be calculated at the level of car parts and other salient regions. Finally, we aggregate the global appearance and local features together to improve the feature performance further. The PGAN combines part-guided bottom-up and top-down attention, global and local visual features in an end-to-end framework. Extensive experiments demonstrate that the proposed method achieves new state-of-the-art vehicle IR performance on four large-scale benchmark datasets.1
Xinyu Zhang 0015, Rufeng Zhang, Jiewei Cao, Dong Gong, Mingyu You, Chunhua Shen
IEEE Trans. Intell. Transp. Syst.5
2021 Diverse Knowledge Distillation for End-to-End Person Search
abstract
Person search aims to localize and identify a specific person from a gallery of images. Recent methods can be categorized into two groups, i.e., two-step and end-to-end approaches. The former views person search as two independent tasks and achieves dominant results using separately trained person detection and re-identification (Re-ID) models. The latter performs person search in an end-to-end fashion. Although the end-to-end approaches yield higher inference efficiency, they largely lag behind those two-step counterparts in terms of accuracy. In this paper, we argue that the gap between the two kinds of methods is mainly caused by the Re-ID sub-networks of end-to-end methods. To this end, we propose a simple yet strong end-to-end network with diverse knowledge distillation to break the bottleneck. We also design a spatial-invariant augmentation to assist model to be invariant to inaccurate detection results. Experimental results on the CUHK-SYSU and PRW datasets demonstrate the superiority of our method against existing approaches -- it achieves on par accuracy with state-of-the-art two-step methods while maintaining high efficiency due to the single joint model. Code is available at: https://git.io/DKD-PersonSearch.
Xinyu Zhang 0015, Jiawang Bian, Chunhua Shen, Mingyu You
AAAI5
2021 Fully integer-based quantization for mobile convolutional neural network inference
Mingyu You, Weisheng Xu
Neurocomputing2
2020 Mask Encoding for Single Shot Instance Segmentation
abstract
To date, instance segmentation is dominated by two-stage methods, as pioneered by Mask R-CNN. In contrast, one-stage alternatives cannot compete with Mask R-CNN in mask AP, mainly due to the difficulty of compactly representing masks, making the design of one-stage methods very challenging. In this work, we propose a simple single-shot instance segmentation framework, termed mask encoding based instance segmentation (MEInst). Instead of predicting the two-dimensional mask directly, MEInst distills it into a compact and fixed-dimensional representation vector, which allows the instance segmentation task to be incorporated into one-stage bounding-box detectors and results in a simple yet efficient instance segmentation framework. The proposed one-stage MEInst achieves 36.4% in mask AP with single-model (ResNeXt-101-FPN backbone) and single-scale testing on the MS-COCO benchmark. We show that the much simpler and flexible one-stage instance segmentation method, can also achieve competitive performance. This framework can be easily adapted for other instance-level recognition tasks. Code is available at: git.io/AdelaiDet
Rufeng Zhang, Zhi Tian, Chunhua Shen, Mingyu You, Youliang Yan
CVPR4
2020 Systematic evaluation of deep face recognition methods
Mingyu You, Yangliu Xu, Li Li 0008
Neurocomputing1
2019 Self-Training With Progressive Augmentation for Unsupervised Cross-Domain Person Re-Identification
abstract
Person re-identification (Re-ID) has achieved great improvement with deep learning and a large amount of labelled training data. However, it remains a challenging task for adapting a model trained in a source domain of labelled data to a target domain of only unlabelled data available. In this work, we develop a self-training method with progressive augmentation framework (PAST) to promote the model performance progressively on the target dataset. Specially, our PAST framework consists of two stages, namely, conservative stage and promoting stage. The conservative stage captures the local structure of target-domain data points with triplet-based loss functions, leading to improved feature representations. The promoting stage continuously optimizes the network by appending a changeable classification layer to the last layer of the model, enabling the use of global information about the data distribution. Importantly, we propose a new self-training strategy that progressively augments the model capability by adopting conservative and promoting stages alternately. Furthermore, to improve the reliability of selected triplet samples, we introduce a ranking-based triplet loss in the conservative stage, which is a label-free objective function based on the similarities between data pairs. Experiments demonstrate that the proposed method achieves state-of-the-art person Re-ID performance under the unsupervised cross-domain setting.
Xinyu Zhang 0015, Jiewei Cao, Chunhua Shen, Mingyu You
ICCV4
2018 PepAls: Performance Prediction and Algorithm Selection Framework for Data Mining Applications
abstract
With the explosive growth of digitalized data, knowledge discovery from data is attracting more attentions from different domains. Lots of data analysis and mining algorithms have been proposed in machine learning field. But most of these works concentrated on outperforming previous ones on more datasets. Performance on a given application dataset remains unclear. To evaluating all the available algorithms against a particular dataset can be computationally expensive and tedious. As a result, it's difficult for domain experts to choose a suitable data analysis algorithm. The frequently applied methods are always limited to several popular ones. This paper investigates the characteristics of dataset with 23 proposed quantitative indicators. With these quantitative descriptions, a novel framework PepAls is proposed as an initial attempt to recommend a suitable analysis algorithm for a given application dataset. The highest analysis result of the dataset, which could be reached by up to date methods, is also predicted by PepAls. Using PepAls, domain engineers can roughly preview the analysis result and quickly get the suitable algorithm. Newly proposed methods in machine learning can be easily introduced to application. PepAls is a framework, which can serve different kinds of learning problem when embedded different recommendation models. This paper uses multi-label learning problem as an example to illustrate the kernel idea. Elaborated experiments show that our approach is practical and helpful to narrow the gap of algorithm design and application.
Mingyu You, Xuanhui Xu
IEEE BigData1
2018 Reading car license plates using deep neural networks
Hui Li 0031, Peng Wang 0015, Mingyu You, Chunhua Shen
Image Vis. Comput.3
2018 An Extended Filtered Channel Framework for Pedestrian Detection
abstract
Pedestrian detection is an important example of object detection and has attracted much attention. Many works have shown that good image features provide high detection accuracy, and a few works have investigated enhancing low-level features (e.g., gradient and color features) using a filtered layer (i.e., convolutional layer) to obtain enhanced features or filtered channel features. To investigate whether these features are saturated, this paper adopts the concept of filtered channel features and strengthens them by adding more convolutional layers. Acting as convolution kernels, multilayer filters are applied to low-level features to obtain the extended filtered channel features, providing a powerful feature extractor with multilayer transformation for pedestrian detection. The proposed extended filtered channel framework (ExtFCF) achieves competitive performance on widely used benchmark datasets (Caltech, INRIA, and KITTI datasets), using only histogram of oriented gradient (HOG) and CIE-LUV [a color space composing of luminance (L) and two chrominance (UV) components by International Commission on Illumination] color features (HOG+LUV) as low-level features. One representative ExtFCF implementation achieves the best result compared with the current best traditional pedestrian detection methods on the Caltech dataset.
Mingyu You, Yubin Zhang, Chunhua Shen, Xinyu Zhang 0015
IEEE Trans. Intell. Transp. Syst.1
2017 Novel feature extraction method for cough detection using NMF
abstract
Cough is a common symptom in respiratory diseases. To provide valuable clinical information for cough diagnosis and monitoring, objectively evaluating the quantity and intensity of cough based on cough detection by pattern recognition technologies is needed. Cough detection aims to extract the boundaries of cough events from an audio stream. From spectral visualisation, it is found that the energy spectrum of cough signal spreads widely in the whole frequency band, which is very different from a speech signal. However, almost all feature extraction methods for cough detection in the previous work are derived from speech recognition region. In this study, to find the difference of cough and other audios in a more compact representation, non‐negative matrix factorisation (NMF) is exploited to extract the spectral structure from signals. Furthermore, the spectral structure from cough signal can be used as filter banks of feature extraction methods, which makes the filter banks more suitable for cough detection than manually designed ones. Besides, parameterisation for the spectral structure also provides an optimising strategy for the authors’ NMF‐based feature extraction method. Experiments are conducted on real data. The results demonstrate that NMF‐based feature extraction method has considerable potential in improving performance for cough detection.
Mingyu You, Zeqin Liu, Jia-Ming Liu, Xianghuai Xu, Zhongmin Qiu
IET Signal Process.1
2017 Sequential data feature selection for human motion recognition via Markov blanket
Hongjun Zhou, Mingyu You, Chao Zhuang
Pattern Recognit. Lett.2
2015 Audio signals encoding for cough classification using convolutional neural networks: A comparative study
abstract
Cough detection has considerable clinical value, which can provide an objective basis for assessment and diagnosis of respiratory diseases. Motivated by the great achievements of convolutional neural networks (CNNs) in recent years, we adopted 5 different ways to encode audio signals as images and treated them as the input of CNNs, so that image processing technology could be applied to analyze audio signals. In order to explore the optimal audio signals encoding method, we performed comparative experiments on medical dataset containing 70000 audio segments from 26 patients. Experimental results show that RASTA-PLP spectrum is the best method to encode audio signals as images with respect to cough classification task, which gives an average accuracy of 0.9965 in 200 iterations on test batches and a F1-score of 0.9768 on samples re-sampled from the test set. Therefore, the image processing based method is shown to be a promising choice for the process of audio signals.
Hui-Hui Wang, Jia-Ming Liu, Mingyu You, Guo-Zheng Li 0001
BIBM3
2014 Cough detection using deep neural networks
abstract
Cough detection and assessment have crucial clinical value for respiratory diseases. Subjective assessments are widely adopted in clinical measurement nowadays, but they are neither accurate nor reliable. An automatic and objective system for cough assessment is strongly expected. Automatic cough detection from audio signal has been studied by peer works. But they are still facing some difficulties like unsatisfactory detection accuracy or lacking large scale validation. In this paper, deep neural networks (DNN) are applied to model acoustic features in cough detection. A two step cough detection system is proposed based on deep neural networks(DNN) and hidden markov model(HMM). The experimental data set contains audio recordings from 20 patients with each recording lasting for about 24 hours. The performances of the newly proposed system were evaluated via sensitivity, specificity, F1 measure and macro average of recall. Different configurations of deep neural networks are evaluated. Experimental results show that many of the DNN configurations outperform Gaussian Mixture Model (GMM) on sensitivity, specificity and F1 measure respectively. On macro average of recall, 13.38% and 22.0% relative error reduction are achieved. The newly proposed system provides better performance and potential capacity for modeling big audio data on the cough detection task.
Jia-Ming Liu, Mingyu You, Guo-Zheng Li 0001, Xianghuai Xu, Zhongmin Qiu
BIBM2
2013 LEVIS: A hypertension dataset in traditional Chinese medicine
abstract
Traditional Chinese Medicine (TCM) is a significant complementary medicine in modern medicine. Hypertension is one of the important causes of Heart cerebrovascular diseases. Research on hypertension in TCM is an important and attractive topic. Considering the exhausted consumption when data collection, few TCM clinical datasets are publicly available. Public providing TCM clinical datasets will not only promote the development of TCM itself, but also encourage the data analysis researches. TCM syndrome information collection form and SOP (standard operating procedure) of hypertension TCM differentiations are designed firstly. Essential hypertension is focused in this study. With strict control measures, 908 reliable TCM hypertension clinical cases are recorded. 167 features, including 129 TCM symptoms from Inspection, auscultation and olfaction, inquiry and palpation, 22 laboratory indicators, and 16 common indexes, are investigated and collected in this database. 13 labels (TCM syndrome) are concentrated in the database. Web address to visit the LEVIS TCM hypertension sharing system is provided. Database structure and table organization is explained in detail. Two examples are given to demonstrate the customized data download methods. LEVIS Hypertension TCM Database is published publicly in the paper. According to the research objectives, different items can be organized into a dataset and downloaded. Academic and noncommercial users can access it at http://levis.tongji.edu.cn/datasets/.
Aihua Ou, Xiaozhong Lin, Guo-Zheng Li 0001, Mingyu You
BIBM4
2013 Customized management of clinical data in traditional Chinese medicine
abstract
Clinical records are the firsthand and vital evidences for Traditional Chinese Medicine (TCM) research. With deep Chinese culture background and years of clinical experience, a TCM clinical expert usually has his or her unique diagnostic methods. Preserving clinical cases of high-experienced TCM practitioners in their original tastes and exploring the in-depth knowledge is thus an important but arduous task. A novel system ISMAC (Intelligent System for Management and Analysis of Clinical cases in TCM) is introduced for customized management and intelligent analysis of TCM clinical data in this paper.
Mingyu You, Shixing Yan, Guo-Zheng Li 0001, Qing-Ce Zhao, Liaoyu Xu, Suying Huang
BIBM1
2012 Model selection for partial least squares based dimension reduction
Guo-Zheng Li 0001, Hai-Ni Qu, Mingyu You
Pattern Recognit. Lett.4
2009 An Enhanced Lipschitz Embedding Classifier for Multi-Emotion Speech Analysis
abstract
This paper proposes an Enhanced Lipschitz Embedding based Classifier (ELEC) for the classification of multi-emotions from speech signals. ELEC adopts geodesic distance to preserve the intrinsic geometry at all scales of speech corpus, instead of Euclidean distance. Based on the minimal geodesic distance to vectors of different emotions, ELEC maps the high dimensional feature vectors into a lower space. Through analyzing the class labels of the neighbor training vectors in the compressed low space, ELEC classifies the test data into six archetypal emotional states, i.e. neutral, anger, fear, happiness, sadness and surprise. Experimental results on clear and noisy data set demonstrate that compared with the traditional methods of dimensionality reduction and classification, ELEC achieves 15% improvement on average for speaker-independent emotion recognition and 11% for speaker-dependent.
Mingyu You, Guo-Zheng Li 0001, Jack Y. Yang, Mary Yang
Int. J. Pattern Recognit. Artif. Intell.1
2008 A Novel Classifier Based on Enhanced Lipschitz Embedding for Speech Emotion Recognition
Mingyu You, Guo-Zheng Li 0001, Luonan Chen, Jianhua Tao 0001
ICIC (1)1
2008 A robust multimodal approach for emotion recognition
Mingli Song, Mingyu You, Chun Chen 0001
Neurocomputing2
2007 Speech Emotion Recognition using an Enhanced Co-Training Algorithm
abstract
In previous systems of speech emotion recognition, supervised learning are frequently employed to train classifiers on lots of labeled examples. However, the labeling of abundant data requires much time and many human efforts. This paper presents an enhanced co-training algorithm to utilize a large amount of unlabeled speech utterances for building a semi-supervised learning system. It uses two conditionally independent attribute views(i.e. temporal features and statistic features) of unlabeled examples to augment a much smaller set of labeled examples. Our experimental results demonstrate that compared with the method based on the supervised training, the proposed system makes 9.0% absolute improvement on female model and 7.4% on male model in terms of average accuracy. Moreover, the enhanced co-training algorithm achieves comparable performance to the co-training prototype, while it can reduce the classification noise which is produced by error labeling in the process of semi-supervised learning.
Jia Liu 0030, Chun Chen 0001, Jiajun Bu, Mingyu You, Jianhua Tao 0001
ICME4
2006 Emotion Recognition from Noisy Speech
abstract
This paper presents an emotion recognition system from clean and noisy speech. Geodesic distance was adopted to preserve the intrinsic geometry of emotional speech. Based on the geodesic distance estimation, an enhanced Lipschitz embedding was developed to embed the 64-dimensional acoustic features into a six-dimensional space. In order to avoid the problems brought by noise reduction, emotion recognition from noisy speech was performed directly. Linear discriminant analysis (LDA), principal component analysis (PCA) and feature selection by sequential forward selection (SFS) with support vector machine (SVM) were also included to compress acoustic features before classifying the emotional states of clean and noisy speech. Experimental results demonstrate that compared with other methods, the proposed system makes approximately 10% improvement. The performance of our system is also robust when speech data is corrupted by increasing noise
Mingyu You, Chun Chen 0001, Jiajun Bu, Jia Liu 0030, Jianhua Tao 0001
ICME1
2005 CHAD: A Chinese Affective Database
Mingyu You, Chun Chen 0001, Jiajun Bu
ACII1
2004 Audio-visual based emotion recognition using tripled hidden Markov model
abstract
Emotion recognition is one of the latest challenges in intelligent human/machine communication. Most of previous work on emotion recognition focused on extracting emotions from visual or audio information separately. A novel approach is presented in this paper to recognize the human emotion which uses both visual and audio from video clips. A tripled hidden Markov model is introduced to perform the recognition which allows the state asynchrony of the audio and visual observation sequences while preserving their natural correlation over time. The experimental results show that this approach outperforms only using visual or audio separately.
Mingli Song, Chun Chen 0001, Mingyu You
ICASSP (5)3
2004 Speech Emotion Recognition and Intensity Estimation
Mingli Song, Chun Chen 0001, Jiajun Bu, Mingyu You
ICCSA (4)4
2004 Speech Driven Facial Animation Using Chinese Mandarin Pronunciation Rules
Mingyu You, Jiajun Bu, Chun Chen 0001, Mingli Song
ICCSA (3)1