Ming Dong 0001

dblp:22/2379-1 · DBLP profile ↗
← Back
100ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0001-8133-7809ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 38 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 14 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ESM-AnatTractNet: Advanced deep learning model of true positive eloquent white matter tractography to improve preoperative evaluation of pediatric epilepsy surgery
abstract
Accurate preoperative identification of true positive white matter pathways involved in critical eloquent functions such as motor, language, and vision plays a vital role in minimizing the risk of postoperative functional deficits and improving postoperative functional outcomes in pediatric epilepsy surgery. This study proposes a novel deep learning model: "ESM-AnatTractNet" that can accurately classify true positive eloquent white matter pathways across preoperative diffusion weighted imaging tractography data of 85 drug-resistant epilepsy patients (age: 10.70 ± 4.41 years). To enhance geometric and anatomical consistency of true positive tract classification, the ESM-AnatTractNet integrated two features in a point-cloud-based framework, 1) electro-physiologically confirmed spatial coordinates using electrical stimulation mapping (ESM) and 2) anatomically-contexted labels of the end-to-end neural connection using a standard brain atlas. Its overall performance was validated by accurately classifying 14 eloquent functional areas in whole brain, objectively optimizing resection margins to preserve eloquent functions using Kalman filter, and precisely predicting postoperative language outcomes using canonical correlation. Our ESM-AnatTractNet outperformed other baseline models, achieving an accuracy of 97% in correctly classifying eloquent areas within 10mm spatial resolution of clinical subdural grid electroencephalography. The Kalman filter analysis achieved 94% accuracy in predicting no deficits when the ESM-AnatTractNet-defined preservation zones were not resected. Postoperative decrease in language-related white matter connection efficacy defined by the ESM-AnatTractNet analysis was significantly associated with worse postoperative language outcome (R=0.73, p < 0.001). Our findings demonstrate that the ESM-AnatTractNet improves non-invasive localization of true positive eloquent white matter pathways, supporting its potential to enhance current preoperative evaluation of pediatric epilepsy surgery.
Min-Hee Lee, Bohan Xiao, Soumyanil Banerjee, Hiroshi Uda, Yoon Ho Hwang, Csaba Juhász, Eishi Asano, Ming Dong 0001, Jeong-Won Jeong
Medical Image Anal.8
2025 Deterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual Approximators
abstract
Image-to-Image (I2I) translation involves converting an image from one domain to another. Deterministic I2I translation, such as in image super-resolution, extends this concept by guaranteeing that each input generates a consistent and predictable output, closely matching the ground truth (GT) with high fidelity. In this paper, we propose a denoising Brownian bridge model with dual approximators (Dual-approx Bridge), a novel generative model that exploits the Brownian bridge dynamics and two neural network-based approximators (one for forward and one for reverse process) to produce faithful output with negligible variance and high image quality in I2I translations. Our extensive experiments on benchmark datasets including image generation and super-resolution demonstrate the consistent and superior performance of Dual-approx Bridge in terms of image quality and faithfulness to GT when compared to both stochastic and deterministic baselines. Project page and code: https://github.com/bohan95/dual-app-bridge
Bohan Xiao, Peiyong Wang, Qisheng He, Ming Dong 0001
CVPR4
2025 Label correlation preserving visual-semantic joint embedding for multi-label zero-shot learning
Zhongchen Ma, Guangchen Wang, Qirong Mao, Ming Dong 0001
Multim. Tools Appl.5
2025 Infrared-Visible Image Fusion Using Dual-Branch Auto-Encoder With Invertible High-Frequency Encoding
abstract
In the field of Infrared-Visible Image Fusion (IVIF), the preservation of details, edges, and texture is crucial for generating high-quality fused images. However, a major challenge arises due to the inevitable loss of high-frequency information during feature extraction, resulting in fused images that lack significant details. In this paper, we propose a dual-branch auto-encoder by exploiting an invertible high-frequency branch for detailed feature preservation and a transformer-based low-frequency branch for global dependencies modeling. First, the high-frequency branch employs the wavelet transforms and an Invertible Neural Networks (INN)-based encoder to model high-frequency features through an invertible transformation, including a forward process for image fusion and an inverse process for original image reconstruction. Additionally, a high-frequency loss is designed to enhance the high-frequency feature representation for high-quality image fusion. Second, a low-frequency branch based on a transformer encoder and an adaptive fusion module is introduced to capture the global contextual features of the infrared and visible images. Finally, the decoder integrates the low- and high-frequency features from both branches to generate the final fused image. Image fusion, object detection, and semantic segmentation experiments conducted on public datasets such as TNO, MFNet, and M3FD, show that our method outperforms the state-of-the-art (SOTA) image fusion methods.
Qirong Mao, Ming Dong 0001, Yongzhao Zhan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Cascaded Network with Hierarchical Self-Distillation for Sparse Point Cloud Classification
abstract
Incomplete point clouds that are generally scanned with flaws or large gaps can cause to inaccurate predictions by learning-based classification methods. To address this issue, in this paper, we propose an end-to-end architecture that compensates for and identifies partial point clouds on the fly. First, we propose a cascaded solution that integrates both the upstream and downstream networks simultaneously, allowing the task-oriented downstream to identify the points generated by the completion-oriented upstream. These two streams complement each other, resulting in improved performance for both completion and downstream-dependent tasks. Second, to explicitly understand the predicted points’ pattern, we introduce hierarchical self-distillation (HSD), which can be applied to arbitrary hierarchy-based point cloud methods. On the classification task, our proposed method performs competitively on the synthetic dataset and achieves superior results on the challenging real-world benchmark when compared to the state-of-the-art models.
Kaiyue Zhou, Ming Dong 0001, Peiyuan Zhi, Shengjin Wang
ICME2
2024 Conditional Diffusion Model with Spatial Attention and Latent Embedding for Medical Image Segmentation
Behzad Hejrati, Soumyanil Banerjee, Carri Glide-Hurst, Ming Dong 0001
MICCAI (9)4
2024 Score-Based Image-to-Image Brownian Bridge
abstract
Image-to-image translation is defined as the process of learning a mapping between images from a source domain and images from a target domain. The probabilistic structure that maps a fixed initial state to a pinned terminal state through a standard Wiener process is a Brownian bridge. In this paper, we propose a score-based Stochastic Differential Equation (SDE) approach via the Brownian bridges, termed the Amenable Brownian Bridges (A-Bridges), to image-to-image translation tasks as an unconditional diffusion model. Our framework embraces a large family of Brownian bridge models, while the discretization of the linear A-Bridge exploits its advantage that provides the explicit solution in a closed form and thus facilitates the model training. Our model enables the accelerated sampling and has achieved record-breaking performance in sample quality and diversity on benchmark datasets following the guidance of its SDE structure.
Peiyong Wang, Bohan Xiao, Qisheng He, Carri Glide-Hurst, Ming Dong 0001
ACM Multimedia5
2024 On Local Temporal Embedding for Semi-Supervised Sound Event Detection
abstract
Semi-supervised sound event detection (SSED) task requires recognizing the categories of events and marking each event's onset and offset times in a mixed audio recording using a small amount of weakly labeled and a large scale of unlabeled data. So, exploring local temporal information, i.e., local discrimination and local correlations in the time domain, is essential for SSED, and in particular, for precise event boundary detection. Besides, as manual-labeled datasets are scarce, SSED tasks require effectively exploiting unlabelled data to reduce overfitting, typically through regularization techniques. Recently, self-supervised learning provided a viable solution to leverage unlabeled data for effective feature learning in various downstream tasks. In this paper, we propose LTE-Net, a novel multitask framework, to learn the Local Temporal Embedding for SSED. Specifically, LTE-Net first locally down-samples the input spectrogram and learns the token embeddings with a high temporal resolution (i.e., local discrimination). Then, LTE-Net effectively models the local correlations among the token embeddings through self-supervised masked spectrogram modeling. Finally, a novel joint (self- and semi-supervision) regularization framework is employed for the training of LTE-Net to effectively leverage unlabeled data in SSED. Extensive experiments on DCASE 2019, 2020 and 2021 SSED datasets show that LTE-Net significantly outperformed existing methods and achieved 2.1% to 8.7%, 2.1% to 3.9% and 1.2% to 6.1% performance gains on the evaluation set in 2019, 2020 and 2021 datasets, respectively.
Lijian Gao, Qirong Mao, Ming Dong 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Exploring Prototype-Anchor Contrast for Semantic Segmentation
abstract
Pixel-wise contrastive learning recently offers a new training paradigm in semantic segmentation by directly shaping the pixel embedding space. Compared with pixel-pixel contrast that often requires large memory and high computation cost, pixel-prototype contrast exploits the semantic correlations among pixels in a more efficient way by pulling positive pixel-prototype pairs close and pushing negative pairs apart. However, most existing work treats pixels as anchors to form contrast, either failing to capture the intra-class variance or introducing extra computational overhead. In this work, we propose Prototype-Anchor Contrast (ProAC), a novel prototypical contrastive learning paradigm that strengthens pixel-prototype associations in a simple yet effective fashion. First, ProAC pre-defines class prototypes (serving as cluster centroids) by exploiting the uniformity on the hypersphere in the feature space and thus requires no prototype updating during network optimization, which greatly simplifies the network training process. Second, by treating prototypes as anchors, ProAC builds a novel prototype-to-pixel learning path, where a large amount of negative pixels can naturally be generated to describe rich semantic information without relying on auxiliary sample augmentation techniques. Finally, as a plug-and-play regularization term, ProAC can be attached to most existing segmentation models and assist the network optimization by directly shaping the pixel embedding space. Extensive experiments on different benchmarks show that our ProAC brings an mIoU increase from 1.4% to 2.0% for fully-supervised models and from 0.9% to 6.0% for domain-adaptive models, respectively. It also leads to a gain of mIoU, ranging from 1.8% to 2.7% in more challenging cases, including different resolutions, diverse illuminations and masked scenarios.
Qinghua Ren, Shijian Lu, Qirong Mao, Ming Dong 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Transferring Lottery Tickets in Computer Vision Models: a Dynamic Pruning Approach
abstract
Deep neural networks can achieve state-of-the-art results on small size datasets by transferring the backbone from a network pre-trained on large datasets. Recent work has shown that pruned networks can also be used as pre-trained models in transfer learning. In this paper, we proposed a novel framework, Transferring Lottery Ticket (TLT), to adapt both masks and weights of a pre-trained and pruned network dynamically during the knowledge transfer to downstream tasks. We show that the lottery tickets of downstream tasks are dramatically different from each other and from the one obtained from the pre-trained network. Thus, both masks and weights need to be learned to better adapt a pre-trained model to the target domain. Our extensive experiments on multiple computer vision tasks, such as image classification and segmentation, show that the transferred networks with adapted masks outperform the ones with original masks at various pruning ratios.
Qisheng He, Ming Dong 0001
IEEE Big Data2
2023 Learning Pruned Structure and Weights Simultaneously from Scratch: an Attention based Approach
abstract
As a deep learning model typically contains millions of trainable weights, there has been a growing demand for a more efficient network structure with reduced storage space and improved run-time efficiency. Pruning is one of the most popular network compression techniques. In this paper, we propose a novel unstructured pruning pipeline, Attention-based Simultaneous sparse structure and Weight Learning (ASWL). In ASWL, an efficient algorithm is proposed to calculate the pruning ratios layer-wisely from attentions, and both weights for the dense network and the sparse network are tracked so that the pruned structure is simultaneously learned from randomly initialized weights. Our experiments on MNIST, Cifar10, and ImageNet show that ASWL achieves superior pruning results in terms of accuracy, pruning ratio and operating efficiency when compared with state-of-the-art network pruning methods.
Qisheng He, Weisong Shi, Ming Dong 0001
IEEE Big Data3
2023 Joint-Former: Jointly Regularized and Locally Down-sampled Conformer for Semi-supervised Sound Event Detection
Lijian Gao, Qirong Mao, Ming Dong 0001
INTERSPEECH3
2023 TE-KWS: Text-Informed Speech Enhancement for Noise-Robust Keyword Spotting
abstract
Keyword spotting (KWS) presents a formidable challenge, particularly in high-noise environments. Traditional denoising algorithms that rely solely on speech have difficulty recovering speech that has been severely corrupted by noise. In this investigation, we develop an adaptive text-informed denoising model to bolster reliable keyword identification in the presence of considerable noise degradation. The whole proposed TE-KWS incorporates a tripartite branch structure, where the speech branch (SB) takes noisy speech as input which provides the raw speech information, the alignment branch (AB) accommodates aligned text input which facilitates accurate restoration of the corresponding speech when text with alignment is preserved, and the text branch (TB) handles unaligned text which prompts the model to autonomously learn the alignment between speech and text. To make the proposed denoising model more beneficial for KWS, following the training of the whole model,the alignment branch (AB) is frozen, and the model is fine-tuned by leveraging its speech restoration and forced alignment capabilities. Subsequently, the input for the text branch (TB) is supplanted with designated keywords, and a heavier denoising penalty is applied on the keywords period, thereby explicitly intensifying the speech restoration ability of the model for keywords. Finally, the Combined Adversarial Domain Adaptation (CADA) is implemented to enhance the robustness of KWS with regard to data pre-and post-speech enhancement (SE). Experimental results indicate that our approach not only markedly ameliorates highly corrupted speech, achieving SOTA performance for marginally corrupted speech, but also bolsters the efficacy and generalizability of prevailing mainstream KWS models.
Dong Liu 0037, Qirong Mao, Lijian Gao, Qinghua Ren, Zhenghan Chen, Ming Dong 0001
ACM Multimedia6
2023 Infrastructure-level Support for GPU-Enabled Deep Learning in DATAVIEW
Junwen Liu, Ziyun Xiao, Shiyong Lu, Dunren Che, Ming Dong 0001, Changxin Bai
Future Gener. Comput. Syst.5
2023 Automated Identification of Toxic Code Reviews Using ToxiCR
abstract
Toxic conversations during software development interactions may have serious repercussions on a Free and Open Source Software (FOSS) development project. For example, victims of toxic conversations may become afraid to express themselves, therefore get demotivated, and may eventually leave the project. Automated filtering of toxic conversations may help a FOSS community maintain healthy interactions among its members. However, off-the-shelf toxicity detectors perform poorly on a software engineering dataset, such as one curated from code review comments. To counter this challenge, we presentToxiCR, a supervised learning based toxicity identification tool for code review interactions. ToxiCR includes a choice to select one of the 10 supervised learning algorithms, an option to select text vectorization techniques, eight preprocessing steps, and a large-scale labeled dataset of 19,651 code review comments. Two out of those eight preprocessing steps are software engineering domain specific. With our rigorous evaluation of the models with various combinations of preprocessing steps and vectorization techniques, we have identified the best combination for our dataset that boosts 95.8% accuracy and an 88.9% F1-score in identifying toxic texts. ToxiCR significantly outperforms existing toxicity detectors on our dataset. We have released our dataset, pre-trained models, evaluation results, and source code publicly, which is available at https://github.com/WSU-SEAL/ToxiCR .
Jaydeb Sarker, Asif Kamal Turzo, Ming Dong 0001, Amiangshu Bosu
ACM Trans. Softw. Eng. Methodol.3
2022 "Zero-Shot" Point Cloud Upsampling
abstract
Recent supervised point cloud upsampling methods are re-stricted by the size of training data and are limited in terms of covering all object shapes. Besides the challenges faced due to data acquisition, the networks also struggle to gener-alize on unseen records. In this paper, we present an internal point cloud upsampling approach at a holistic level referred to as “Zero-Shot” Point Cloud Upsampling (ZSPU). Our approach is data agnostic and relies solely on the internal infor-mation provided by a particular point cloud without patching in both self-training and testing phases. This single-stream design significantly reduces the training time by learning the relation between low resolution (LR) point clouds and their high (original) resolution (HR) counterparts. This association will then provide super resolution (SR) outputs when origi-nal point clouds are loaded as input. ZSPU achieves com-petitive/superior quantitative and qualitative performances on benchmark datasets when compared with other upsampling methods.
Kaiyue Zhou, Ming Dong 0001, Suzan Arslanturk
ICME2
2022 Adaptive Hierarchical Pooling for Weakly-supervised Sound Event Detection
abstract
In Weakly-supervised Sound Event Detection (WSED), the ground truth of training data contains the presence or absence of each sound event only at the clip-level (i.e., no frame-level annotations). Recently, WSED has been formulated under the multi-instance learning framework, and a critical component within this formulation is the design of the temporal pooling function. In this paper, we propose an adaptive hierarchical pooling (HiPool) for WSED, which combines the advantages of max pooling in audio tagging and weighted average pooling in audio localization through a novel hierarchical structure and learns event-wise optimal pooling functions through continuous relaxation-based joint optimization. Extensive experiments on benchmark datasets show that HiPool outperforms the current pooling methods and greatly improves the performance of WSED. HiPool also has great generality - ready to be plugged into any WSED models.
Lijian Gao, Qirong Mao, Ming Dong 0001
ACM Multimedia4
2021 SA-GAN: Structure-Aware GAN for Organ-Preserving Synthetic CT Generation
Hajar Emami, Ming Dong 0001, Siamak P. Nejad-Davarani, Carri Glide-Hurst
MICCAI (6)2
2021 Reproducibility Companion Paper: On Learning Disentangled Representation for Acoustic Event Detection
abstract
This companion paper is provided to describe the major experiments reported in our paper "On Learning Disentangled Representation for Acoustic Event Detection" published in ACM Multimedia 2019. To make the replication of our work easier, we first give an introduction of the computing environment where all of our experiments are conducted. Furthermore, we provide an environmental configuration file to setup the compiling environment and other artifacts including the source code, datasets and the files generated during our experiments. Finally, we summarize the structure and usage of the source code. For more details, please consult the README file in the archive of artifacts on GitHub: https://github.com/mastergofujs/SED_PyTorch.
Lijian Gao, Qirong Mao, Ming Dong 0001, Ratna Babu Chinnam, Lucile Sassatelli, Miguel Fabián Romero Rondón, Ujjwal Sharma 0001
ACM Multimedia4
2021 Deep Relational Reasoning for the Prediction of Language Impairment and Postoperative Seizure Outcome Using Preoperative DWI Connectome Data of Children With Focal Epilepsy
abstract
Prolonged seizures in children with focal epilepsy (FE) may impair language functions and often reoccur after surgical intervention. This study is aimed at developing a novel deep relational reasoning network to investigate whether conventional diffusion-weighted imaging connectome analysis can be improved when predicting expressive and receptive scores of preoperative language impairments and classifying postoperative seizure outcomes (seizure freedom or recurrence) in individual FE children. To deeply reason the dependencies of axonal connections that are sparsely distributed in the whole brain, this study proposes the "dilated CNN + RN", a dilated convolutional neural network (CNN) combined with a relation network (RN). The performance of the dilated CNN + RN was evaluated using whole brain connectome data from 51 FE children. It was found that when compared with other state-of-the-art algorithms, the dilated CNN + RN led to an average improvement of 90.2% and 97.3% in predicting expressive and receptive language scores, and 2.2% and 4% improvement in classifying seizure freedom and seizure recurrence, respectively. These improvements were independent of the prefixed connectome densities. Also, the dilated CNN + RN could provide an explainable artificial intelligence (AI) model by computing gradient-based regression/classification activation maps. This mapping analysis revealed left superior-medial frontal cortex, bilateral hippocampi, and cerebellum as crucial hubs, facilitating important connections that were most predictive of language function and seizure refractoriness after surgery.
Soumyanil Banerjee, Ming Dong 0001, Min-Hee Lee, Nolan O'Hara, Csaba Juhász, Eishi Asano, Jeong-Won Jeong
IEEE Trans. Medical Imaging2
2021 SPA-GAN: Spatial Attention GAN for Image-to-Image Translation
abstract
Image-to-image translation is to learn a mapping between images from a source domain and images from a target domain. In this paper, we introduce the attention mechanism directly to the generative adversarial network (GAN) architecture and propose a novel spatial attention GAN model (SPA-GAN) for image-to-image translation tasks. SPA-GAN computes the attention in its discriminator and use it to help the generator focus more on the most discriminative regions between the source and target domains, leading to more realistic output images. We also find it helpful to introduce an additional feature map loss in SPA-GAN training to preserve domain specific features during translation. Compared with existing attention-guided GAN models, SPA-GAN is a lightweight model that does not need additional attention networks or supervision. Qualitative and quantitative comparison against state-of-the-art methods on benchmark datasets demonstrates the superior performance of SPA-GAN.
Hajar Emami, Majid Moradi Aliabadi, Ming Dong 0001, Ratna Babu Chinnam
IEEE Trans. Multim.3
2019 Prioritization of Multi-Level Risk Factors for Obesity
abstract
Obesity has become a significant threat to health. Identifying and understanding the underlying obesity risk factors (ORFs) are crucial for optimizing prevention, intervention and treatment for obesity. Most existing methodological approaches to risk factor analysis are employed within the single task learning (STL) framework to learn a ranked list of ORFs for a whole population. However, obesity is a multi-faced health outcome. Some ORFs are highly specific to a certain subpopulation and others are universal to the entire population. Multi-task learning (MTL) framework offers a solution to connect multiple related tasks. Within the MTL framework, we implement two tailor-made models, i.e., multi-task feature learning (MTFL) and clustered multi-task learning (CMTL), to conduct ORFs analysis. The former is capable of finding the universal ORFs for all subpopulations without sacrificing the uniqueness of each subpopulation. The latter uncovers the grouping structure and conducts multi-level ORFs analysis simultaneously. Experiments on a public behavioral dataset demonstrate a superior performance of our methods in prioritizing multi-level ORFs.
Lu Wang 0004, Ming Dong 0001, Elizabeth Towner, Dongxiao Zhu
BIBM2
2019 On Learning Disentangled Representation for Acoustic Event Detection
abstract
Polyphonic Acoustic Event Detection (AED) is a challenging task as the sounds are mixed with the signals from different events, and the features extracted from the mixture do not match well with features calculated from sounds in isolation, leading to suboptimal AED performance. In this paper, we propose a supervised β-VAE model for AED, which adds a novel event-specific disentangling loss in the objective function of disentangled learning. By incorporating either latent factor blocks or latent attention in disentangling, supervised β-VAE learns a set of discriminative features for each event. Extensive experiments on benchmark datasets show that our approach outperforms the current state-of-the-arts (top-1 performers in the Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 AED challenge). Supervised β-VAE has great success in challenging AED tasks with a large variety of events and imbalanced data.
Lijian Gao, Qirong Mao, Ming Dong 0001, Yu Jing, Ratna Babu Chinnam
ACM Multimedia3
2019 Triple attention network for sentimental visual question answering
Nelson Ruwa, Qirong Mao, Heping Song, Hongjie Jia, Ming Dong 0001
Comput. Vis. Image Underst.5
2019 Mood-aware visual question answering
Nelson Ruwa, Qirong Mao, Liangjun Wang, Jianping Gou, Ming Dong 0001
Neurocomputing5
2019 Objective Detection of Eloquent Axonal Pathways to Minimize Postoperative Deficits in Pediatric Epilepsy Surgery Using Diffusion Tractography and Convolutional Neural Networks
abstract
Convolutional neural networks (CNNs) have recently been used in biomedical imaging applications with great success. In this paper, we investigated the classi?cation performance of CNN models on diffusion weighted imaging (DWI) streamlines de?ned by functional MRI (fMRI) and electrical stimulation mapping (ESM). To learn a set of discriminative and interpretable features from the extremely unbalanced dataset, we evaluated different CNN architectures with multiple loss functions (e.g., focal loss and center loss) and a soft attention mechanism, and compared our models with current state-ofthe-art methods. Through extensive experiments on streamlines collected from 70 healthy children and 70 children with focal epilepsy, we demonstrated that our deep CNN model with focal and central losses and soft attention outperforms all existing models in the literature and provides clinically acceptable accuracy (73 -100%) for the objective detection of functionally-important white matter pathways including ESM determined eloquent areas such as primary motor, aphasia, speech arrest, auditory, and visual functions. The ?ndings of this study encourage further investigations to determine if DWICNN analysis can serve as a noninvasive diagnostic tool during pediatric presurgical planning by estimating not only the location of essential cortices at the gyral level, but also the underlying ?bers connecting these cortical areas, to minimize or predict postsurgical functional de?cits. This study translates an advanced CNN model to clinical practice in the pediatric population where currently available approaches (e.g., ESM, fMRI) are suboptimal. The implementation will be released at https://github. com/HaotianMXu/Brain-?ber-classi?cation-using-CNNs.
Ming Dong 0001, Min-Hee Lee, Noel O'Hara, Eishi Asano, Jeong-Won Jeong
IEEE Trans. Medical Imaging2
2018 Coupled End-to-End Transfer Learning With Generalized Fisher Information
abstract
In transfer learning, one seeks to transfer related information from source tasks with sufficient data to help with the learning of target task with only limited data. In this paper, we propose a novel Coupled End-to-end Transfer Learning (CETL) framework, which mainly consists of two convolutional neural networks (source and target) that connect to a shared decoder. A novel loss function, the coupled loss, is used for CETL training. From a theoretical perspective, we demonstrate the rationale of the coupled loss by establishing a learning bound for CETL. Moreover, we introduce the generalized Fisher information to improve multi-task optimization in CETL. From a practical aspect, CETL provides a unified and highly flexible solution for various learning tasks such as domain adaption and knowledge distillation. Empirical result shows the superior performance of CETL on cross-domain and cross-task image classification.
Shixing Chen, Caojin Zhang, Ming Dong 0001
CVPR3
2018 Cascaded Multi-level Transformed Dirichlet Process for Multi-pose Facial Expression Recognition
abstract
As an essential way of human emotional behavior understanding, facial expression recognition (FER) has been studied extensively in recent years. However, the existing methods of FER are typically based on near-frontal face data. High-recognition accuracy for multi-pose FER continues to be a challenge. In this paper, we present a novel cascaded multi-level Transformed Dirichlet Process (cml-TDP) model for multi-pose FER. The top-level structure of the cml-TDP model has been carefully designed to make coarse-to-fine prediction, and the outputs of the model are fused for robust and accurate estimation at each level. There are three primary merits to cml-TDP. First, pose is explicitly introduced into cml-TDP so that separate training and parameter tuning for each pose is not required. Second, cml-TDP describes an image by its detected positions and appearance features to implicitly construct geometric constraints. Third, cml-TDP can learn an intermediate facial expression representation subject to geometric constraints. By sharing the pool of spatially coherent features over expressions and poses, we provide a scalable solution for multi-pose FER. The proposed model has been evaluated on two benchmark databases, BU-3DFE and RAFD, and achieved 79.33% and 75.00% FER accuracy on these two datasets, respectively, which has outperformed current state-of-the-art FER methods.
Qirong Mao, Feifei Zhang 0001, Liangjun Wang, Sidian Luo, Ming Dong 0001
Comput. J.5
2018 Multinomial classification with class-conditional overlapping sparse feature groups
Xiangrui Li, Dongxiao Zhu, Ming Dong 0001
Pattern Recognit. Lett.3
2018 Deep Age Estimation: From Classification to Ranking
abstract
Human age is considered an important biometric trait for human identification or search. Recent research shows that the aging features deeply learned from large-scale data lead to significant performance improvement on facial image-based age estimation. However, age-related ordinal information is totally ignored in these approaches. In this paper, we propose a novel convolutional neural network (CNN)-based framework, ranking-CNN, for age estimation. Ranking-CNN contains a set of basic CNNs, each of which is trained with ordinal age labels. Then, their binary outputs are aggregated for the final age prediction. From a theoretical perspective, we obtain an approximation for the final ranking error, show that it is controlled by the maximum error produced among subranking problems, and thus find a new error bound, which provides helpful guidance for the training and analysis of deep rankers. Based on the new error bound, we theoretically give an explicit formula for the learning of ranking-CNN and demonstrate its convergence using the stochastic approximation method. Moreover, we rigorously prove that ranking-CNN, by considering ordinal relation between ages, is more likely to get smaller estimation errors when compared with multiclass classification approaches. Through extensive experiments, we show that ranking-CNN outperforms other state-of-the-art feature extractors and age estimators on benchmark datasets.
Shixing Chen, Caojin Zhang, Ming Dong 0001
IEEE Trans. Multim.3
2018 Spatially Coherent Feature Learning for Pose-Invariant Facial Expression Recognition
abstract
Feature learning has enjoyed much attention and achieved good performance in recent studies of image processing. Unlike the required training conditions often assumed there, far less labeled data is available for training emotion classification systems. In addition, current feature learning is typically performed on an entire face image without considering the dependency between features. These approaches ignore the fact that faces are structured and the neighboring features are dependent. Thus, the learned features lack the power to describe visually coherent facial images. Our method is therefore designed with the goal of simplifying the problem domain by removing expression-irrelevant factors from the input images, with a key region-based mechanism, which is an effort to reduce the amount of data required to effectively train the feature-learning methods. Meanwhile, we can construct geometric constraints between the key regions and its detected positions. To this end, we introduce a Spatially Coherent featurelearning method for Pose-invariant Facial Expression Recognition (SC-PFER). In our model, we first perform face frontalization through a 3D pose-normalization technique, which could normalize poses while preserving the identity information through synthesizing frontal faces for facial images with arbitrary views. Subsequently, we select a sequence of key regions around 51 key points in the synthetic frontal face images for efficient unsupervised feature learning. Finally, we introduce a linkage structure over the learning-based features and the corresponding geometry information of each key region to encode the dependencies of the regions. Our method, on the whole, does not require training multiple models for each specific pose and avoids separating training and parameter tuning for each pose. The proposed framework has been evaluated on two benchmark databases, BU-3DFE and SFEW, for pose-invariant Facial Expression Recognition (FER). The experimental results demonstrate that our algorithm outperforms current state-of-the-art FER methods. Specifically, our model achieves an improvement of 1.72% and 1.11% FER accuracy, on average, on BU-3DFE and SFEW, respectively.
Feifei Zhang 0001, Qirong Mao, Xiangjun Shen, Yongzhao Zhan 0001, Ming Dong 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2017 Using Ranking-CNN for Age Estimation
abstract
Human age is considered an important biometric trait for human identification or search. Recent research shows that the aging features deeply learned from large-scale data lead to significant performance improvement on facial image-based age estimation. However, age-related ordinal information is totally ignored in these approaches. In this paper, we propose a novel Convolutional Neural Network (CNN)-based framework, ranking-CNN, for age estimation. Ranking-CNN contains a series of basic CNNs, each of which is trained with ordinal age labels. Then, their binary outputs are aggregated for the final age prediction. We theoretically obtain a much tighter error bound for ranking-based age estimation. Moreover, we rigorously prove that ranking-CNN is more likely to get smaller estimation errors when compared with multi-class classification approaches. Through extensive experiments, we show that statistically, ranking-CNN significantly outperforms other state-of-the-art age estimation models on benchmark datasets.
Shixing Chen, Caojin Zhang, Ming Dong 0001, Jialiang Le, Mike Rao
CVPR3
2017 Directionally Convolutional Networks for 3D Shape Segmentation
abstract
Previous approaches on 3D shape segmentation mostly rely on heuristic processing and hand-tuned geometric descriptors. In this paper, we propose a novel 3D shape representation learning approach, Directionally Convolutional Network (DCN), to solve the shape segmentation problem. DCN extends convolution operations from images to the surface mesh of 3D shapes. With DCN, we learn effective shape representations from raw geometric features, i.e., face normals and distances, to achieve robust segmentation. More specifically, a two-stream segmentation framework is proposed: one stream is made up by the proposed DCN with the face normals as the input, and the other stream is implemented by a neural network with the face distance histogram as the input. The learned shape representations from the two streams are fused by an element-wise product. Finally, Conditional Random Field (CRF) is applied to optimize the segmentation. Through extensive experiments conducted on benchmark datasets, we demonstrate that our approach outperforms the current state-of-the-arts (both classic and deep learning-based) on a large variety of 3D shapes.
Ming Dong 0001, Zichun Zhong
ICCV2
2017 Modeling Over-Dispersion for Network Data Clustering
abstract
Over-dispersed network data mining has emerged as a central theme in data science, evident by a sharp increase in the volume of real-world network data with imbalanced clusters.While most of existing clustering methods are designed for discovering the number of clusters and class specific connectivity patterns, few methods are available to uncover the imbalanced clusters,commonly existing in network communities and image segments.In this paper, we propose a generalized probabilistic modeling framework,SizeConnectivity, to estimate over-dispersed cluster size distribution together with class specific connectivity patterns from network data.We performed extensive synthetic and real-world experiments on clustering social network data and image data for detecting network communities and image segments.Our results demonstrate a superior performance of our SizeConnectivity clustering method in recovering the hidden structure of network data via modeling over-dispersion.
Lu Wang 0004, Dongxiao Zhu, Ming Dong 0001, Yan Li 0052
ICMLA3
2017 Object tracking via Dirichlet process-based appearance models
Raed Almomani, Ming Dong 0001, Dongxiao Zhu
Neural Comput. Appl.2
2017 Hierarchical Bayesian Theme Models for Multipose Facial Expression Recognition
abstract
As an essential way of human emotional behavior understanding, facial expression recognition (FER) has attracted a great deal of attention in multimedia research. Most of studies are conducted in a “lab-controlled” environment, and their real-world performance degenerates greatly due to factors such as head pose variations. In this paper, we propose a pose-based hierarchical Bayesian theme model to address challenging issues in multipose FER. Local appearance features and global geometry information are combined in our model to learn an intermediate face representation before recognizing expressions. By sharing a pool of features with various poses, our model provides a unified solution for multipose FER, bypassing the separate training and parameter tuning for each pose, and thus is scalable to a large number of poses. Experiments on both benchmark facial expression databases and Internet images show the superior/highly competitive performance of our system when compared with the current state of the art.
Qirong Mao, Qiyu Rao, Ming Dong 0001
IEEE Trans. Multim.4
2016 A Bayesian hierarchical appearance model for robust object tracking
abstract
In tracking, one of the major challenges comes from handling appearance variations caused by changes in scale, pose, illumination and occlusion. In this paper, we propose a novel Bayesian Hierarchical Appearance Model (BHAM) for robust object tracking. Our idea is to model the appearance of a target as a combination of multiple appearance models, each covering the target appearance changes under a given view angle. Specifically, target instances are modeled by Dirichlet Process and dynamically clustered based on their visual similarity. Thus, BHAM provides an infinite nonparametric mixture of distributions that can grow automatically with the complexity of the appearance data. We built an object tracking system by integrating BHAM with background subtraction and the KLT tracker. Our experimental results on real-world videos show that our system has superior performance when compared with several state-of-the-art trackers.
Raed Almomani, Ming Dong 0001, Dongxiao Zhu
ICME2
2016 Poisson-Markov Mixture Model and Parallel Algorithm for Binning Massive and Heterogenous DNA Sequencing Reads
Lu Wang 0004, Dongxiao Zhu, Yan Li 0052, Ming Dong 0001
ISBRA4
2016 Multi-pose Facial Expression Recognition Using Transformed Dirichlet Process
abstract
Driven by recent advances in human-centered computing, Facial Expression Recognition (FER) has attracted significant attention in many applications. In this paper, we propose a novel graphical model, multi-level Transformed Dirichlet Process (ml-TDP), for multi-pose FER. In our approach, pose is explicitly introduced into ml-TDP so that separate training and parameter tuning for each pose is not required. In addition, ml-TDP can learn an intermediate facial expression representation subject to geometric constraints. By sharing the pool of spatially-coherent features over expressions and poses, we provide a scalable solution for multi-pose FER. Extensive experimental result on benchmark facial expression databases shows the superior performance of ml-TDP.
Feifei Zhang 0001, Qirong Mao, Ming Dong 0001, Yongzhao Zhan 0001
ACM Multimedia3
2016 A study of the effectiveness of machine learning methods for classification of clinical interview fragments into a large number of categories
Alexander Kotov 0001, April Idalski Carcone, Ming Dong 0001, Sylvie Naar, Kathryn Brogan Hartlieb
J. Biomed. Informatics4
2015 Interpretable Probabilistic Latent Variable Models for Automatic Annotation of Clinical Text
Alexander Kotov 0001, April Idalski Carcone, Ming Dong 0001, Sylvie Naar, Kathryn Brogan Hartlieb
AMIA4
2015 Multi-level Approximate Spectral Clustering
abstract
Clustering is a task of finding natural groups in datasets based on measured or perceived similarity between data points. Spectral clustering is a well-known graph-theoretic approach, which is capable of capturing non-convex geometries of datasets. However, it generally becomes infeasible for analyzing large datasets due to relatively high time and space complexity. In this paper, we propose Multi-level Approximate Spectral (MAS) clustering to enable efficient analysis of large datasets. By integrating a series of low-rank matrix approximations (i.e., approximations to the affinity matrix and its subspace, as well as those for the Laplacian matrix and the Laplacian subspace), MAS achieves great computational and spacial efficiency. MAS provides a general framework for fast and accurate spectral clustering, which works with any kernels, various fast sampling strategies and different low-rank approximation algorithms. In addition, it can be easily extended for distributed computing. From a theoretical perspective, we provide rigorous analysis of its approximation error in addition to its correctness and computational complexity. Through extensive experiments we demonstrate superior performance of the proposed method relative to several well-known approximate spectral clustering algorithms.
Ming Dong 0001, Alexander Kotov 0001
ICDM2
2015 Exemplar-based low-rank matrix decomposition for data clustering
Ming Dong 0001
Data Min. Knowl. Discov.2
2014 Learning Good Features to Track
abstract
Object tracking is an important task within the field of computer vision. Tracking accuracy depends mainly on finding good discriminative features to estimate the target location. In this paper, we introduce online feature learning in tracking and propose to learn good features to track generic objects using online convolutional neural networks (OCNN). OCNN has two feature mapping layers that are trained offline based on unlabeled data. In tracking, the collected positive and negative samples from the previously tracked frames are used to learn good features for a specific target. OCNN is also augmented with a classifier to provide a decision. We build a tracking system by combining OCNN and a color-based multi-appearance model. Our experimental results on publicly available video datasets show that the tracking system has superior performance when compared with several state of-the-art trackers.
Raed Almomani, Ming Dong 0001
ICMLA2
2014 Detection of Abnormal Human Behavior Using a Matrix Approximation-Based Approach
abstract
Automatic detection of abnormal events is one of central tasks in video surveillance. In this paper we present a matrix approximation-based method to detect abnormal human behavior. In our model, a behavior pattern is represented by a motion matrix obtained through object tracking. We model typical motions associated with normal behaviors with a set of motion subspaces, computed through low-rank matrix approximation. Then, abnormal human behaviors are identified by the motion deviations from the representative subspaces. Our method does not require a complicated classification procedure, and can fast detect abnormal events in complex scenes. In addition, through the adaptive learning module, our model is built on the observed data, and can be expanded by incorporating new behavior patterns during the detection process. The results on simulated surveillance videos show the effectiveness of our method.
Ming Dong 0001
ICMLA2
2014 Speech Emotion Recognition Using CNN
abstract
Deep learning systems, such as Convolutional Neural Networks (CNNs), can infer a hierarchical representation of input data that facilitates categorization. In this paper, we propose to learn affect-salient features for Speech Emotion Recognition (SER) using semi-CNN. The training of semi-CNN has two stages. In the first stage, unlabeled samples are used to learn candidate features by contractive convolutional neural network with reconstruction penalization. The candidate features, in the second step, are used as the input to semi-CNN to learn affect-salient, discriminative features using a novel objective function that encourages the feature saliency, orthogonality and discrimination. Our experiment results on benchmark datasets show that our approach leads to stable and robust recognition performance in complex scenes (e.g., with speaker and environment distortion), and outperforms several well-established SER features.
Zhengwei Huang, Ming Dong 0001, Qirong Mao, Yongzhao Zhan 0001
ACM Multimedia2
2014 Learning Salient Features for Speech Emotion Recognition Using Convolutional Neural Networks
abstract
As an essential way of human emotional behavior understanding, speech emotion recognition (SER) has attracted a great deal of attention in human-centered signal processing. Accuracy in SER heavily depends on finding good affect- related , discriminative features. In this paper, we propose to learn affect-salient features for SER using convolutional neural networks (CNN). The training of CNN involves two stages. In the first stage, unlabeled samples are used to learn local invariant features (LIF) using a variant of sparse auto-encoder (SAE) with reconstruction penalization. In the second step, LIF is used as the input to a feature extractor, salient discriminative feature analysis (SDFA), to learn affect-salient, discriminative features using a novel objective function that encourages feature saliency, orthogonality, and discrimination for SER. Our experimental results on benchmark datasets show that our approach leads to stable and robust recognition performance in complex scenes (e.g., with speaker and language variation, and environment distortion) and outperforms several well-established SER features.
Qirong Mao, Ming Dong 0001, Zhengwei Huang, Yongzhao Zhan 0001
IEEE Trans. Multim.2
2013 SegTrack: A novel tracking system with improved object segmentation
abstract
Most tracking methods depend on a rectangle or an ellipse mask to segment and track objects. Typically, using a larger or smaller mask will lead to loss of tracked objects. In this paper, we propose an object tracking system (SegTrack) that deals with partial and full occlusions by employing improved segmentation methods. Our improved mixture of Gaussians segments foreground objects from the background and solves stop-then-move and move-then-stop problems. Then, the KLT tracker tracks objects in consecutive frames and detects partial and full occlusions. In partial occlusion, a novel silhouette segmentation algorithm evolves the silhouettes of occluded objects by matching the location and appearance of occluded objects between successive frames. In full occlusion, one or more feature vectors for each tracked object are used to re-identify the object after reappearing. Our experimental results show that SegTrack provides more accurate and robust tracking when compared to other state-of-the-art trackers.
Raed Almomani, Ming Dong 0001
ICIP2
2012 Multi-instance rendering based on dynamic differential surface propagation
abstract
High-quality rendering of a complex system usually depends on accurate segmentation of the corresponding objects. In medical imaging data, it is difficult to extract and visualize multiple related objects simultaneously due to the noise, shape variance, and resolution difference of the objects. In this paper we present a viable method to adaptively extract the isosurfaces of different yet informatively related tissues simultaneously from the multimodality imaging data of the brain based on the statistical Partial Differential Equation (PDE) deformable models. Our system and experiments demonstrate the power of using explicit PDE models in extracting and rendering of multiple objects.
Zhaoqiang Lai, Ming Dong 0001, Jing Hua 0001
ICIP3
2012 Real-time detection of abnormal crowd behavior using a matrix approximation-based approach
abstract
Automatic detection of abnormal crowd activities is one of central tasks in video surveillance. In this paper we present a matrix approximation-based method to detect abnormal crowd behavior. In our approach, we model typical motions associated with normal crowd behaviors with a set of motion subspaces, computed through low-rank matrix approximation. Then, abnormal crowd behaviors are identified by the motion deviations from the representative subspaces. Our method does not require complicated tracking or classification method, and can fast detect abnormal events in complex crowd scenes. In addition, through the adaptive learning module, our model is built on the observed data, and can be expanded by incorporating new crowd behavior patterns during the detection process. The results on simulated crowd scenes show the effectiveness of our method.
Ming Dong 0001
ICIP2
2012 Multi-level Low-rank Approximation-based Spectral Clustering for image segmentation
Ming Dong 0001
Pattern Recognit. Lett.2
2012 Low-Rank Kernel Matrix Factorization for Large-Scale Evolutionary Clustering
abstract
Traditional clustering techniques are inapplicable to problems where the relationships between data points evolve over time. Not only is it important for the clustering algorithm to adapt to the recent changes in the evolving data, but it also needs to take the historical relationship between the data points into consideration. In this paper, we propose ECKF, a general framework for evolutionary clustering large-scale data based on low-rank kernel matrix factorization. To the best of our knowledge, this is the first work that clusters large evolutionary data sets by the amalgamation of low-rank matrix approximation methods and matrix factorization-based clustering. Since the low-rank approximation provides a compact representation of the original matrix, and especially, the near-optimal low-rank approximation can preserve the sparsity of the original data, ECKF gains computational efficiency and hence is applicable to large evolutionary data sets. Moreover, matrix factorization-based methods have been shown to effectively cluster high-dimensional data in text mining and multimedia data analysis. From a theoretical standpoint, we mathematically prove the convergence and correctness of ECKF, and provide detailed analysis of its computational efficiency (both time and space). Through extensive experiments performed on synthetic and real data sets, we show that ECKF outperforms the existing methods in evolutionary clustering.
Manjeet Rege, Ming Dong 0001, Yongsheng Ding
IEEE Trans. Knowl. Data Eng.3
2011 On the clustering of large-scale data: A matrix-based approach
abstract
Nowadays, the analysis of large amounts of digital documents become a hot research topic since the libraries and database are converted electronically, such as PUBMED and IEEE publications. The ubiquitous phenomenon of massive data and sparse information imposes considerable challenges in data mining research. In this paper, we propose a theoretical framework, Exemplar-based Low-rank sparse Matrix Decomposition (ELMD), to cluster large-scale datasets. Specifically, given a data matrix, ELMD first computes a representative data subspace and a near-optimal low-rank approximation. Then, the cluster centroids and indicators are obtained through matrix decomposition, in which we require that the cluster centroids lie within the representative data subspace. From a theoretical perspective, we show the correctness and convergence of the ELMD algorithm, and provide detailed analysis on its efficiency. Through extensive experiments performed on both synthetic and real datasets, we demonstrate the superior performance of ELMD for clustering large-scale data.
Ming Dong 0001
IJCNN2
2010 Non-Negative Matrix Factorization for Semisupervised Heterogeneous Data Coclustering
abstract
Coclustering heterogeneous data has attracted extensive attention recently due to its high impact on various important applications, such us text mining, image retrieval, and bioinformatics. However, data coclustering without any prior knowledge or background information is still a challenging problem. In this paper, we propose a Semisupervised Non-negative Matrix Factorization (SS-NMF) framework for data coclustering. Specifically, our method computes new relational matrices by incorporating user provided constraints through simultaneous distance metric learning and modality selection. Using an iterative algorithm, we then perform trifactorizations of the new matrices to infer the clusters of different data types and their correspondence. Theoretically, we prove the convergence and correctness of SS-NMF coclustering and show the relationship between SS-NMF with other well-known coclustering models. Through extensive experiments conducted on publicly available text, gene expression, and image data sets, we demonstrate the superior performance of SS-NMF for heterogeneous data coclustering.
Ming Dong 0001
IEEE Trans. Knowl. Data Eng.3
2009 Distributed Faulty Sensor Detection in Sensor Networks
Xuanwen Luo, Ming Dong 0001
ICANN (2)2
2009 Image co-clustering with multi-modality features and user feedbacks
abstract
In Content-based Image Retrieval (CBIR) research, advanced technology that fuses the heterogeneous information into image clustering has drawn extensive attention recently. However, using multiple features for co-clustering images without any user feedbacks is a challenging problem. In this paper, we propose a Semi-Supervised Non-negative Matrix Factorization (SS-NMF) framework for image co-clustering. Our method computes new relational matrices by incorporating user provided feedbacks into images through simultaneous distance metric learning and feature selection for different low-level visual features. Using an iterative algorithm, we perform tri-factorizations of the new matrices to infer image clusters. Theoretically, we show the convergence and correctness of SS-NMF co-clustering and the advantages of SS-NMF co-clustering over existing approaches. Through extensive experiments conducted on image data sets, we demonstrate that SS-NMF provides an effective and efficient solution for image co-clustering.
Ming Dong 0001
ACM Multimedia2
2009 Semi-supervised Document Clustering with Simultaneous Text Representation and Categorization
Ming Dong 0001
ECML/PKDD (1)3
2009 Simultaneous Localized Feature Selection and Model Detection for Gaussian Mixtures
abstract
In this paper, we propose a novel approach of simultaneous localized feature selection and model detection for unsupervised learning. In our approach, local feature saliency, together with other parameters of Gaussian mixtures, are estimated by Bayesian variational learning. Experiments performed on both synthetic and real-world data sets demonstrate that our approach is superior over both global feature selection and subspace clustering methods.
Yuanhong Li, Ming Dong 0001, Jing Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Exemplar-based Visualization of Large Document Corpus (InfoVis2009-1115)
abstract
With the rapid growth of the World Wide Web and electronic information services,text corpus is becoming available on-line at an incredible rate.By displaying text data in a logical layout (e.g., color graphs),text visualization presents a direct way to observe the documents as well as understand the relationship between them.In this paper, we propose a novel technique, Exemplar-based Visualization (EV), to visualize an extremely large text corpus. Capitalizing on recent advances in matrix approximation and decomposition, EV presents a probabilistic multidimensional projection model in the low-rank text subspace with a sound objective function. The probability of each document proportion to the topics is obtained through iterative optimization and embedded to a low dimensional space using parameter embedding.By selecting the representative exemplars, we obtain a compact approximation of the data. This makes the visualization highly efficient and flexible. In addition, the selected exemplars neatly summarize the entire data set and greatly reduce the cognitive overload in the visualization, leading to an easier interpretation of large text corpus. Empirically, we demonstrate the superior performance of EV through extensive experiments performed on the publicly available text data sets.
Ming Dong 0001, Jing Hua 0001
IEEE Trans. Vis. Comput. Graph.3
2009 Intrinsic Geometric Scale Space by Shape Diffusion
abstract
This paper formalizes a novel, intrinsic geometric scale space (IGSS) of 3D surface shapes. The intrinsic geometry of a surface is diffused by means of the Ricci flow for the generation of a geometric scale space. We rigorously prove that this multiscale shape representation satisfies the axiomatic causality property. Within the theoretical framework, we further present a feature-based shape representation derived from IGSS processing, which is shown to be theoretically plausible and practically effective. By integrating the concept of scale-dependent saliency into the shape description, this representation is not only highly descriptive of the local structures, but also exhibits several desired characteristics of global shape representations, such as being compact, robust to noise and computationally efficient. We demonstrate the capabilities of our approach through salient geometric feature detection and highly discriminative matching of 3D scans.
Guangyu Zou, Jing Hua 0001, Zhaoqiang Lai, Xianfeng Gu, Ming Dong 0001
IEEE Trans. Vis. Comput. Graph.5
2008 A matrix-based approach for semi-supervised document co-clustering
abstract
In order to derive high quality information from text, the field of text mining has advanced swiftly from simple document clustering to co-clustering documents and words. However, document co-clustering without any prior knowledge or background information is a challenging problem. In this paper, we propose a Semi-Supervised Non-negative Matrix Factorization (SS-NMF) based framework for document co-clustering. Our method computes a new word-document matrix by incorporating user provided constraints through distance metric learning. Using an iterative algorithm, we perform tri-factorization of the new matrix to infer the document and word clusters. Through extensive experiments conducted on publicly available data sets, we demonstrate the superior performance of SS-NMF for document co-clustering.
Ming Dong 0001
CIKM3
2008 Salient region detection and feature extraction in 3D visual data
abstract
Saliency detection and local feature extraction for 2D images have received extensive attention recently. In this paper, we propose saliency detection and feature extraction techniques for 3D visual data. Our algorithm directly works in 3D scale space and detects interesting regions in different scales. We then extract a local descriptor based on gradient location-orientation histogram which is invariant to scale and rotation of the 3D object. The proposed methodology has been tested on 3D synthetic and Magnetic Resonance Imaging (MRI) data sets. The performance of the algorithm is evaluated based on the repeatability of saliency detection and descriptor matching, after 3D transformation and in the presence of noise.
Ming Dong 0001
ICIP1
2008 Localized feature selection for Gaussian mixtures using variational learning
abstract
Typical unsupervised feature selection algorithms select a common feature subset for all the clusters. Consequently, clusters embedded in different feature subspaces are not discovered. In this paper, we propose a novel approach of simultaneous localized feature selection and model detection for unsupervised learning. In our approach, local feature saliency, together with other parameters of Gaussian mixtures, are estimated by Bayesian variational learning. Experiments performed on real-world datasets illustrate that our approach is superior over both global feature selection and subspace clustering methods.
Yuanhong Li, Ming Dong 0001, Yunqian Ma
ICPR2
2008 Feature selection for clustering with constraints using Jensen-Shannon divergence
abstract
In semi-supervised clustering, domain knowledge can be converted to constraints and used to guide the clustering. In this paper we propose a feature selection algorithm for semi-supervised clustering. In our method, features are conditionally independent. Feature saliency is first computed in unsupervised clustering using the expectation maximization model. Then, it is refined in the tuning step to minimize the feature-wise constraint violation measure, calculated based on the Jensen-Shannon divergence. Experimental results show that a small amount of supervision can improve the performance of clustering and feature selection.
Yuanhong Li, Ming Dong 0001, Yunqian Ma
ICPR2
2008 Graph theoretical framework for simultaneously integrating visual and textual features for efficient web image clustering
abstract
With the explosive growth of Web and the recent development in digital media technology, the number of images on the Web has grown tremendously. Consequently, Web image clustering has emerged as an important application. Some of the initial efforts along this direction revolved around clustering Web images based on the visual features of images or textual features by making use of the text surrounding the images. However, not much work has been done in using multimodal information for clustering Web images. In this paper, we propose a graph theoretical framework for simultaneously integrating visual and textual features for efficient Web image clustering. Specifically, we model visual features, images and words from surrounding text using a tripartite graph. Partitioning this graph leads to clustering of the Web images. Although, graph partitioning approach has been adopted before, the main contribution of this work lies in a new algorithm that we propose- Consistent Isoperimetric High-order Co-clustering (CIHC), for partitioning the tripartite graph. Computationally, CIHC is very quick as it requires a simple solution to a sparse system of linear equations. Our theoretical analysis and extensive experiments performed on real Web images demonstrate the performance of CIHC in terms of the quality, efficiency and scalability in partitioning the visual feature-image-word tripartite graph.
Manjeet Rege, Ming Dong 0001, Jing Hua 0001
WWW2
2008 Knowledge discovery in corporate events by neural network rule extraction
Ming Dong 0001, Xu-Shen Zhou
Appl. Intell.1
2008 Bipartite isoperimetric graph partitioning for data co-clustering
Manjeet Rege, Ming Dong 0001, Farshad Fotouhi
Data Min. Knowl. Discov.2
2008 Surface matching with salient keypoints in geodesic scale space
abstract
Abstract This paper develops a new salient keypoints‐based shape description which extracts the salient surface keypoints with detected scales. Salient geometric features can then be defined collectively on all the detected scale normalized local patches to form a shape descriptor for surface matching purpose. The saliency‐driven keypoints are computed as local extrema of the difference of Gaussian function defined over a curved surface in geodesic scale space. This method can properly function on either manifold or non‐manifold surface without resorting to any surface mapping or parameterization procedures. Therefore, it has a wide utility in many applications such as shape matching, classification, and recognition. Our experiments on 3D shapes demonstrate that the salient keypoints and local feature descriptors are robust and stable to noisy input and insensitive to resolution change. We have applied our technique to the tasks of 3D shape matching, and the experimental results showed good performance and the effectiveness of this new method. Copyright © 2008 John Wiley & Sons, Ltd.
Guangyu Zou, Jing Hua 0001, Ming Dong 0001, Hong Qin 0001
Comput. Animat. Virtual Worlds3
2008 Non-negative matrix factorization for semi-supervised data clustering
Manjeet Rege, Ming Dong 0001, Jing Hua 0001
Knowl. Inf. Syst.3
2008 Localized feature selection for clustering
Yuanhong Li, Ming Dong 0001, Jing Hua 0001
Pattern Recognit. Lett.2
2008 Geodesic Distance-weighted Shape Vector Image Diffusion
abstract
This paper presents a novel and efficient surface matching and visualization framework through the geodesic distance-weighted shape vector image diffusion. Based on conformal geometry, our approach can uniquely map a 3D surface to a canonical rectangular domain and encode the shape characteristics (e.g., mean curvatures and conformal factors) of the surface in the 2D domain to construct a geodesic distance-weighted shape vector image, where the distances between sampling pixels are not uniform but the actual geodesic distances on the manifold. Through the novel geodesic distance-weighted shape vector image diffusion presented in this paper, we can create a multiscale diffusion space, in which the cross-scale extrema can be detected as the robust geometric features for the matching and registration of surfaces. Therefore, statistical analysis and visualization of surface properties across subjects become readily available. The experiments on scanned surface models show that our method is very robust for feature extraction and surface matching even under noise and resolution change. We have also applied the framework on the real 3D human neocortical surfaces, and demonstrated the excellent performance of our approach in statistical analysis and integrated visualization of the multimodality volumetric data over the shape vector image.
Jing Hua 0001, Zhaoqiang Lai, Ming Dong 0001, Xianfeng Gu, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.3
2007 Incorporating User Provided Constraints into Document Clustering
abstract
Document clustering without any prior knowledge or background information is a challenging problem. In this paper, we propose SS-NMF: a semi-supervised non- negative matrix factorization framework for document clustering. In SS-NMF, users are able to provide supervision for document clustering in terms of pairwise constraints on a few documents specifying whether they "must" or "cannot" be clustered together. Through an iterative algorithm, we perform symmetric tri-factorization of the document- document similarity matrix to infer the document clusters. Theoretically, we show that SS-NMF provides a general framework for semi-supervised clustering and that existing approaches can be considered as special cases of SS-NMF. Through extensive experiments conducted on publicly available data sets, we demonstrate the superior performance of SS-NMF for clustering documents.
Manjeet Rege, Ming Dong 0001, Jing Hua 0001
ICDM3
2007 Localized Feature Selection for Clustering and its Application in Image Grouping
abstract
In clustering, global feature selection algorithms attempt to select a common feature subset that is relevant for all clusters. Consequently, they are not able to identify individual clusters that exist in different feature subspaces. In this paper, we propose a localized feature selection algorithm for clustering. The proposed algorithm computes adjusted and normalized scatter separability for individual clusters. A sequential backward search is then applied to find the optimal (maybe local) feature subsets for each cluster. Experiment results on both synthetic data clustering and content-based image grouping show the need for feature selection in clustering and the benefits of selecting features locally.
Yuanhong Li, Ming Dong 0001, Jing Hua 0001
ICME2
2007 Gene Expression Clustering: a Novel Graph Partitioning Approach
abstract
In order to help understand how the genes are affected by different disease conditions in a biological system, clustering is typically performed to analyze gene expression data. In this paper, we propose to solve the clustering problem using a graph theoretical approach, and apply a novel graph partitioning model -isoperimetric graph partitioning (IGP), to group biological samples from gene expression data. The IGP algorithm has several advantages compared to the well-established spectral graph partitioning (SGP) model. First, IGP requires a simple solution to a sparse system of linear equations instead of the eigen-problem in the SGP model. Second, IGP avoids degenerate cases produced by spectral approach to achieve a partition with higher accuracy. Moreover, we integrate unsupervised gene selection into the proposed approach through two-way ordering of gene expression data, such that we can eliminate irrelevant or redundant genes in the data and obtain an improved clustering result. We evaluate our approach on several well-known problems involving gene expression profiles of colon cancer and leukemia subtypes. Our experiment results demonstrate that IGP constantly outperforms SGP and produces a better result that is closer to the original labeling of sample sets provided by domain experts. Furthermore, the clustering accuracy is improved significantly when IGP is integrated with the unsupervised gene (feature) selection.
Ming Dong 0001, Manjeet Rege
IJCNN2
2007 Deriving semantics for image clustering from accumulated user feedbacks
abstract
Image clustering solely based on visual features without any knowledge or background information suffers from the problem of semantic gap. In this paper, we propose SS-NMF: a semi-supervised non-negative matrix factorization framework for image clustering. Accumulated relevance feedback in a CBIR system is treated as user provided supervision for guiding the image clustering. We consider the set of positive images in the feedback as constraints on the clustering specifying that the images "must" be clustered together. Similarly, negative images provide constraints specifying that they "cannot" be clustered along with the positive images. Through an iterative algorithm, we perform symmetric tri-factorization of the image-image similarity matrix to infer the clustering. Theoretically, we prove the correctness of SS-NMF by showing that the algorithm is guaranteed to converge. Through experiments conducted on general purpose image datasets, we demonstrate the superior performance of SS-NMF for clustering images effectively.
Manjeet Rege, Ming Dong 0001, Farshad Fotouhi
ACM Multimedia3
2007 Clustering web images with multi-modal features
abstract
Web image clustering has drawn significant attention in the research community recently. However, not much work has been done in using multi-modal information for clustering Web images. In this paper, we address the problem of Web image clustering by simultaneous integration of visual and textual features from a graph partitioning perspective. In particular, we modelled visual features, images, and words from the surrounding text of the images using a tripartite graph. This graph is actually considered as a fusion of two bipartite graphs that are partitioned simultaneously by the proposed Consistent Isoperimetric High-order Co-clustering(CIHC) framework. Although a similar approach has been adopted before, the main contribution of this work lies in the computational efficiency, quality in Web image clustering and scalability to large image repositories that CIHC is able to achieve. We demonstrate this through experimental results performed on real Web images.
Manjeet Rege, Ming Dong 0001, Jing Hua 0001
ACM Multimedia2
2007 Integrative Information Visualization of Multimodality Neuroimaging Data
abstract
This paper presents a novel integrative information visualization framework for cross-subject neuroimaging data analysis. The framework can integrate multimodal information captured by different imaging modalities and population-based statistical information presented by different subjects. In this framework, accurate registration of cortical structures is the foundation for the information integration across population. We present a non-rigid intersubject brain surface registration method using conformal structure and spherical thin-plate splines. Spherical thin-plate splines are designed to explicitly match prominent homologous landmarks, and meanwhile, interpolate a global deformation field on the spherical domain, registering brain surfaces in a transformed space. Subsequently, an approach for the integrative information fusion and visualization is presented to handle multimodality neuroimaging data. The entire framework demonstrates its usefulness in multimodality neuroimaging data analysis across subjects.
Guangyu Zou, Jing Hua 0001, Ming Dong 0001
PG3
2007 Building a user-centered semantic hierarchy in image databases
Manjeet Rege, Ming Dong 0001, Farshad Fotouhi
Multim. Syst.2
2007 3D reconstruction from 2D images with hierarchical continuous simplices
Yunhao Tan, Jing Hua 0001, Ming Dong 0001
Vis. Comput.3
2006 Region-based Image Annotation using Asymmetrical Support Vector Machine-based Multiple-Instance Learning
abstract
In region-based image annotation, keywords are usually associated with images instead of individual regions in the training data set. This poses a major challenge for any learning strategy. In this paper, we formulate image annotation as a supervised learning problem under Multiple-Instance Learning (MIL) framework. We present a novel Asymmetrical Support Vector Machine-based MIL algorithm (ASVM-MIL), which extends the conventional Support Vector Machine (SVM) to the MIL setting by introducing asymmetrical loss functions for false positives and false negatives. The proposed ASVM-MIL algorithm is evaluated on both image annotation data sets and the benchmark MUSK data sets.
Changbo Yang, Ming Dong 0001, Jing Hua 0001
CVPR (2)2
2006 Co-clustering Documents and Words Using Bipartite Isoperimetric Graph Partitioning
abstract
In this paper, we present a novel graph theoretic approach to the problem of document-word co-clustering. In our approach, documents and words are modeled as the two vertices of a bipartite graph. We then propose isoperimetric co-clustering algorithm (ICA) - a new method for partitioning the document-word bipartite graph. ICA requires a simple solution to a sparse system of linear equations instead of the eigenvalue or SVD problem in the popular spectral co-clustering approach. Our extensive experiments performed on publicly available datasets demonstrate the advantages of ICA over spectral approach in terms of the quality, efficiency and stability in partitioning the document-word bipartite graph.
Manjeet Rege, Ming Dong 0001, Farshad Fotouhi
ICDM2
2006 Co-Clustering Image Features and Semantic Concepts
abstract
In this paper, we present a novel idea of co-clustering image features and semantic concepts. We accomplish this by modelling user feedback logs and low-level features using a bipartite graph. Our experiments demonstrate that (1) incorporating semantic information achieves better image clustering and (2) feature selection in co-clustering narrows the semantic gap, thus enabling efficient image retrieval.
Manjeet Rege, Ming Dong 0001, Farshad Fotouhi
ICIP2
2006 Finding a Semantic Structure Interactively in Image Databases
abstract
We present a new approach to organize an image database by finding a semantic structure interactively based on multi-user relevance feedback. By treating user relevance feedbacks as weak classifiers and combining them together, we are able to capture the categories in the users' mind and build a semantic structure in the image database. Experiments performed on an image database consisting of general purpose images demonstrate that our system outperforms some of the other conventional methods
Manjeet Rege, Ming Dong 0001, Farshad Fotouhi
ICME2
2006 Localized Support Vector Machines for Classification
abstract
Support vector machines (SVMs) have been promising methods in pattern recognition because of their solid mathematical foundation. In this paper, we propose a localized SVM classification scheme (LSVM). In which we first cluster the training data in each category, and then train a set of SVMs based on these dusters. The SVMs trained from the clusters in each category that are nearest to the given input pattern are then selected for the final classification. Our experiments on six UCI datasets show that LSVM outperforms the traditional SVM.
Ming Dong 0001
IJCNN1
2006 S-IRAS: An Interactive Semantic Image Retrieval and Annotation System
abstract
Relevance feedback and semantic retrieval have received extensive attention recently in the computer vision community. In this article, we present a semantic image query system with integrated feedback mechanism. Our system has two major components: the low-level feature space and the semantic space. In the low-level feature space, images are described by multidimensional vectors and are clustered based on the similarity of their contents. In the semantic space, the relationship among keywords is captured by a semantic hierarchy built by the aid of WordNet. Based on our system architecture, we propose a novel feedback solution for semantic retrieval called semantic feedback, which allows our system to interact with users directly at the semantic level. The short-term and long-term learning process of the semantic feedback substantially improves the image retrieval and annotation performance of our system. We demonstrate the effectiveness of our approach with experiments using 5,000 images from Corel database. The significant contribution of this article is in the scenario of having a relatively small training data set compared to the testing data set. Some of the previous work in the same direction has chosen a very large training set and a very small testing set. Clearly, the problem that we try to solve is more realistic and challenging.
Changbo Yang, Ming Dong 0001, Farshad Fotouhi
Int. J. Semantic Web Inf. Syst.2
2006 On Distributed Fault-Tolerant Detection in Wireless Sensor Networks
abstract
In this paper, we consider two important problems for distributed fault-tolerant detection in wireless sensor networks: 1) how to address both the noise-related measurement error and sensor fault simultaneously in fault-tolerant detection and 2) how to choose a proper neighborhood size n for a sensor node in fault correction such that the energy could be conserved. We propose a fault-tolerant detection scheme that explicitly introduces the sensor fault probability into the optimal event detection process. We mathematically show that the optimal detection error decreases exponentially with the increase of the neighborhood size. Experiments with both Bayesian and Neyman-Pearson approaches in simulated sensor networks demonstrate that the proposed algorithm is able to achieve better detection and better balance between detection accuracy and energy usage. Our work makes it possible to perform energy-efficient fault-tolerant detection in a wireless sensor network.
Xuanwen Luo, Ming Dong 0001, Yinlun Huang
IEEE Trans. Computers2
2005 ARGDYP: an Adaptive Region Growing and DYnamic Programming Algorithm for Stenosis Detection in MRI
abstract
In this paper, a novel image analysis algorithm adaptive region growing and dynamic programming (ARGDYP) is proposed for stenosis detection in magnetic resonance (MR) images. ARGDYP combines an adaptive region growing method for 3D vessel tracking and a dynamic programming approach for 2D vessel boundary detection. Our experiments based on both real and simulated MR data show that the proposed algorithm is able to accurately measure the cross sectional area of a blood vessel and thus determine whether or not there is a stenosis.
Jing Jiang 0014, Ming Dong 0001, E. Mark Haacke
ICASSP (2)2
2005 Image content annotation using Bayesian framework and complement components analysis
abstract
In this paper, we consider image annotation as a problem of image classification, in which each keyword is treated as a distinct class label. We then build a Bayesian model to solve the classification problem. To preserve the in-variation in the training data and reduce the noises, we also propose to estimate the class conditional probabilities in the feature subspace constructed by complement components analysis (CCA). We demonstrate the effectiveness of our approach through experiments in terms of annotation precision and recall.
Changbo Yang, Ming Dong 0001, Farshad Fotouhi
ICIP (1)2
2005 XML path based relevance model for automatic image annotation
abstract
This is the first paper that proposes automatic image annotation using the semantics of XML. In this paper, we propose XPRM-XML path based relevance model for automatic image annotation. Our experimental results show that the proposed model has considerable advantage over single word annotations in performing automatic semantic annotation.
Manjeet Rege, Ming Dong 0001, Farshad Fotouhi
ICME2
2005 I2A: an interactive image annotation system
abstract
In this paper, we propose an interactive image annotation system. The proposed system has two connected components, the low-level feature space and the semantic space. Experiments show that our system is able to provide accurate annotation for images by learning the connections between the two spaces through statistical modelling, natural language processing, and users' interaction.
Changbo Yang, Ming Dong 0001, Farshad Fotouhi
ICME2
2005 Semantic feedback for interactive image retrieval
abstract
In this paper we present a semantic image retrieval system with integrated feedback mechanism. In our system, we propose a novel feedback solution for semantic retrieval: semantic feedback, which allows our system to interact with users directly at the semantic level. The learning process of the semantic feedback substantially improves the image retrieval performance of the proposed system. We demonstrate the effectiveness of our approach with experiments using 5,000 images from Corel database.
Changbo Yang, Ming Dong 0001, Farshad Fotouhi
ACM Multimedia2
2005 Region based image annotation through multiple-instance learning
abstract
In an annotated image database, keywords are usually associated with images instead of individual regions, which poses a major challenge for any region based image annotation algorithm. In this paper, we propose to learn the correspondence between image regions and keywords through Multiple-Instance Learning (MIL). After a representative image region has been learned for a given keyword, we consider image annotation as a problem of image classification, in which each keyword is treated as a distinct class label. The classification problem is then addressed using the Bayesian framework. The proposed image annotation method is evaluated on an image database with 5,000 images.
Changbo Yang, Ming Dong 0001, Farshad Fotouhi
ACM Multimedia2
2005 Classifiability-based omnivariate decision trees
abstract
Top-down induction of decision trees is a simple and powerful method of pattern classification. In a decision tree, each node partitions the available patterns into two or more sets. New nodes are created to handle each of the resulting partitions and the process continues. A node is considered terminal if it satisfies some stopping criteria (for example, purity, i.e., all patterns at the node are from a single class). Decision trees may be univariate, linear multivariate, or nonlinear multivariate depending on whether a single attribute, a linear function of all the attributes, or a nonlinear function of all the attributes is used for the partitioning at each node of the decision tree. Though nonlinear multivariate decision trees are the most powerful, they are more susceptible to the risks of overfitting. In this paper, we propose to perform model selection at each decision node to build omnivariate decision trees. The model selection is done using a novel classifiability measure that captures the possible sources of misclassification with relative ease and is able to accurately reflect the complexity of the subproblem at each node. The proposed approach is fast and does not suffer from as high a computational burden as that incurred by typical model selection algorithms. Empirical results over 26 data sets indicate that our approach is faster and achieves better classification accuracy compared to statistical model select algorithms.
Yuanhong Li, Ming Dong 0001, Ravi Kothari
IEEE Trans. Neural Networks2
2004 Discovering Document Semantics QBYS: A System for Querying the WWW by Semantics
Farshad Fotouhi, Sorin Draghici, Ming Dong 0001
Multim. Tools Appl.4
2003 Evolution based approaches to the preservation of endangered natural languages
abstract
Cultural algorithms, a form of evolutionary programming, employ a dual inheritance mechanism at population and knowledge levels to support problem solving, reasoning and knowledge extraction. Domain knowledge is extracted and separated from individuals within a population and is placed in a belief space. Hierarchical structures employed in the belief space help to accelerate and guide population evolution. The structure of the cultural algorithm lends itself well to a data rich, but knowledge poor distributed environment. In this paper we investigate the use of cultural algorithms to collect and mediate information collected from Web searches and Web services related to the task of acquiring and preserving knowledge about endangered languages where this knowledge about endangered languages where this knowledge is stored in a number of disparate sites.
Jeffrey M. Stefan, Robert G. Reynolds, F. Fatouhi, Anthony Aristar, Shiyong Lu, Ming Dong 0001
IEEE Congress on Evolutionary Computation6
2003 Analyzing dividend events with neural network rule extraction
abstract
Over the last two decades, artificial neural networks (ANN) have been applied to solve a variety of problems such as pattern classification and function approximation. In many applications, it is desirable to extract knowledge from trained neural networks for the users to gain a better understanding of the network's solution. In this paper, we apply REFANN (rule extraction from function approximating neural networks) in dividend events study. Based on our study of 1530 dividend initiations and 692 resumptions events from April 1965 to December 2000, we find that the positive relation between the short-term price reaction and the ratio of annualized dividend amount to stock price is primarily limited to 96 firms that have high dividend ratio and small firm size. The results suggest that the degree of short-term stock price underreaction to dividend events may not be as dramatic as previously believed. The results also show that the relations between the stock price response and firm size is also different across different types of firms. It is suggested that drawing the conclusions from the whole dividend events data may leave some important information unexamined. Our rule extraction method may shed some lights on further empirical research in corporate events studies because more information can be drawn from the data.
Ming Dong 0001, Xu-Shen Zhou
IJCNN1
2003 Classifiability based omnivariate decision trees
abstract
Decision trees represent a simple and powerful method of induction from labeled examples. Univariate decision trees consider the value of a single attribute at each node, leading to the splits that are parallel to the axes. In linear multivariate decision trees, all the attributes are used and the partition at each node is based on a linear discriminate (a hyperplane). Nonlinear multivariate decision trees are able to divide the input space arbitrarily based on higher order parameterizations of the discriminate, though one should be aware of the increase of the complexity and the decrease in the number of examples available as moves further down the tree. In omnivariate decision trees, the decision node may be univariate, linear, or nonlinear. Such architecture frees the designer from choosing the appropriate tree type for a given problem. In this paper, we propose to do the model selection at each decision node based on a novel classifiability measure when building omnivariate decision trees. The classifiability measure captures the possible sources of misclassification with relative ease and is able to accurately reflect the complexity of subproblems at each node. The proposed approach does not require the time consuming statistic tests at each node and therefore does not suffer from as high computational burden as typical model selection algorithm. Our simulation results over several data sets indicate that our approach can achieve at least as good classification accuracy as statistical tests based model select algorithms, but in much faster speed.
Yuanhong Li, Ming Dong 0001
IJCNN2
2003 Feature subset selection using a new definition of classifiability
Ming Dong 0001, Ravi Kothari
Pattern Recognit. Lett.1
2001 Look-ahead based fuzzy decision tree induction
abstract
Decision tree induction is typically based on a top-down greedy algorithm that makes locally optimal decisions at each node. Due to the greedy and local nature of the decisions made at each node, there is considerable possibility of instances at the node being split along branches such that instances along some or all of the branches require a large number of additional nodes for classification. In this paper, we present a computationally efficient way of incorporating look-ahead into fuzzy decision tree induction. Our algorithm is based on establishing the decision at each internal node by jointly optimizing the node splitting criterion (information gain or gain ratio) and the classifiability of instances along each branch of the node. Simulations results confirm that the use of the proposed look-ahead method leads to smaller decision trees and as a consequence better test performance.
Ming Dong 0001, Ravi Kothari
IEEE Trans. Fuzzy Syst.1
1999 Neighborhood induced stochastic resonance
abstract
Spatial interactions over a local neighborhood between neurons is common in many models of neural networks. The spatial interaction is often shaped to allow for improved response in the application for which the network is constructed. In this paper, we study an additional effect that can occur when such spatial interactions are imposed. Specifically, we find that the net input received by a neuron from neighboring neurons can induce stochastic resonance thereby increasing the signal-to-noise ratio of the neuron output. We also find that stochastic resonance is enhanced with increasing variance between the neighborhood signals received by the neuron. Thus the net signal received by a neuron due to neighborhood neurons plays a role similar to that played by noise in a classical stochastic resonance situations. We believe that these results, when combined with the more traditional use of spatial interaction, motivate the design of new and enhanced neural detectors.
Ravi Kothari, Ming Dong 0001, Dinesh K. Bhatia
IJCNN2