VLDB 2026 Research / reviewers in the wild / expert
Haojin Yang 0001
dblp:94/10762
· DBLP profile ↗
47ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-8733-5772ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 22 · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Image Token Matters: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent EditingabstractLarge Vision-Language Models (LVLMs) with discrete image tokenizers unify multimodal representations by encoding visual inputs into a finite set of tokens. Despite their effectiveness, we find that these models still hallucinate non-existent objects. We hypothesize that one reason is due to visual priors induced during training: when certain image tokens frequently co-occur in the same spatial regions and represent shared objects, they become strongly associated with the verbalizations of those objects. As a result, the model may hallucinate by evoking visually absent tokens that often co-occur with present ones. To test this assumption, we construct a co-occurrence graph of image tokens using a segmentation dataset and employ a Graph Neural Network (GNN) with contrastive learning followed by a clustering method to group tokens that frequently co-occur in similar visual contexts. We find that hallucinations predominantly correspond to clusters whose tokens dominate the input, and more specifically, that the visually absent tokens in those clusters show much higher correlation with hallucinated objects compared to tokens present in the image. Based on this observation, we propose a hallucination mitigation method that suppresses the influence of visually absent tokens by modifying latent image embeddings during generation. Experiments show our method reduces hallucinations while preserving expressivity. Weixing Wang 0005, Zifeng Ding, Jindong Gu, Christoph Meinel, Gerard de Melo, Haojin Yang 0001 |
NeurIPS | 7 |
| 2024 | Enhancing Optimization Robustness in 1-Bit Neural Networks Through Stochastic Sign Descent
Nianhui Guo, Christoph Meinel, Haojin Yang 0001 |
ECCV (31) | 4 |
| 2024 | Otem-IGCD: An Optimal Transport-based EM Framework for Imbalanced Generalized Category DiscoveryabstractGeneralized Class Discovery (GCD) seeks to identify both known and unknown categories within an unlabeled dataset, utilizing the knowledge from a labeled dataset of known classes. Existing research implicitly/explicitly assumes that the frequency of occurrence for each category, whether known or unknown, is approximately the same in the unlabeled data. However, real-world scenarios often exhibit a long-tailed distribution of visual classes, where known or common classes appear more frequently than unknown or rare ones. Addressing this discrepancy, we introduce a new challenge: Imbalanced Generalized Category Discovery (IGCD), which deals with an imbalanced distribution in unlabeled data, favoring known over unknown classes. To tackle this, we propose a novel Optimal Transport-based Expectation Maximization framework for Imbalanced Generalized Category Discovery (Otem-IGCD) by aligning the marginal class prior distribution. Otem-IGCD also incorporates a systematic mechanism for estimating the imbalanced class prior distribution under the GCD setup. Our comprehensive experiments reveal that Otem-IGCD surpasses previous state-of-the-art GCD methods by achieving an improvement of approximately 2 - 4% on CIFAR100 and 15 - 19% on ImageNet-100, indicating its superior effectiveness in solving the Imbalanced GCD problem. Ben Dai, Christoph Meinel, Haojin Yang 0001 |
IJCNN | 4 |
| 2024 | Low-bit CUTLASS GEMM Template Auto-tuning using Neural NetworkabstractOptimizing General Matrix Multiplication (GEMM) on GPU platforms has become increasingly important due to the scaling demands of modern deep neural network research. While substantial progress has been made in accelerating high-precision GEMM, optimizing lower-bit GEMM remains an open problem. The CUTLASS library offers highly optimized low-bit GEMM based on tensor cores, but performance varies significantly with tile and pipeline settings across different GPUs. We introduce a novel auto-tuning framework for low-bit CUTLASS GEMM that employs a neural network model to predict optimal GEMM template parameters for target GPUs. This model was trained on a synthetic dataset featuring various matrix sizes from different Ampere GPUs and evaluated on these GPUs. In the test dataset, our method achieved an accuracy of up to 92.9%. Real-time evaluations of low-bit data types on the A100 GPU demonstrated speedups of up to 2.03× for GEMM and 1.44× for the linear layer compared to the default templates. Nianhui Guo, Christoph Meinel, Haojin Yang 0001 |
ISPA | 4 |
| 2024 | Guided Cluster Aggregation: A Hierarchical Approach to Generalized Category DiscoveryabstractDespite advances in image recognition, recognizing novel categories in unlabeled data remains challenging for machine learning methods, even though humans can perform this task with ease. A recently developed setting to tackle this problem is Generalized Category Discovery (GCD), in which the task is to, given a labeled dataset, classify an unlabeled dataset, where the unlabeled dataset contains both known classes and novel classes that do not appear in the labeled data. Existing GCD methods mostly focus on learning strong image representations, on which they then apply a clustering algorithm such as k-means. Despite obtaining good performance, they do not fully exploit the potential of the learned features due to the simple nature of the clustering mechanism. To address this issue, we make use of the fact that local neighborhoods in self-supervised feature spaces are highly homogeneous. We leverage this observation to develop Guided Cluster Aggregation (GCA), a hierarchical approach that first groups the data into small clusters of high purity, then aggregates them into larger clusters. Experiments show that GCA outperforms semi-supervised k-means in most cases, especially in fine-grained classification tasks. Code available at https://github.com/J-L-O/guidedcluster-aggregation. Jona Otholt, Christoph Meinel, Haojin Yang 0001 |
WACV | 3 |
| 2024 | A flexible BERT model enabling width- and depth-dynamic inferenceabstractFine-tuning and inference on Large Language Models like BERT have become increasingly expensive regarding memory cost and computation resources. The recently proposed computation-flexible BERT models facilitate their deployment in varied computational environments. Training such flexible BERT models involves jointly optimizing multiple BERT subnets, which will unavoidably interfere with one another. Besides, the performance of large subnets is limited by the performance gap between the smallest subnet and the supernet, despite efforts to enhance the smaller subnets. In this regard, we propose layer-wise Neural grafting to boost BERT subnets, especially the larger ones. The proposed method improves the average performance of the subnets on six GLUE tasks and boosts the supernets on all GLUE tasks and the SQuAD data set. Based on the boosted subnets, we further build an inference framework enabling practical width- and depth-dynamic inference regarding different inputs by combining width-dynamic gating modules and early exit off-ramps in the depth dimension. Experimental results show that the proposed framework achieves a better dynamic inference range than other methods in terms of trading off performance and computational complexity on four GLUE tasks and SQuAD. In particular, our best-tradeoff inference result outperforms other fixed-size models with similar amount of computations. Compared to BERT-Base, the proposed inference framework yields a 1.3-point improvement in the average GLUE score and a 2.2-point increase in the F1 score on SQuAD, while reducing computations by around 45%. Christoph Meinel, Haojin Yang 0001 |
Comput. Speech Lang. | 3 |
| 2023 | Boosting Bert Subnets with Neural GraftingabstractPre-trained Language Models in Natural Language Processing have become increasingly computationally expensive and memory demanding. The recently proposed computation-adaptive BERT models facilitate their deployment in practical applications. Training such a BERT model involves jointly optimizing subnets of varying sizes, which is not easy due to their mutual interference with one another. The larger-size subnets in particular could deteriorate when there is a large performance gap between the smallest subnet and the super-net. In this work, we propose Neural grafting to boost BERT subnets, especially the larger ones. Specifically, we regard the less important sub-modules of a BERT model as less active and reactivate them via layer-wise Neural grafting. Experimental results show that the proposed method improves the average performance of BERT subnets on six datasets of GLUE benchmark. The subnet performing comparable to the supernet BERT-Base reduces around 67% and 70% inference latency on GPU and CPU, respectively. Moreover, we compare two Neural grafting strategies under varied experimental settings, hoping to shed light on the application scenarios of Neural grafting. Christoph Meinel, Haojin Yang 0001 |
ICASSP | 3 |
| 2023 | Flexible BERT with Width- and Depth-dynamic InferenceabstractPre-trained Language Models bring about an in-creasing computational and memory cost. The recently proposed computation-flexible BERT models facilitate their deployment in varied computational environments. Training such flexible BERT models involves jointly optimizing multiple BERT subnets that inevitably interfere with one another. Besides, the performance of large sub nets is limited when there is a significant performance gap between the smallest sub net and the supernet, despite methods managing to enhance the smaller subnets. We propose layer-wise Neural grafting to boost BERT subnets, particularly the larger ones. The proposed method improves the average performance of BERT sub nets on six out of eight GLUE tasks. Furthermore, we build a flexible BERT model that enables practical width- and depth-dynamic inference regarding different inputs by combining width-dynamic gating modules and early exit off-ramps in the depth dimension. Experimental results demonstrate that the proposed framework achieves a better dynamic inference range than other methods in the trade-off between performance and computational complexity on four GLUE tasks and the SQuAD data set. Our optimal-tradeoff inference result, in particular, outperforms related fixed-size models with comparable computational complexity. Compared to the supernet, BERT-Base, this inference result improves the average GLUE score and Fl score on SQuAD by 1.3 and 2.2 absolute points, respectively, and decreases computations by around 45%. Christoph Meinel, Haojin Yang 0001 |
IJCNN | 3 |
| 2023 | SMKD: Selective Mutual Knowledge DistillationabstractMutual knowledge distillation (MKD) is a technique used to transfer knowledge between multiple models in a collaborative manner. However, it is important to note that not all knowledge is accurate or reliable, particularly under challenging conditions such as label noise, which can lead to models that memorize undesired information. This problem can be addressed by improving the reliability of the knowledge source, as well as selectively selecting reliable knowledge for distillation. While making a model more reliable is a widely studied topic, selective MKD has received less attention. To address this, we propose a new framework called selective mutual knowledge distillation (SMKD). The key component of SMKD is a generic knowledge selection formulation, which allows for either static or progressive selection thresholds. Additionally, SMKD covers two special cases: using no knowledge and using all knowledge, resulting in a unified MKD framework. We present extensive experimental results to demonstrate the effectiveness of SMKD and justify its design. Xinshao Wang, Neil Robertson 0002, David A. Clifton, Christoph Meinel, Haojin Yang 0001 |
IJCNN | 6 |
| 2022 | Synthesis in Style: Semantic Segmentation of Historical Documents using Synthetic DataabstractOne of the most pressing problems in the automated analysis of historical documents is the availability of annotated training data. The problem is that labeling samples is a time-consuming task because it requires human expertise and thus, cannot be automated well. In this work, we propose a novel method to construct synthetic labeled datasets for historical documents where no annotations are available. We train a StyleGAN model to synthesize document images that capture the core features of the original documents. While originally, the StyleGAN architecture was not intended to produce labels, it indirectly learns the underlying semantics to generate realistic images. Using our approach, we can extract the semantic information from the intermediate feature maps and use it to generate ground truth labels. To investigate if our synthetic dataset can be used to segment the text in historical documents, we use it to train multiple supervised segmentation models and evaluate their performance. We also train these models on another dataset created by a state-of-the-art synthesis approach to show that the models trained on our dataset achieve better results while requiring even less human annotation effort. Christian Bartz, Hendrik Rätz, Jona Otholt, Christoph Meinel, Haojin Yang 0001 |
ICPR | 5 |
| 2021 | One Model to Reconstruct Them All: A Novel Way to Use the Stochastic Noise in StyleGAN
Christian Bartz, Joseph Bethge, Haojin Yang 0001, Christoph Meinel |
BMVC | 3 |
| 2021 | Denoising AutoEncoder Based Delete and Generate Approach for Text Style Transfer
Haojin Yang 0001, Christoph Meinel |
ICANN (3) | 2 |
| 2021 | MeliusNet: An Improved Network Architecture for Binary Neural NetworksabstractBinary Neural Networks (BNNs) are neural networks which use binary weights and activations instead of the typical 32-bit floating point values. They have reduced model sizes and allow for efficient inference on mobile or embedded devices with limited power and computational resources. However, the binarization of weights and activations leads to feature maps of lower quality and lower capacity and thus a drop in accuracy compared to their 32-bit counterparts. Previous work has increased the number of channels or used multiple binary bases to alleviate these problems. In this paper, we instead present an architectural approach: MeliusNet. It consists of alternating a DenseBlock, which increases the feature capacity, and our proposed ImprovementBlock, which increases the feature quality. Experiments on the ImageNet dataset demonstrate the superior performance of our MeliusNet over a variety of popular binary architectures with regards to both computation savings and accuracy. Furthermore, BNN models trained with our method can match the accuracy of the popular compact network MobileNet-v1 in terms of model size and number of operations. Our code is published online: https://github.com/hpi-xnor/BMXNet-v2. Joseph Bethge, Christian Bartz, Haojin Yang 0001, Christoph Meinel |
WACV | 3 |
| 2020 | Best Student Forcing: A Simple Training Mechanism in Adversarial Language GenerationabstractLanguage models trained with Maximum Likelihood Estimation (MLE) have been considered as a mainstream solution in Natural Language Generation (NLG) for years. Recently, various approaches with Generative Adversarial Nets (GANs) have also been proposed. While offering exciting new prospects, GANs in NLG by far are nevertheless reportedly suffering from training instability and mode collapse, and therefore outperformed by conventional MLE models. In this work, we propose techniques for improving GANs in NLG, namely Best Student Forcing (BSF), a novel yet simple adversarial training mechanism in which generated sequences of high quality are selected as temporary ground-truth to further train the generator. We also use an ensemble of discriminators to increase training stability and sample diversity. Evaluation shows that the combination of BSF and multiple discriminators consistently performs better than previous GAN approaches over various metrics, and outperforms a baseline MLE in terms of Fr ́ech ́et Distance, a recently proposed metric capturing both sample quality and diversity. Jonathan Sauder, Xiaoyin Che, Gonçalo Mordido, Haojin Yang 0001, Christoph Meinel |
LREC | 5 |
| 2020 | BMXNet 2: An Open Source Framework for Low-bit Networks - Reproducing, Understanding, Designing and ShowcasingabstractBinary and quantized neural networks are a promising technique to run convolutional neural networks on mobile or embedded devices. BMXNet 2 is an open-source framework that provides a broad basis for academia and industry. It provides a modern implementation of binary and quantized layers with a wide array of implemented state-of-the-art models. Our implementation fosters reproducibility of other works and our own work through publishing model code, hyperparameters, detailed model graphs, and training logs. Furthermore, we implement several applications for BNNs, including demo applications, which can run on a smartphone or a Raspberry Pi. The code can be found online: https://github.com/hpi-xnor/BMXNet-v2 Joseph Bethge, Christian Bartz, Haojin Yang 0001, Christoph Meinel |
ACM Multimedia | 3 |
| 2020 | microbatchGAN: Stimulating Diversity with Multi-Adversarial DiscriminationabstractWe propose to tackle the mode collapse problem in generative adversarial networks (GANs) by using multiple discriminators and assigning a different portion of each minibatch, called microbatch, to each discriminator. We gradually change each discriminator's task from distinguishing between real and fake samples to discriminating samples coming from inside or outside its assigned microbatch by using a diversity parameter α. The generator is then forced to promote variety in each minibatch to make the micro-batch discrimination harder to achieve by each discriminator. Thus, all models in our framework benefit from having variety in the generated set to reduce their respective losses. We show evidence that our solution promotes sample diversity since early training stages on multiple datasets. Gonçalo Mordido, Haojin Yang 0001, Christoph Meinel |
WACV | 2 |
| 2020 | Recurrent generative adversarial network for learning imbalanced medical image semantic segmentation
Mina Rezaei, Haojin Yang 0001, Christoph Meinel |
Multim. Tools Appl. | 2 |
| 2019 | Training Accurate Binary Neural Networks from ScratchabstractBinary neural networks are a promising approach to execute convolutional neural networks on devices with low computational power. Previous work on this subject often quantizes pretrained full-precision models and uses complex training strategies. In our work, we focus on increasing the performance of binary neural networks by training from scratch with a simple training strategy. In our experiments we show that we are able to achieve state-of-the-art results on standard benchmark datasets. Further, we analyze how full-precision network structures can be adapted for efficient binary networks and adopt a network architecture based on a DenseNet for binary networks, which lets us improve the state-of-the-art even further. Our source code can be found online: https://github.com/hpi-xnor/BMXNet-v2. Joseph Bethge, Haojin Yang 0001, Christoph Meinel |
ICIP | 2 |
| 2019 | Conditional Generative Adversarial Refinement Networks for Unbalanced Medical Image Semantic SegmentationabstractWe propose a new generative adversarial architecture to mitigate imbalance data problem in medical image semantic segmentation where the majority of pixels belongs to a healthy region and few belong to lesion or non-health region. A model trained with imbalanced data tends to bias towards healthy data which is not desired in clinical applications and predicted outputs by these networks have high precision and low sensitivity. We propose a new conditional generative refinement network with three components: a generative, a discriminative, and a refinement networks to mitigate imbalanced data problem through ensemble learning. The generative network learns to the segment at the pixel level by getting feedback from the discriminative network according to the true positive and true negative maps. On the other hand, the refinement network learns to predict the false positive and the false negative masks produced by the generative network that has significant value, especially in medical application. The final semantic segmentation masks are then composed by the output of the three networks. The proposed architecture shows state-of-the-art results on LiTS-2017 for simultaneous liver and lesion segmentation, and MDA231 for microscopic cell segmentation. We have achieved competitive results on BraTS-2017 for brain tumor segmentation. Mina Rezaei, Haojin Yang 0001, Konstantin Harmuth, Christoph Meinel |
WACV | 2 |
| 2018 | SEE: Towards Semi-Supervised End-to-End Scene Text RecognitionabstractDetecting and recognizing text in natural scene images is a challenging, yet not completely solved task. In recent years several new systems that try to solve at least one of the two sub-tasks (text detection and text recognition) have been proposed. In this paper we present SEE, a step towards semi-supervised neural networks for scene text detection and recognition, that can be optimized end-to-end. Most existing works consist of multiple deep neural networks and several pre-processing steps. In contrast to this, we propose to use a single deep neural network, that learns to detect and recognize text from natural images, in a semi-supervised way. SEE is a network that integrates and jointly learns a spatial transformer network, which can learn to detect text regions in an image, and a text recognition network that takes the identified text regions and recognizes their textual content. We introduce the idea behind our novel approach and show its feasibility, by performing a range of experiments on standard benchmark datasets, where we achieve competitive results. Christian Bartz, Haojin Yang 0001, Christoph Meinel |
AAAI | 2 |
| 2018 | Instance Tumor Segmentation using Multitask Convolutional Neural NetworkabstractAutomatic tumor segmentation is an important and challenging clinical task because tumors have different sizes, shapes, contrasts, and locations. In this paper, we present an automatic instance semantic segmentation method based on deep neural networks (DNNs). The proposed networks are tailored to picture tumors in magnetic resonance imaging (MRI) and computed tomography (CT) images. We present an end-to-end multitask learning architecture comprising three stages, namely detection, segmentation, and classification. This paper introduces a new technique for tumor detection based on high-level extracted features from convolutional neural networks (CNNs) using the Hough transform technique. The detected tumor(s), are segmented with a set of fully connected (FC) layers, and the segmented mask is classified through FCs. The proposed architecture gives promising results on the popular medical image benchmarks. Our framework is generalized in the sense that it can be used in different types of medical images in varied sizes, such as the Liver Tumor Segmentation (LiTS-2017) challenge, and the Brain Tumor Segmentation (BraTS-2016) benchmark. Mina Rezaei, Haojin Yang 0001, Christoph Meinel |
IJCNN | 2 |
| 2018 | Image Captioning with Deep Bidirectional LSTMs and Multi-Task LearningabstractGenerating a novel and descriptive caption of an image is drawing increasing interests in computer vision, natural language processing, and multimedia communities. In this work, we propose an end-to-end trainable deep bidirectional LSTM (Bi-LSTM (Long Short-Term Memory)) model to address the problem. By combining a deep convolutional neural network (CNN) and two separate LSTM networks, our model is capable of learning long-term visual-language interactions by making use of history and future context information at high-level semantic space. We also explore deep multimodal bidirectional models, in which we increase the depth of nonlinearity transition in different ways to learn hierarchical visual-language embeddings. Data augmentation techniques such as multi-crop, multi-scale, and vertical mirror are proposed to prevent overfitting in training deep models. To understand how our models “translate” image to sentence, we visualize and qualitatively analyze the evolution of Bi-LSTM internal states over time. The effectiveness and generality of proposed models are evaluated on four benchmark datasets: Flickr8K, Flickr30K, MSCOCO, and Pascal1K datasets. We demonstrate that Bi-LSTM models achieve highly competitive performance on both caption generation and image-sentence retrieval even without integrating an additional mechanism (e.g., object detection, attention model). Our experiments also prove that multi-task learning is beneficial to increase model generality and gain performance. We also demonstrate the performance of transfer learning of the Bi-LSTM model significantly outperforms previous methods on the Pascal1K dataset. Cheng Wang 0002, Haojin Yang 0001, Christoph Meinel |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | Automatic Lecture Subtitle Generation and How It HelpsabstractIn this paper we propose an integrated framework of automatic bilingual subtitle generation for lecture videos, especially for MOOCs. The framework consists of Automatic Speech Recognition (ASR), Sentence Boundary Detection (SBD), and Machine Translation (MT). Then we quantitatively evaluate the auto-generated subtitles, the manually produced subtitles from scratch, and the auto-generated subtitles with manual modification in term of accuracy and time expenditure, in both original and target languages. The result shows that the auto-generated subtitles in the original language (English) are fairly accurate already. By using them as the draft, human subtitle producers can save 54% of the working time and simultaneously reduce the error rate by 54.3%, which is a significant improvement. However, the effectiveness of machine translated subtitles (English to Chinese) is limited. In the end, if the proposed framework is applied, the total working time in preparing bilingual subtitles can be shortened by approximately 1/3, with no decline in quality. Xiaoyin Che, Sheng Luo 0002, Haojin Yang 0001, Christoph Meinel |
ICALT | 3 |
| 2017 | Language Identification Using Deep Convolutional Recurrent Neural Networks
Christian Bartz, Tom Herold, Haojin Yang 0001, Christoph Meinel |
ICONIP (6) | 3 |
| 2017 | Deep Neural Network with l2-Norm Unit for Brain Lesions Detection
Mina Rezaei, Haojin Yang 0001, Christoph Meinel |
ICONIP (4) | 2 |
| 2017 | BMXNet: An Open-Source Binary Neural Network Implementation Based on MXNetabstractBinary Neural Networks (BNNs) can drastically reduce memory size and accesses by applying bit-wise operations instead of standard arithmetic operations. Therefore it could significantly improve the efficiency and lower the energy consumption at runtime, which enables the application of state-of-the-art deep learning models on low power devices. BMXNet is an open-source BNN library based on MXNet, which supports both XNOR-Networks and Quantized Neural Networks. The developed BNN layers can be seamlessly applied with other standard library components and work in both GPU and CPU mode. BMXNet is maintained and developed by the multimedia research group at Hasso Plattner Institute and released under Apache license. Extensive experiments validate the efficiency and effectiveness of our implementation. The BMXNet library, several sample projects, and a collection of pre-trained binary deep models are available for download at https://github.com/hpi-xnor. Haojin Yang 0001, Martin Fritzsche, Christian Bartz, Christoph Meinel |
ACM Multimedia | 1 |
| 2016 | Pre-Course Key Segment Analysis of Online Lecture VideosabstractIn this paper we propose a method to evaluate the importance of lecture video segments in online courses. The video will be first segmented based on the slide transition. Then we evaluate the importance of each segment based on our analysis of the teacher's focus. This focus is mainly identified by exploring features in the slide and the speech. Since the whole analysis process is based on multimedia materials, it could be done before the official start of the course. By setting survey questions and collecting forum statistics in the MOOC "Web Technologies", the proposed method is evaluated. Both the general trend and the high accuracy of selected key segments (over 70%) prove the effectiveness of the proposed method. Xiaoyin Che, Thomas Staubitz, Haojin Yang 0001, Christoph Meinel |
ICALT | 3 |
| 2016 | Action Recognition in Surveillance Video Using ConvNets and Motion History Image
Sheng Luo 0002, Haojin Yang 0001, Cheng Wang 0002, Xiaoyin Che, Christoph Meinel |
ICANN (2) | 2 |
| 2016 | Real-Time Action Recognition in Surveillance Videos Using ConvNets
Sheng Luo 0002, Haojin Yang 0001, Cheng Wang 0002, Xiaoyin Che, Christoph Meinel |
ICONIP (3) | 2 |
| 2016 | Exploring multimodal video representation for action recognitionabstractA video contains rich perceptual information, such as visual appearance, motion and audio, which can be used for understanding the activities in videos. Recent works have shown the combination of appearance (spatial) and motion (temporal) clues can significantly improve human action recognition performance in videos. To further explore the multimodal representation of video in action recognition, We propose a framework to learn a multimodal representations from video appearance, motion as well as audio data. Convolutional Neural Networks (CNN) are trained for each modality respectively. For fusing multiple features extracted with CNNs, we propose to add a fusion layer on the top of CNNs to learn a joint video representation. In fusion phase, we investigate both early fusion and late fusion with Neural Network and Support Vector Machine. Compare to existing works, (1) our work measures the benefits of taking audio information into consideration and (2) implements sophisticated fusion methods. The effectiveness of proposed approach is evaluated on UCF101 and UCF101-50 (selected subset in which each video contains audio data) for action recognition. The experimental results show that different modalities are complementary to each other and multimodal representation can be beneficial for final prediction. Furthermore, proposed fusion approach achieves 85.1% accuracy in fusing spatial-temporal on UCF101 (split 1), which is very competitive to state-of-the-art works. Cheng Wang 0002, Haojin Yang 0001, Christoph Meinel |
IJCNN | 2 |
| 2016 | Sentence Boundary Detection Based on Parallel Lexical and Acoustic Models
Xiaoyin Che, Sheng Luo 0002, Haojin Yang 0001, Christoph Meinel |
INTERSPEECH | 3 |
| 2016 | Punctuation Prediction for Unsegmented Transcript Based on Word Vector
Xiaoyin Che, Cheng Wang 0002, Haojin Yang 0001, Christoph Meinel |
LREC | 3 |
| 2016 | Image Captioning with Deep Bidirectional LSTMsabstractThis work presents an end-to-end trainable deep bidirectional LSTM (Long-Short Term Memory) model for image captioning. Our model builds on a deep convolutional neural network (CNN) and two separate LSTM networks. It is capable of learning long term visual-language interactions by making use of history and future context information at high level semantic space. Two novel deep bidirectional variant models, in which we increase the depth of nonlinearity transition in different way, are proposed to learn hierarchical visual-language embeddings. Data augmentation techniques such as multi-crop, multi-scale and vertical mirror are proposed to prevent overfitting in training deep models. We visualize the evolution of bidirectional LSTM internal states over time and qualitatively analyze how our models "translate" image to sentence. Our proposed models are evaluated on caption generation and image-sentence retrieval tasks with three benchmark datasets: Flickr8K, Flickr30K and MSCOCO datasets. We demonstrate that bidirectional LSTM models achieve highly competitive performance to the state-of-the-art results on caption generation even without integrating additional mechanism (e.g. object detection, attention model etc.) and significantly outperform recent methods on retrieval task Cheng Wang 0002, Haojin Yang 0001, Christian Bartz, Christoph Meinel |
ACM Multimedia | 2 |
| 2016 | SceneTextReg: A Real-Time Video OCR SystemabstractWe showcase a system for real-time video text recognition. The system is based on the standard workflow of text spotting system, which includes text detection and word recognition procedures. We apply deep neural networks in both procedures. In text localization stage, textual candidates are roughly captured by using a Maximally Stable Extremal Regions (MSERs) detector with high recall rate, false alarms are then eliminated by using Convolutional Neural Network (CNN ) verifier. For word recognition, we developed a skeleton based method for segmenting text region from its background, then a CNN based word recognizer is utilized for recognizing texts. Our current implementation demonstrates a real time performance for recognizing scene text by using a standard laptop with webcam. The word recognizer achieves competitive result to state-of-the-art methods by only using synthetical training data. Haojin Yang 0001, Cheng Wang 0002, Christian Bartz, Christoph Meinel |
ACM Multimedia | 1 |
| 2016 | A deep semantic framework for multimodal representation learning
Cheng Wang 0002, Haojin Yang 0001, Christoph Meinel |
Multim. Tools Appl. | 2 |
| 2015 | Reward-based Intermittent Reinforcement in Gamification for E-learningabstractNowadays gamification is a hot topic in the world, a lot of websites, applications and researches adapt this
method to arouse users' motivation. From the past experience, gamification indeed has a positive influence
on users' motivation especially in e-learning field. However, the gamification method either is hard to be
applied to professional content called meaningful gamification or is negative on user's intrinsic motivation
called reward-based gamification. So we study the game addiction mechanism and propose the reward-based
intermittent reinforcement method in gamification to take advantage of user independence feature in
the latter one and eliminate the negative influence on user's intrinsic motivation. In order to investigate the
practicability and integrate effectiveness, we implement this model in our tele-teaching platform. Sheng Luo 0002, Haojin Yang 0001, Christoph Meinel |
CSEDU (1) | 2 |
| 2015 | Does Multilevel Semantic Representation Improve Text Categorization?
Cheng Wang 0002, Haojin Yang 0001, Christoph Meinel |
DEXA (1) | 2 |
| 2015 | Visual-Textual Late Semantic Fusion Using Deep Neural Network for Document Categorization
Cheng Wang 0002, Haojin Yang 0001, Christoph Meinel |
ICONIP (1) | 2 |
| 2015 | Deep Semantic Mapping for Cross-Modal RetrievalabstractCross-Modal mapping plays an essential role in multimedia information retrieval systems. However, most of existing work paid much attention on learning mapping functions but neglected the exploration of high-level semantic representation of modalities. Inspired by recent success of deep learning, in this paper, deep CNN (convolutional neural networks) features and topic features are utilized as visual and textual semantic representation respectively. To investigate the highly non-linear semantic correlation between image and text, we propose a regularized deep neural network(RE-DNN) for semantic mapping across modalities. By imposing intra-modal regularization as supervised pre-training, we finally learn a joint model which captures both intra-modal and inter-modal relationships. Our approach is superior to previous work in follows: (1) it explores high-level semantic correlations, (2) it requires little prior knowledge for model training, (3) it is able to tackle modality missing problem. Extensive experiments on benchmark Wikipedia dataset show RE-DNN outperforms the state-of-the-art approaches in cross-modal retrieval. Cheng Wang 0002, Haojin Yang 0001, Christoph Meinel |
ICTAI | 2 |
| 2015 | An Improved System For Real-Time Scene Text RecognitionabstractIn this paper we showcase a system for real-time text detection and recognition. We apply deep features created by Convolutional Neural Networks (CNNs) for both text detection and word recognition task. For text detection we follow the common localization-verification scheme which already shown its excellent ability in numerous previous work. In text localization stage, textual regions are roughly detected by using a MSERs (Maximally Stable Extremal Regions) detector with high recall rate. False alarms are then eliminated by using a CNNs classifier, and remaining text regions are further grouped into words. In the word recognition stage, we developed an skeleton-based text binarization method for segmenting text from its background. A CNNs based recognizer is then applied for recognizing character. The initial experiments show the powerful ability of deep features for text classification comparing with commonly used visual features. Our current implementation demonstrates real-time performance for recognizing scene text by using a standard PC with webcam. Haojin Yang 0001, Cheng Wang 0002, Xiaoyin Che, Sheng Luo 0002, Christoph Meinel |
ICMR | 1 |
| 2015 | Concept-Based Multimodal Learning for Topic Generation
Cheng Wang 0002, Haojin Yang 0001, Xiaoyin Che, Christoph Meinel |
MMM (1) | 2 |
| 2015 | Table Detection from Slide Images
Xiaoyin Che, Haojin Yang 0001, Christoph Meinel |
PSIVT | 2 |
| 2014 | Improving text recognition by distinguishing scene and overlay textabstractVideo texts are closely related to the content of a video. They provide a valuable source for indexing and interpretation of video data. Text detection and recognition task in images or videos typically distinguished between overlay and scene text. Overlay text is artificially superimposed on the image at the time of editing and scene text is text captured by the recording system. Typically, OCR systems are specialized on one kind of text type. However, in video images both types of text can be found. In this paper, we propose a method to automatically distinguish between overlay and scene text to dynamically control and optimize post processing steps following text detection. Based on a feature combination a Support Vector Machine (SVM) is trained to classify scene and overlay text. We show how this distinction in overlay and scene text improves the word recognition rate. Accuracy of the proposed methods has been evaluated by using publicly available test data sets. Bernhard Quehl, Haojin Yang 0001, Harald Sack |
ICMV | 2 |
| 2014 | A framework for improved video text detection and recognition
Haojin Yang 0001, Bernhard Quehl, Harald Sack |
Multim. Tools Appl. | 1 |
| 2013 | Lecture video segmentation by automatically analyzing the synchronized slidesabstractIn this paper we propose a solution which segments lecture video by analyzing its supplementary synchronized slides. The slides content derives automatically from OCR (Optical Character Recognition) process with an approximate accuracy of 90%. Then we partition the slides into different subtopics by examining their logical relevance. Since the slides are synchronized with the video stream, the subtopics of the slides indicate exactly the segments of the video. Our evaluation reveals that the average length of segments for each lecture is ranged from 5 to 15 minutes, and 45% segments achieved from test datasets are logically reasonable. Xiaoyin Che, Haojin Yang 0001, Christoph Meinel |
ACM Multimedia | 2 |
| 2012 | Automated Extraction of Lecture Outlines from Lecture Videos - A Hybrid Solution for Lecture Video Indexing
Haojin Yang 0001, Franka Grünewald, Christoph Meinel |
CSEDU (1) | 1 |
| 2011 | Automatic Lecture Video Indexing Using Video OCR TechnologyabstractDuring the last years, digital lecture libraries and lecture video portals have become more and more popular. However, finding efficient methods for indexing multimedia still remains a challenging task. Since the text displayed in a lecture video is closely related to the lecture content, it provides a valuable source for indexing and retrieving lecture contents. In this paper, we present an approach for automatic lecture video indexing based on video OCR technology. We have developed a novel video segmenter for automated slide video structure analysis and a weighted DCT (discrete cosines transformation) based text detector. A dynamic image constrast/brightness adaption serves the purpose of enhancing the text image quality to make it processible by existing common OCR software. Time-based text occurence information as well as the analyzed text content are further used for indexing. We prove the accuracy of the proposed approach by evaluation. Haojin Yang 0001, Maria Siebert, Patrick Lühne, Harald Sack, Christoph Meinel |
ISM | 1 |