EDBT 2026 Demo / reviewers in the wild / expert
Wesley De Neve
dblp:62/1393
· DBLP profile ↗
84ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-8190-3839ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 15 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 7 since 2021Databases, data management, data science and information retrieval · 4Security and privacy · 3Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Boosting 3D Liver Shape Datasets with Diffusion Models and Implicit Neural Representations
Khoa Tuan Nguyen, Francesca Tozzi, Wouter Willaert, Joris Vankerschaver, Niki Rashidian, Wesley De Neve |
MICCAI (2) | 6 |
| 2025 | SpurBreast: A Curated Dataset for Investigating Spurious Correlations in Real-World Breast MRI Classification
Jongbum Won, Wesley De Neve, Joris Vankerschaver, Utku Ozbulak |
MICCAI (16) | 2 |
| 2025 | Impact of U2-type introns on splice site prediction in A. thaliana species using deep learningabstractBACKGROUND: Splice site prediction in plant genomes poses substantial challenges that can be addressed using deep learning models. U2-type introns are especially useful for such studies given their ubiquity in plant genomes and the availability of rich datasets. We formulated two hypotheses: one proposing that short introns may enhance prediction effectiveness due to reduced spatial complexity, and another suggesting that sequences with multiple introns provide a richer context for splicing events. RESULTS: Our findings demonstrate that (1) models trained on datasets containing shorter introns achieve improved effectiveness for acceptor splice sites, but not for donor splice sites, indicating a more nuanced relationship between intron length and splice site prediction than initially hypothesized, and (2) models trained on datasets with multiple introns per sequence show higher effectiveness compared to those trained on datasets with a single intron per sequence. Notably, among the 402 bp sequences analyzed, 72% contained single introns while 28% contained multiple introns for donor sites (36,399 versus 13,987 sequences), with similar proportions observed for acceptor sites (37,236 versus 14,112 sequences). These computational insights align with biological observations, particularly regarding the conserved spatial relationship between branch points and acceptor splice sites, as well as the synergistic effects of multiple introns on splicing efficiency. CONCLUSIONS: The obtained results contribute to a deeper understanding of how intronic features influence splice site prediction and suggest that future prediction models should consider factors such as intron length, multiplicity, and the spatial arrangement of splice-related signals. Espoir K. Kabanga, Seonil Jee, Soeun Yun, Stephen Depuydt, Arnout Van Messem, Wesley De Neve |
BMC Bioinform. | 6 |
| 2025 | Re-assessing accuracy degradation: a framework for understanding DNN behavior on similar-but-non-identical test datasetsabstractAbstract Deep Neural Networks (DNNs) often demonstrate remarkable performance when evaluated on the test dataset used during model creation. However, their ability to generalize effectively when deployed is crucial, especially in critical applications. One approach to assess the generalization capability of a DNN model is to evaluate its performance on replicated test datasets, which are created by closely following the same methodology and procedures used to generate the original test dataset. Our investigation focuses on the performance degradation of pre-trained DNN models in multi-class classification tasks when evaluated on these replicated datasets; this performance degradation has not been entirely explained by generalization shortcomings or dataset disparities. To address this, we introduce a new evaluation framework that leverages uncertainty estimates generated by the models studied. This framework is designed to isolate the impact of variations in the evaluated test datasets and assess DNNs based on the consistency of their confidence in their predictions. By employing this framework, we can determine whether an observed performance drop is primarily caused by model inadequacy or other factors. We applied our framework to analyze 564 pre-trained DNN models across the CIFAR-10 and ImageNet benchmarks, along with their replicated versions. Contrary to common assumptions about model inadequacy, our results indicate a substantial reduction in the performance gap between the original and replicated datasets when accounting for model uncertainty. This suggests a previously unrecognized adaptability of models to minor dataset variations. Our findings emphasize the importance of understanding dataset intricacies and adopting more nuanced evaluation methods when assessing DNN model performance. This research contributes to the development of more robust and reliable DNN models, especially in critical applications where generalization performance is of utmost importance. The code to reproduce our experiments will be available at https://github.com/esla/Reassessing_DNN_Accuracy . Esla Timothy Anzaku, Haohan Wang, Ajiboye Babalola, Arnout Van Messem, Wesley De Neve |
Mach. Learn. | 5 |
| 2025 | Interpreting Stress Detection Models Using SHAP and Attention for MuSe-Stress 2022abstractUnderstanding emotional reactions, especially stress, during job interviews holds significant implications for assessing the well-being of candidates and tailoring feedback. However, current techniques, though effective, often lack interpretability. In this study, we investigate emotion recognition by focusing on making sense of machine-learning models. Specifically, our work leverages the power of interpretable methods in detecting stress through multimodal time series. Building upon prior research, our main contribution is a novel method for calculating feature importance scores using Shapley Additive exPlanations (SHAP) and attention. We applied this technique to models from the MuSe 2022 stress detection competition, generating insights into the importance and interplay of various features in Arousal or Valence prediction. Our findings suggest that leveraging SHAP for feature selection can enhance prediction effectiveness while mitigating computational demands. With this, we introduce an advanced, interpretable paradigm for multi-modal emotion recognition in practical stress-detection scenarios. Homin Park, Ganghyun Kim, Jinsung Oh, Arnout Van Messem, Wesley De Neve |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets?abstractMachine learning (ML) research strongly relies on benchmarks in order to determine the relative effectiveness of newly proposed models. Recently, a number of prominent research effort argued that a number of models that improve the state-of-the-art by a small margin tend to do so by winning what they call a "benchmark lottery". An important benchmark in the field of machine learning and computer vision is the ImageNet where newly proposed models are often showcased based on their performance on this dataset. Given the large number of self-supervised learning (SSL) frameworks that has been proposed in the past couple of years each coming with marginal improvements on the ImageNet dataset, in this work, we evaluate whether those marginal improvements on ImageNet translate to improvements on similar datasets or not. To do so, we investigate twelve popular SSL frameworks on five ImageNet variants and discover that models that seem to perform well on ImageNet may experience significant performance declines on similar datasets. Specifically, state-of-the-art frameworks such as DINO and Swav, which are praised for their performance, exhibit substantial drops in performance while MoCo and Barlow Twins displays comparatively good results. As a result, we argue that otherwise good and desirable properties of models remain hidden when benchmarking is only performed on the ImageNet validation set, making us call for more adequate benchmarking. To avoid the "benchmark lottery" on ImageNet and to ensure a fair benchmarking process, we investigate the usage of a unified metric that takes into account the performance of models on other ImageNet variant datasets. Utku Ozbulak, Esla Timothy Anzaku, Solha Kang, Wesley De Neve, Joris Vankerschaver |
IJCNN | 4 |
| 2024 | Assessing the reliability of point mutation as data augmentation for deep learning with genomic dataabstractBACKGROUND: Deep neural networks (DNNs) have the potential to revolutionize our understanding and treatment of genetic diseases. An inherent limitation of deep neural networks, however, is their high demand for data during training. To overcome this challenge, other fields, such as computer vision, use various data augmentation techniques to artificially increase the available training data for DNNs. Unfortunately, most data augmentation techniques used in other domains do not transfer well to genomic data. RESULTS: Most genomic data possesses peculiar properties and data augmentations may significantly alter the intrinsic properties of the data. In this work, we propose a novel data augmentation technique for genomic data inspired by biology: point mutations. By employing point mutations as substitutes for codons, we demonstrate that our newly proposed data augmentation technique enhances the performance of DNNs across various genomic tasks that involve coding regions, such as translation initiation and splice site detection. CONCLUSION: Silent and missense mutations are found to positively influence effectiveness, while nonsense mutations and random mutations in non-coding regions generally lead to degradation. Overall, point mutation-based augmentations in genomic datasets present valuable opportunities for improving the accuracy and reliability of predictive models for DNA sequences. Hyun Jung Lee, Utku Ozbulak, Homin Park, Stephen Depuydt, Wesley De Neve, Joris Vankerschaver |
BMC Bioinform. | 5 |
| 2023 | Mutate and observe: utilizing deep neural networks to investigate the impact of mutations on translation initiationabstractMOTIVATION: The primary regulatory step for protein synthesis is translation initiation, which makes it one of the fundamental steps in the central dogma of molecular biology. In recent years, a number of approaches relying on deep neural networks (DNNs) have demonstrated superb results for predicting translation initiation sites. These state-of-the art results indicate that DNNs are indeed capable of learning complex features that are relevant to the process of translation. Unfortunately, most of those research efforts that employ DNNs only provide shallow insights into the decision-making processes of the trained models and lack highly sought-after novel biologically relevant observations. RESULTS: By improving upon the state-of-the-art DNNs and large-scale human genomic datasets in the area of translation initiation, we propose an innovative computational methodology to get neural networks to explain what was learned from data. Our methodology, which relies on in silico point mutations, reveals that DNNs trained for translation initiation site detection correctly identify well-established biological signals relevant to translation, including (i) the importance of the Kozak sequence, (ii) the damaging consequences of ATG mutations in the 5'-untranslated region, (iii) the detrimental effect of premature stop codons in the coding region, and (iv) the relative insignificance of cytosine mutations for translation. Furthermore, we delve deeper into the Beta-globin gene and investigate various mutations that lead to the Beta thalassemia disorder. Finally, we conclude our work by laying out a number of novel observations regarding mutations and translation initiation. AVAILABILITY AND IMPLEMENTATION: For data, models, and code, visit github.com/utkuozbulak/mutate-and-observe. Utku Ozbulak, Hyun Jung Lee, Jasper Zuallaert, Wesley De Neve, Stephen Depuydt, Joris Vankerschaver |
Bioinform. | 4 |
| 2023 | CRISPR-Cas-Docker: web-based in silico docking and machine learning-based classification of crRNAs with Cas proteinsabstractBACKGROUND: CRISPR-Cas-Docker is a web server for in silico docking experiments with CRISPR RNAs (crRNAs) and Cas proteins. This web server aims at providing experimentalists with the optimal crRNA-Cas pair predicted computationally when prokaryotic genomes have multiple CRISPR arrays and Cas systems, as frequently observed in metagenomic data. RESULTS: CRISPR-Cas-Docker provides two methods to predict the optimal Cas protein given a particular crRNA sequence: a structure-based method (in silico docking) and a sequence-based method (machine learning classification). For the structure-based method, users can either provide experimentally determined 3D structures of these macromolecules or use an integrated pipeline to generate 3D-predicted structures for in silico docking experiments. CONCLUSION: CRISPR-Cas-Docker addresses the need of the CRISPR-Cas community to predict RNA-protein interactions in silico by optimizing multiple stages of computation and evaluation, specifically for CRISPR-Cas systems. CRISPR-Cas-Docker is available at www.crisprcasdocker.org as a web server, and at https://github.com/hshimlab/CRISPR-Cas-Docker as an open-source tool. Homin Park, Jongbum Won, Yunseol Park, Esla Timothy Anzaku, Joris Vankerschaver, Arnout Van Messem, Wesley De Neve, Hyunjin Shim |
BMC Bioinform. | 7 |
| 2021 | Selection of Source Images Heavily Influences the Effectiveness of Adversarial Attacks
Utku Ozbulak, Esla Timothy Anzaku, Wesley De Neve, Arnout Van Messem |
BMVC | 3 |
| 2021 | ToxDL: deep learning using primary structure and domain embeddings for assessing protein toxicityabstractMOTIVATION: Genetically engineering food crops involves introducing proteins from other species into crop plant species or modifying already existing proteins with gene editing techniques. In addition, newly synthesized proteins can be used as therapeutic protein drugs against diseases. For both research and safety regulation purposes, being able to assess the potential toxicity of newly introduced/synthesized proteins is of high importance. RESULTS: In this study, we present ToxDL, a deep learning-based approach for in silico prediction of protein toxicity from sequence alone. ToxDL consists of (i) a module encompassing a convolutional neural network that has been designed to handle variable-length input sequences, (ii) a domain2vec module for generating protein domain embeddings and (iii) an output module that classifies proteins as toxic or non-toxic, using the outputs of the two aforementioned modules. Independent test results obtained for animal proteins and cross-species transferability results obtained for bacteria proteins indicate that ToxDL outperforms traditional homology-based approaches and state-of-the-art machine-learning techniques. Furthermore, through visualizations based on saliency maps, we are able to verify that the proposed network learns known toxic motifs. Moreover, the saliency maps allow for directed in silico modification of a sequence, thus making it possible to alter its predicted protein toxicity. AVAILABILITY AND IMPLEMENTATION: ToxDL is freely available at http://www.csbio.sjtu.edu.cn/bioinf/ToxDL/. The source code can be found at https://github.com/xypan1232/ToxDL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaoyong Pan, Jasper Zuallaert, Hong-Bin Shen, Elda Posada Campos, Denys O. Marushchak, Wesley De Neve |
Bioinform. | 7 |
| 2021 | Investigating the significance of adversarial attacks and their relation to interpretability for radar-based human activity recognition systemsabstractGiven their substantial success in addressing a wide range of computer vision challenges, Convolutional Neural Networks (CNNs) are increasingly being used in smart home applications, with many of these applications relying on the automatic recognition of human activities. In this context, low-power radar devices have recently gained in popularity as recording sensors, given that the usage of these devices allows mitigating a number of privacy concerns, a key issue when making use of conventional video cameras. Another concern that is often cited when designing smart home applications is the resilience of these applications against cyberattacks. It is, for instance, well-known that the combination of images and CNNs is vulnerable against adversarial examples, mischievous data points that force machine learning models to generate wrong classifications during testing time. In this paper, we investigate the vulnerability of radar-based CNNs to adversarial attacks, and where these radar-based CNNs have been designed to recognize human gestures. Through experiments with four unique threat models, we show that radar-based CNNs are susceptible to both white- and black-box adversarial attacks. We also expose the existence of an extreme adversarial attack case, where it is possible to change the prediction made by the radar-based CNNs by only perturbing the padding of the inputs, without touching the frames where the action itself occurs. Moreover, we observe that gradient-based attacks exercise perturbation not randomly, but on important features of the input data. We highlight these important features by making use of Grad-CAM, a popular neural network interpretability method, hereby showing the connection between adversarial perturbation and prediction interpretability. Utku Ozbulak, Baptist Vandersmissen, Azarakhsh Jalalvand, Ivo Couckuyt, Arnout Van Messem, Wesley De Neve |
Comput. Vis. Image Underst. | 6 |
| 2020 | Indoor human activity recognition using high-dimensional sensors and deep neural networks
Baptist Vandersmissen, Nicolas Knudde, Azarakhsh Jalalvand, Ivo Couckuyt, Tom Dhaene, Wesley De Neve |
Neural Comput. Appl. | 6 |
| 2020 | Perturbation analysis of gradient-based adversarial attacks
Utku Ozbulak, Manvel Gasparyan, Wesley De Neve, Arnout Van Messem |
Pattern Recognit. Lett. | 3 |
| 2019 | Not All Adversarial Examples Require a Complex Defense: Identifying Over-optimized Adversarial Examples with IQR-based Logit ThresholdingabstractDetecting adversarial examples currently stands as one of the biggest challenges in the field of deep learning. Adversarial attacks, which produce adversarial examples, increase the prediction likelihood of a target class for a particular data point. During this process, the adversarial example can be further optimized, even when it has already been wrongly classified with 100% confidence, thus making the adversarial example even more difficult to detect. For this kind of adversarial examples, which we refer to as over-optimized adversarial examples, we discovered that the logits of the model provide solid clues on whether the data point at hand is adversarial or genuine. In this context, we first discuss the masking effect of the softmax function for the prediction made and explain why the logits of the model are more useful in detecting over-optimized adversarial examples. To identify this type of adversarial examples in practice, we propose a non-parametric and computationally efficient method which relies on interquartile range, with this method becoming more effective as the image resolution increases. We support our observations throughout the paper with detailed experiments for different datasets (MNIST, CIFAR-10, and ImageNet) and several architectures. Utku Ozbulak, Arnout Van Messem, Wesley De Neve |
IJCNN | 3 |
| 2019 | Impact of Adversarial Examples on Deep Learning Models for Biomedical Image Segmentation
Utku Ozbulak, Arnout Van Messem, Wesley De Neve |
MICCAI (2) | 3 |
| 2018 | Computer-Aided Diagnosis and Localization of Glaucoma Using Deep Learning
Mijung Kim, Homin Park, Jasper Zuallaert, Olivier Janssens, Sofie Van Hoecke, Wesley De Neve |
BIBM | 6 |
| 2018 | Explaining Character-Aware Neural Networks for Word-Level Prediction: Do They Discover Linguistic Rules?abstractCharacter-level features are currently used in different neural network-based natural language processing algorithms.However, little is known about the character-level patterns those models learn.Moreover, models are often compared only quantitatively while a qualitative analysis is missing.In this paper, we investigate which character-level patterns neural networks learn and if those patterns coincide with manually-defined word segmentations and annotations.To that end, we extend the contextual decomposition (Murdoch et al., 2018) technique to convolutional neural networks which allows us to compare convolutional neural networks and bidirectional long short-term memory networks.We evaluate and compare these models for the task of morphological tagging on three morphologically different languages and show that these models implicitly discover understandable linguistic rules. Fréderic Godin, Kris Demuynck, Joni Dambre, Wesley De Neve, Thomas Demeester |
EMNLP | 4 |
| 2018 | AQUa: an adaptive framework for compression of sequencing quality scores with random access functionalityabstractMotivation: The past decade has seen the introduction of new technologies that significantly lowered the cost of genome sequencing. As a result, the amount of genomic data that must be stored and transmitted is increasing exponentially. To mitigate storage and transmission issues, we introduce a framework for lossless compression of quality scores. Results: This article proposes AQUa, an adaptive framework for lossless compression of quality scores. To compress these quality scores, AQUa makes use of a configurable set of coding tools, extended with a Context-Adaptive Binary Arithmetic Coding scheme. When benchmarking AQUa against generic single-pass compressors, file sizes are reduced by up to 38.49% when comparing with GNU Gzip and by up to 6.48% when comparing with 7-Zip at the Ultra Setting, while still providing support for random access. When comparing AQUa with the purpose-built, single-pass, and state-of-the-art compressor SCALCE, which does not support random access, file sizes are reduced by up to 21.14%. When comparing AQUa with the purpose-built, dual-pass, and state-of-the-art compressor QVZ, which does not support random access, file sizes are larger by 6.42-33.47%. However, for one test file, the file size is 0.38% smaller, illustrating the strength of our single-pass compression framework. This work has been spurred by the current activity on genomic information representation (MPEG-G) within the ISO/IEC SC29/WG11 technical committee. Availability and implementation: The software is available on Github: https://github.com/tparidae/AQUa. Contact: [email protected]. Tom Paridaens, Glenn Van Wallendael, Wesley De Neve, Peter Lambert |
Bioinform. | 3 |
| 2018 | SpliceRover: interpretable convolutional neural networks for improved splice site predictionabstractMotivation: During the last decade, improvements in high-throughput sequencing have generated a wealth of genomic data. Functionally interpreting these sequences and finding the biological signals that are hallmarks of gene function and regulation is currently mostly done using automated genome annotation platforms, which mainly rely on integrated machine learning frameworks to identify different functional sites of interest, including splice sites. Splicing is an essential step in the gene regulation process, and the correct identification of splice sites is a major cornerstone in a genome annotation system. Results: In this paper, we present SpliceRover, a predictive deep learning approach that outperforms the state-of-the-art in splice site prediction. SpliceRover uses convolutional neural networks (CNNs), which have been shown to obtain cutting edge performance on a wide variety of prediction tasks. We adapted this approach to deal with genomic sequence inputs, and show it consistently outperforms already existing approaches, with relative improvements in prediction effectiveness of up to 80.9% when measured in terms of false discovery rate. However, a major criticism of CNNs concerns their 'black box' nature, as mechanisms to obtain insight into their reasoning processes are limited. To facilitate interpretability of the SpliceRover models, we introduce an approach to visualize the biologically relevant information learnt. We show that our visualization approach is able to recover features known to be important for splice site prediction (binding motifs around the splice site, presence of polypyrimidine tracts and branch points), as well as reveal new features (e.g. several types of exclusion patterns near splice sites). Availability and implementation: SpliceRover is available as a web service. The prediction tool and instructions can be found at http://bioit2.irc.ugent.be/splicerover/. Supplementary information: Supplementary data are available at Bioinformatics online. Jasper Zuallaert, Fréderic Godin, Mijung Kim, Arne Soete, Yvan Saeys, Wesley De Neve |
Bioinform. | 6 |
| 2018 | On the application of reservoir computing networks for noisy image recognition
Azarakhsh Jalalvand, Kris Demuynck, Wesley De Neve, Jean-Pierre Martens |
Neurocomputing | 3 |
| 2018 | Dual Rectified Linear Units (DReLUs): A replacement for tanh activation functions in Quasi-Recurrent Neural Networks
Fréderic Godin, Jonas Degrave, Joni Dambre, Wesley De Neve |
Pattern Recognit. Lett. | 4 |
| 2018 | Indoor Person Identification Using a Low-Power FMCW RadarabstractContemporary surveillance systems mainly use video cameras as their primary sensor. However, video cameras possess fundamental deficiencies, such as the inability to handle low-light environments, poor weather conditions, and concealing clothing. In contrast, radar devices are able to sense in pitch-dark environments and to see through walls. In this paper, we investigate the use of micro-Doppler (MD) signatures retrieved from a low-power radar device to identify a set of persons based on their gait characteristics. To that end, we propose a robust feature learning approach based on deep convolutional neural networks. Given that we aim at providing a solution for a real-world problem, people are allowed to walk around freely in two different rooms. In this setting, the IDentification with Radar data data set is constructed and published, consisting of 150 min of annotated MD data equally spread over five targets. Through experiments, we investigate the effectiveness of both the Doppler and time dimension, showing that our approach achieves a classification error rate of 24.70% on the validation set and 21.54% on the test set for the five targets used. When experimenting with larger time windows, we are able to further lower the error rate. Baptist Vandersmissen, Nicolas Knudde, Azarakhsh Jalalvand, Ivo Couckuyt, André Bourdoux, Wesley De Neve, Tom Dhaene |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2017 | Interpretable convolutional neural networks for effective translation initiation site predictionabstractThanks to rapidly evolving sequencing techniques, the amount of genomic data at our disposal is growing increasingly large. Determining the gene structure is a fundamental requirement to effectively interpret gene function and regulation. An important part in that determination process is the identification of translation initiation sites. In this paper, we propose a novel approach for automatic prediction of translation initiation sites, leveraging convolutional neural networks that allow for automatic feature extraction. Our experimental results demonstrate that we are able to improve the state-of-the-art approaches with a decrease of 75.2% in false positive rate and with a decrease of 24.5% in error rate on chosen datasets. Furthermore, an in-depth analysis of the decision-making process used by our predictive model shows that our neural network implicitly learns biologically relevant features from scratch, without any prior knowledge about the problem at hand, such as the Kozak consensus sequence, the influence of stop and start codons in the sequence and the presence of donor splice site patterns. In summary, our findings yield a better understanding of the internal reasoning of a convolutional neural network when applying such a neural network to genomic data. Jasper Zuallaert, Mijung Kim, Yvan Saeys, Wesley De Neve |
BIBM | 4 |
| 2017 | AFRESh: an adaptive framework for compression of reads and assembled sequences with random access functionalityabstractMOTIVATION: The past decade has seen the introduction of new technologies that lowered the cost of genomic sequencing increasingly. We can even observe that the cost of sequencing is dropping significantly faster than the cost of storage and transmission. The latter motivates a need for continuous improvements in the area of genomic data compression, not only at the level of effectiveness (compression rate), but also at the level of functionality (e.g. random access), configurability (effectiveness versus complexity, coding tool set …) and versatility (support for both sequenced reads and assembled sequences). In that regard, we can point out that current approaches mostly do not support random access, requiring full files to be transmitted, and that current approaches are restricted to either read or sequence compression. RESULTS: We propose AFRESh, an adaptive framework for no-reference compression of genomic data with random access functionality, targeting the effective representation of the raw genomic symbol streams of both reads and assembled sequences. AFRESh makes use of a configurable set of prediction and encoding tools, extended by a Context-Adaptive Binary Arithmetic Coding scheme (CABAC), to compress raw genetic codes. To the best of our knowledge, our paper is the first to describe an effective implementation CABAC outside of its' original application. By applying CABAC, the compression effectiveness improves by up to 19% for assembled sequences and up to 62% for reads. By applying AFRESh to the genomic symbols of the MPEG genomic compression test set for reads, a compression gain is achieved of up to 51% compared to SCALCE, 42% compared to LFQC and 44% compared to ORCOM. When comparing to generic compression approaches, a compression gain is achieved of up to 41% compared to GNU Gzip and 22% compared to 7-Zip at the Ultra setting. Additionaly, when compressing assembled sequences of the Human Genome, a compression gain is achieved up to 34% compared to GNU Gzip and 16% compared to 7-Zip at the Ultra setting. AVAILABILITY AND IMPLEMENTATION: A Windows executable version can be downloaded at https://github.com/tparidae/AFresh . CONTACT: [email protected]. Tom Paridaens, Glenn Van Wallendael, Wesley De Neve, Peter Lambert |
Bioinform. | 3 |
| 2017 | Effective and efficient human action recognition using dynamic frame skipping and trajectory rejection
Jeong-Jik Seo, Hyungil Kim, Wesley De Neve, Yong Man Ro |
Image Vis. Comput. | 3 |
| 2016 | Leveraging CABAC for No-Reference Compression of Genomic Data with Random Access SupportabstractIn previous work, the authors developed a modular no-reference framework that compresses FASTA files by applying a predict-and-residue method, as used in video coding. We extended this framework with support for Context-Adaptive Binary Arithmetic Coding (CABAC), while at the same time preserving random access functionality and offering support for the full IUB/IUPAC nucleic acid codes alphabet. Tom Paridaens, Jens Panneel, Wesley De Neve, Peter Lambert, Rik Van de Walle |
DCC | 3 |
| 2016 | Normalized Semantic Web Distance
Tom De Nies, Christian Beecks, Fréderic Godin, Wesley De Neve, Grzegorz Stepien, Dörthe Arndt, Laurens De Vocht, Ruben Verborgh, Thomas Seidl 0001, Erik Mannens, Rik Van de Walle |
ESWC | 4 |
| 2016 | Towards using Reservoir Computing Networks for noise-robust image recognitionabstractReservoir Computing Network (RCN) is a special type of the single layer recurrent neural networks, in which the input and the recurrent connections are randomly generated and only the output weights are trained. Besides the ability to process temporal information, the key points of RCN are easy training and robustness against noise. Recently, we introduced a simple strategy to tune the parameters of RCN resulted in an effective and noise-robust RCN-based model for speech recognition. The aim of this work is to extend that study to the field of image processing. In particular, we investigate the potential of RCNs in achieving a competitive performance on the well-known MNIST dataset by following the aforementioned parameter optimizing strategy. Moreover, we achieve good noise robust recognition by utilizing such a network to denoise images and supplying them to a recognizer that is solely trained on clean images. The conducted experiments demonstrate that the proposed RCN-based handwritten digit recognizer achieves an error rate of 0.81 percent on the clean test data of the MNIST benchmark and that the proposed RCN-based denoiser can effectively reduce the error rate on the various types of noise. Azarakhsh Jalalvand, Wesley De Neve, Rik Van de Walle, Jean-Pierre Martens |
IJCNN | 2 |
| 2016 | An Automated End-To-End Pipeline for Fine-Grained Video Annotation using Deep Neural NetworksabstractThe searchability of video content is often limited to the descriptions authors and/or annotators care to provide. The level of description can range from absolutely nothing to fine-grained annotations at the level of frames. Based on these annotations, certain parts of the video content are more searchable than others. Baptist Vandersmissen, Lucas Sterckx, Thomas Demeester, Azarakhsh Jalalvand, Wesley De Neve, Rik Van de Walle |
ICMR | 5 |
| 2016 | Towards robust and reliable multimedia analysis through semantic integration of services
Ben De Meester, Ruben Verborgh, Pieter Pauwels, Wesley De Neve, Erik Mannens, Rik Van de Walle |
Multim. Tools Appl. | 4 |
| 2015 | Hybrid Books for Interactive Digital Storytelling: Connecting Story Entities and Emotions to Smart Environments
Hajar Ghaem Sigarchian, Ben De Meester, Frank Salliau, Wesley De Neve, Sara Logghe, Ruben Verborgh, Erik Mannens, Rik Van de Walle, Dimitri Schuurman |
ICIDS | 4 |
| 2015 | Hyperspectral Image Classification with Convolutional Neural NetworksabstractHyperspectral image (HSI) classification is one of the most widely used methods for scene analysis from hyperspectral imagery. In the past, many different engineered features have been proposed for the HSI classification problem. In this paper, however, we propose a feature learning approach for hyperspectral image classification based on convolutional neural networks (CNNs). The proposed CNN model is able to learn structured features, roughly resembling different spectral band-pass filters, directly from the hyperspectral input data. Our experimental results, conducted on a commonly-used remote sensing hyperspectral dataset, show that the proposed method provides classification results that are among the state-of-the-art, without using any prior knowledge or engineered features. Viktor Slavkovikj, Steven Verstockt, Wesley De Neve, Sofie Van Hoecke, Rik Van de Walle |
ACM Multimedia | 3 |
| 2014 | Sub-sampled dictionaries for coarse-to-fine sparse representation-based human action recognitionabstractAutomatic human action recognition is a core functionality of systems for video surveillance and human-object interaction. However, the diverse nature of human actions and the noisy nature of most video content make it difficult to achieve effective human action recognition. To overcome the aforementioned problems, Sparse Representation (SR) has recently attracted substantial research attention. However, although SR-based approaches have proven to be reasonably effective, the computational complexity of the testing stage prohibits their usage by applications requiring support for real-time operation and a vast number of human action classes. In this paper, we propose a novel method for human action recognition, leveraging coarse-to-fine sparse representations that have been obtained through dictionary sub-sampling. Comparative experimental results obtained for the UCF50 dataset demonstrate that the proposed method is able to achieve efficient human action recognition, at no substantial loss in recognition accuracy. Hyunseok Min, Jeong-Jik Seo, Wesley De Neve, Yong Man Ro |
ICME | 4 |
| 2014 | Image-Based Road Type ClassificationabstractThe ability to automatically determine the road type from sensor data is of great significance for automatic annotation of routes and autonomous navigation of robots and vehicles. In this paper, we present a novel algorithm for content-based road type classification from images. The proposed method learns discriminative features from training data in an unsupervised manner, thus not requiring domain-specific feature engineering. This is an advantage over related road surface classification algorithms which are only able to make a distinction between pre-specified uniform terrains. In order to evaluate the proposed approach, we have constructed a challenging road image dataset of 20,000 samples from real-world road images in the paved and unpaved road classes. Experimental results on this dataset show that the proposed algorithm can achieve state-of-the-art performance in road type classification. Viktor Slavkovikj, Steven Verstockt, Wesley De Neve, Sofie Van Hoecke, Rik Van de Walle |
ICPR | 3 |
| 2014 | Visually weighted neighbor voting for image tag relevance learning
Sihyoung Lee, Wesley De Neve, Yong Man Ro |
Multim. Tools Appl. | 2 |
| 2013 | Sparse Representation-Based Human Action Recognition Using an Action Region-Aware DictionaryabstractAutomatic human action recognition is a core functionality of systems for video surveillance and human-object interaction. Conventional vision-based systems for human action recognition require the use of segmentation in order to achieve an acceptable level of recognition effectiveness. However, generic techniques for automatic segmentation are currently not available yet. Therefore, in this paper, we propose a novel sparse representation-based method for human action recognition, taking advantage of the observation that, although the location and size of the action region in a test video clip is unknown, the construction of a dictionary can leverage information about the location and size of action regions in training video clips. That way, we are able to segment, implicitly, action and context information in a test video clip, thus improving the effectiveness of classification. That way, we are also able to develop a context-adaptive classification strategy. As shown by comparative experimental results obtained for the UCF Sports Action data set, the proposed method facilitates effective human action recognition, even when testing does not rely on explicit segmentation. Hyunseok Min, Wesley De Neve, Yong Man Ro |
ISM | 2 |
| 2013 | Improved License Plate Recognition for Low-Resolution CCTV Forensics by Integrating Sparse Representation-Based Super-Resolution
Hyunseok Min, Seung-Ho Lee, Wesley De Neve, Yong Man Ro |
IWDW | 3 |
| 2013 | Towards fusion of collective knowledge and audio-visual content features for annotating broadcast videoabstractBroadcasters produce vast collections of video content. However, the lack of fine-grained annotations makes it difficult to retrieve video fragments of interest from these vast collections. Indeed, manual annotation of video content is labour-intensive and time-consuming. Moreover, the applicability of algorithms for automatic annotation of video content is limited, given that too many prerequisites need to be fulfilled and that a lot of concepts are unidentifiable. At the same time, people are using social media to share their thoughts about the content they view on television. Therefore, in this Ph.D. research, we plan to investigate novel machine learning-based approaches towards the task of fine-grained annotation of broadcast video content, fusing the collective knowledge present in social media with the output of audio-visual content analysis algorithms. Fréderic Godin, Wesley De Neve, Rik Van de Walle |
ICMR | 2 |
| 2012 | Video Copy Detection Using Inclined Video Tomography and Bag-of-Visual-WordsabstractTechniques for video fingerprinting are helpful in managing vast libraries of video clips. Recent advances have shown that video tomography and Bag-of-Visual-Words (BoVW) can be successfully used for the purpose of video fingerprinting. In this paper, we introduce a novel video signature (i.e., a novel video fingerprint) that takes advantage of both video tomography and BoVW. Specifically, the proposed video signature is created by first extracting inclined tomography images from the video content, and by subsequently applying the BoVW approach to the inclined tomography images obtained. The key to our approach is that we make the angle of inclination of the tomography images dependent on the amount of motion in the video content. That way, the proposed video signature is able to capture both spatial and temporal information. Experimental results obtained for the publicly available TREVID-2009 video set indicate that video copy detection by means of the proposed video signature is robust against spatial and temporal transformations. Hyunseok Min, Semin Kim 0001, Wesley De Neve, Yong Man Ro |
ICME | 3 |
| 2012 | Parallel deblocking filtering in H.264/AVC using multiple CPUs and GPUsabstractDeblocking filtering in the H.264/AVC standard is a computationally complex process because of the filter's high content adaptivity. Furthermore, the deblocking filter introduces a significant number of data dependencies, making parallel processing not obvious. Our previous works analyzed the dependencies of the filter and proposed a massively-parallel implementation, specifically tailored for execution on a single GPU. In this paper, we extend this work by proposing a parallel processing scheme for accelerating deblocking filtering using multiple CPU cores or GPUs. This scheme allows for standard-compliant filtering, regardless of slice configuration. Results show that our multi-GPU implementation using our proposed scheme achieves faster-than real-time deblocking at over 3794 frames per second for 1080p video pictures by using three GPUs. A multi-core CPU implementation using 8 CPU cores allows 1080p deblocking filtering of up to 695 frames per second. Bart Pieters, Charles-Frederik Hollemeersch, Jan De Cock, Wesley De Neve, Peter Lambert, Rik Van de Walle |
ACM Multimedia | 4 |
| 2012 | Near-Duplicate Video Clip Detection Using Model-Free Semantic Concept Detection and Adaptive Semantic Distance MeasurementabstractMotivated by the observation that content transformations tend to preserve the semantic information conveyed by video clips, this paper introduces a novel technique for near-duplicate video clip (NDVC) detection, leveraging model-free semantic concept detection and adaptive semantic distance measurement. In particular, model-free semantic concept detection is realized by taking advantage of the collective knowledge in an image folksonomy (which is an unstructured collection of user-contributed images and tags), facilitating the use of an unrestricted concept vocabulary. Adaptive semantic distance measurement is realized by means of the signature quadratic form distance (SQFD), making it possible to flexibly measure the similarity between video shots that contain a varying number of semantic concepts, and where these semantic concepts may also differ in terms of relevance and nature. Experimental results obtained for the MIRFLICKR-25000 image set (used as a source of collective knowledge) and the TRECVID 2009 video set (used to create query and reference video clips) demonstrate that model-free semantic concept detection and SQFD can be successfully used for the purpose of identifying NDVCs. Hyunseok Min, Wesley De Neve, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Improving image tag recommendation using favorite image contextabstractTag recommendation allows mitigating the amount of user effort needed to annotate images. Assuming that favorite images and their associated tags are indicative of the visual and topical interests of users, this paper proposes a personalized image tag recommendation technique that makes use of favorite image context. Specifically, to recommend tags for a newly uploaded image, we propose to take advantage of the tags assigned to favorite images of the user who uploaded the image, fusing tag statistics and visual similarity. Experimental results obtained for images and tags retrieved from Flickr compare the use of favorite image context to the use of personal and collective context for the purpose of tag recommendation, showing that the use of favorite image context is promising. Wonyong Eom, Sihyoung Lee, Wesley De Neve, Yong Man Ro |
ICIP | 3 |
| 2011 | Towards a better understanding of model-free semantic concept detection for annotation and near-duplicate video clip detectionabstractGiven the observation that content transformations tend to preserve semantic information, we demonstrated in previous research that model-free semantic concept detection can be successfully leveraged for identifying NDVCs. In this paper, we seek a better understanding of the usefulness of model-free semantic concept detection for both the task of annotation and NDVC detection. In particular, through extensive experiments, we demonstrate that the problem of detecting semantic concepts for the goal of identifying NDVCs is more relaxed than the problem of detecting semantic concepts for annotation purposes: whereas incorrectly detected semantic concepts negatively affect the effectiveness of annotation, they do not negatively affect the effectiveness of NDVC detection, as long as the same incorrect semantic concepts are detected for both the reference and near-duplicate video clips. This observation has practical implications for the design of a video management system that makes use of model-free semantic concept detection for both the purpose of annotation and NDVC detection. Hyunseok Min, Wesley De Neve, Yong Man Ro |
ICIP | 3 |
| 2011 | Leveraging an image folksonomy and the Signature Quadratic Form Distance for semantic-based detection of near-duplicate video clipsabstractBeing able to detect near-duplicate video clips (NDVCs) is a prerequisite for a plethora of multimedia applications. Given the observation that content transformations tend to preserve semantic information, techniques for NDVC detection may benefit from the use of a semantic approach. This paper discusses how an image folksonomy (i.e., community-contributed images and metadata) and the Signature Quadratic Form Distance (SQFD) can be leveraged for the purpose of identifying NDVCs. Experimental results obtained for the MIRFLICKR-25000 image set and the TRECVID 2009 video set indicate that an image folksonomy and SQFD can be successfully used for detecting NDVCs. In addition, our findings show that model-free NDVC detection (i.e., NDVC detection using an image folksonomy) has a higher semantic coverage than model-based NDVC detection (i.e., NDVC detection using the VIREO-374 semantic concept models). Hyunseok Min, Wesley De Neve, Yong Man Ro |
ICME | 3 |
| 2011 | Contribution of Non-scrambled Chroma Information in Privacy-Protected Face Images to Privacy Leakage
Hosik Sohn, Dohyoung Lee, Wesley De Neve, Konstantinos N. Plataniotis, Yong Man Ro |
IWDW | 3 |
| 2011 | Bimodal fusion of low-level visual features and high-level semantic features for near-duplicate video clip detection
Hyunseok Min, Wesley De Neve, Yong Man Ro |
Signal Process. Image Commun. | 3 |
| 2011 | Parallel Deblocking Filtering in MPEG-4 AVC/H.264 on Massively Parallel ArchitecturesabstractThe deblocking filter in the MPEG-4 AVC/H.264 standard is computationally complex because of its high content adaptivity, resulting in a significant number of data dependencies. These data dependencies interfere with parallel filtering of multiple macroblocks (MBs) on massively parallel architectures. In this letter, we introduce a novel MB partitioning scheme for concurrent deblocking in the MPEG-4 AVC/H.264 standard, based on our idea of deblocking filter independency, a corrected version of the limited error propagation effect proposed in the letter. Our proposed scheme enables concurrent MB deblocking of luma samples with limited synchronization effort, independently of slice configuration, and is compliant with the MPEG-4 H.264/AVC standard. We implemented the method on the massively parallel architecture of the graphics processing unit (GPU). Experimental results show that our GPU implementation achieves faster-than real-time deblocking at 1309 frames per second for 1080p video pictures. Both software-based deblocking filters and state-of-the-art GPU-enabled algorithms are outperformed in terms of speed by factors up to 10.2 and 19.5, respectively, for 1080p video pictures. Bart Pieters, Charles-Frederik Hollemeersch, Jan De Cock, Peter Lambert, Wesley De Neve, Rik Van de Walle |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | Privacy Protection in Video Surveillance Systems: Analysis of Subband-Adaptive Scrambling in JPEG XRabstractThis paper discusses a privacy-protected video surveillance system that makes use of JPEG extended range (JPEG XR). JPEG XR offers a low-complexity solution for the scalable coding of high-resolution images. To address privacy concerns, face regions are detected and scrambled in the transform domain, taking into account the quality and spatial scalability features of JPEG XR. Experiments were conducted to investigate the performance of our surveillance system, considering visual distortion, bit stream overhead, and security aspects. Our results demonstrate that subband-adaptive scrambling is able to conceal privacy-sensitive face regions with a feasible level of protection. In addition, our results show that subband-adaptive scrambling of face regions outperforms subband-adaptive scrambling of frames in terms of coding efficiency, except when low video bit rates are in use. Hosik Sohn, Wesley De Neve, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Collaborative Face Recognition for Improved Face Annotation in Personal Photo Collections Shared on Online Social NetworksabstractUsing face annotation for effective management of personal photos in online social networks (OSNs) is currently of considerable practical interest. In this paper, we propose a novel collaborative face recognition (FR) framework, improving the accuracy of face annotation by effectively making use of multiple FR engines available in an OSN. Our collaborative FR framework consists of two major parts: selection of FR engines and merging (or fusion) of multiple FR results. The selection of FR engines aims at determining a set of personalized FR engines that are suitable for recognizing query face images belonging to a particular member of the OSN. For this purpose, we exploit both social network context in an OSN and social context in personal photo collections. In addition, to take advantage of the availability of multiple FR results retrieved from the selected FR engines, we devise two effective solutions for merging FR results, adopting traditional techniques for combining multiple classifier results. Experiments were conducted using 547 991 personal photos collected from an existing OSN. Our results demonstrate that the proposed collaborative FR method is able to significantly improve the accuracy of face annotation, compared to conventional FR approaches that only make use of a single FR engine. Further, we demonstrate that our collaborative FR framework has a low computational cost and comes with a design that is suited for deployment in a decentralized OSN. Wesley De Neve, Konstantinos N. Plataniotis, Yong Man Ro |
IEEE Trans. Multim. | 2 |
| 2010 | Exploiting collective knowledge in an image folksonomy for semantic-based near-duplicate video detectionabstractAn increasing number of duplicates and near-duplicates can be found on websites for video sharing. These duplicates and near-duplicates often infringe copyright or clutter search results. Consequently, a high need exists for techniques that allow identifying duplicates and near-duplicates. In this paper, we propose a semantic-based approach towards the task of identifying near-duplicates. Our approach makes use of semantic video signatures that are constructed by detecting semantic concepts along the temporal axis of video sequences. Specifically, we make use of an image folksonomy (i.e., a set of user-contributed images annotated with user-supplied tags) to detect semantic concepts in video sequences, making it possible to exploit an unrestricted concept vocabulary. Comparative experiments using the MUSCLE-VCD-2007 dataset and folksonomy images retrieved from Flickr show that our approach is successful in identifying near-duplicates. Hyunseok Min, Wesley De Neve, Yong Man Ro |
ICIP | 2 |
| 2010 | Image tag refinement along the 'what' dimension using tag categorization and neighbor votingabstractOnline sharing of images is increasingly becoming popular, resulting in the availability of vast collections of user-contributed images that have been annotated with usersupplied tags. However, user-supplied tags are often not related to the actual image content, affecting the performance of multimedia applications that rely on tag-based retrieval of user-contributed images. This paper proposes a modular approach towards tag refinement, taking into account the nature of tags. First, tags are automatically categorized in five categories using WordNet: ‘where’, ‘when’, ‘who’, ‘what’, and ‘how’. Next, as a start towards a full implementation of our modular tag refinement approach, we use neighbor voting to learn the relevance of tags along the ‘what’ dimension. Our experimental results show that the proposed tag refinement technique is able to successfully differentiate correct tags from noisy tags along the ‘what’ dimension. In addition, we demonstrate that the proposed tag refinement technique is able to improve the effectiveness of image tag recommendation for non-tagged images. Sihyoung Lee, Wesley De Neve, Yong Man Ro |
ICME | 2 |
| 2010 | Towards using semantic features for near-duplicate video detectionabstractAn increasing number of near-duplicate video clips (NDVCs) can be found on websites for video sharing. These NDVCs often infringe copyright or clutter search results. Consequently, a high need exists for techniques that allow identifying NDVCs. NDVC detection techniques represent a video clip with a unique set of features. Conventional video signatures typically make use of low-level visual features (e.g., color histograms, local interest points). However, low-level visual features are sensitive to transformations of the video content. In this paper, given the observation that transformations preserve the semantic information in the video content, we study the use of semantic features for the purpose of identifying NDVCs. Experimental results obtained for the MUSCLE-VCD-2007 dataset indicate that semantic features have a high level of robustness against transformations and different keyframe selection strategies. In addition, when relying on the temporal variation of semantic features, semantic video signatures are characterized by a high degree of uniqueness, even when a vocabulary with a low number of semantic concepts is in use (for a query video clip that is sufficiently long). Hyunseok Min, Wesley De Neve, Yong Man Ro |
ICME | 2 |
| 2010 | Semantic Concept Detection for User-Generated Video Content Using a Refined Image Folksonomy
Hyunseok Min, Sihyoung Lee, Wesley De Neve, Yong Man Ro |
MMM | 3 |
| 2010 | Format-independent and metadata-driven media resource adaptation using semantic web technologies
Davy Van Deursen, Wim Van Lancker, Sarah De Bruyne, Wesley De Neve, Erik Mannens, Rik Van de Walle |
Multim. Syst. | 4 |
| 2010 | NinSuna: a fully integrated platform for format-independent multimedia content adaptation and delivery using Semantic Web technologies
Davy Van Deursen, Wim Van Lancker, Wesley De Neve, Tom Paridaens, Erik Mannens, Rik Van de Walle |
Multim. Tools Appl. | 3 |
| 2010 | MAP-based image tag recommendation using a visual folksonomy
Sihyoung Lee, Wesley De Neve, Konstantinos N. Plataniotis, Yong Man Ro |
Pattern Recognit. Lett. | 2 |
| 2010 | Tag refinement in an image folksonomy using visual similarity and tag co-occurrence statistics
Sihyoung Lee, Wesley De Neve, Yong Man Ro |
Signal Process. Image Commun. | 2 |
| 2010 | Automatic Face Annotation in Personal Photo Collections Using Context-Based Unsupervised Clustering and Face Information FusionabstractIn this paper, a novel face annotation framework is proposed that systematically leverages context information such as situation awareness information with current face recognition (FR) solutions. In particular, unsupervised situation and subject clustering techniques have been developed that are aided by context information. Situation clustering groups together photos that are similar in terms of capture time and visual content, allowing for the reliable use of visual context information during subject clustering. The aim of subject clustering is to merge multiple face images that belong to the same individual. To take advantage of the availability of multiple face images for a particular individual, we propose effective FR methods that are based on face information fusion strategies. The performance of the proposed annotation method has been evaluated using a variety of photo sets. The photo sets were constructed using 1385 photos from the MPEG-7 Visual Core Experiment 3 (VCE-3) data set and approximately 20000 photos collected from well-known photo-sharing websites. The reported experimental results show that the proposed face annotation method significantly outperforms traditional face annotation solutions at no additional computational cost, with accuracy gains of up to 25% for particular cases. Wesley De Neve, Yong Man Ro, Konstantinos N. Plataniotis |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Privacy Protection in Video Surveillance Systems Using Scalable Video CodingabstractThanks to high-speed Internet access and feature-rich mobile devices, the demand for ubiquitous and secure surveillance systems has increased. In this paper, we propose a privacy-protected video surveillance system that makes use of scalable video coding (SVC). SVC can be used to fulfill the requirement of omnipresence. Further, to address privacy concerns, we detect face regions and subsequently scramble these regions-of-interest (ROIs) in the compressed domain. To demonstrate the feasibility of the proposed video surveillance system, simulation results are provided. The results show that our system is able to provide a good level of security, while offering access to surveillance video content in heterogeneous usage environments. Hosik Sohn, Esla Timothy Anzaku, Wesley De Neve, Yong Man Ro, Konstantinos N. Plataniotis |
AVSS | 3 |
| 2009 | Semantic annotation of personal video content using an image folksonomyabstractThe increasing popularity of user-generated content (UGC) requires effective annotation techniques in order to facilitate precise content search and retrieval. In this paper, we propose a new approach for the semantic annotation of personal video content, taking advantage of user-contributed tags available in an image folksonomy. Video shots and folksonomy images are first represented by a semantic vector. Next, the semantic vectors are used to measure the semantic similarity between each video shot and the folksonomy images. Tags assigned to semantically similar folksonomy images are then used to annotate the video shots. To verify the effectiveness of the proposed annotation method, experiments were performed with video sequences retrieved from YouTube and images downloaded from Flickr. Our experimental results demonstrate that the proposed method is able to successfully annotate personal video content with user-contributed tags retrieved from an image folksonomy. In addition, the size of our tag vocabulary is significantly higher than the size of the tag vocabulary used by conventional annotation methods. Hyunseok Min, Wesley De Neve, Yong Man Ro, Konstantinos N. Plataniotis |
ICIP | 3 |
| 2009 | Near-Duplicate Video Detection Using Temporal Patterns of Semantic ConceptsabstractMethods for video copy detection are typically based on the use of low-level visual features. However, low-level features may vary significantly for near-duplicates, which are video sequences that have been the subject of spatial or temporal modifications. As such, the use of low-level visual features may be inadequate for detecting near-duplicates. In this paper, we present a new video copy detection method that aims to identify near-duplicates for a given query video sequence. More specifically, the proposed method is based on identifying semantic concepts along the temporal axis of a particular video sequence, resulting in the construction of a so-called semantic video signature. The semantic video signature is then used for the purpose of similarity measurement. The main advantage of the proposed method lies in the fact that the presence of semantic concepts is highly robust to spatial and temporal video transformations. Our experimental results show that the use of a semantic video signature allows for the efficient and effective detection of near-duplicates. Hyunseok Min, Wesley De Neve, Yong Man Ro |
ISM | 3 |
| 2009 | A Statistical and Iterative Method for Data Hiding in Palette-Based Images
Semin Kim 0001, Wesley De Neve, Yong Man Ro |
IWDW | 2 |
| 2009 | Region-of-interest scrambling for scalable surveillance video using JPEG XRabstractPresent-day video surveillance systems are often required not to intrude upon the privacy of the general public. In this paper, we discuss a privacy-protected video surveillance system that makes use of the JPEG XR standard. This standard offers a low-complexity solution for the scalable coding of high-resolution images. To address privacy concerns, face regions are detected and subsequently scrambled in the transform domain, taking into account the spatial and quality scalability features of JPEG XR. A number of experiments were conducted in order to investigate the efficiency of our video surveillance system, considering bit stream overhead and security aspects. Hosik Sohn, Wesley De Neve, Yong Man Ro |
ACM Multimedia | 2 |
| 2009 | Semantic adaptation of synchronized multimedia streams in a format-independent wayabstractToday, users often want to consume personalized versions of multimedia content (e.g., a selection of the most interesting video scenes). These user preferences can be fulfilled by performing semantic adaptations, which are typically guided by semantic metadata. In this paper, we discuss how high-level semantic adaptations along the temporal axis can be performed using format-independent adaptation engines. Further, the impact of semantic adaptations applied to synchronized multimedia streams (e.g., synchronized audio and video streams) is investigated. H.264/AVC and AAC are used to evaluate our proposed approach, which relies on a format-independent adaptation framework based on MPEG-B BSDL and STX. Davy Van Deursen, Wesley De Neve, Wim Van Lancker, Rik Van de Walle |
PCS | 2 |
| 2009 | Improved BSDL-based content adaptation for JPEG 2000 and HD Photo (JPEG XR)
Wesley De Neve, Davy Van Deursen, Wim Van Lancker, Yong Man Ro, Rik Van de Walle |
Signal Process. Image Commun. | 1 |
| 2008 | NinSuna: A Format-Independent Multimedia Content Adaptation Platform Based on Semantic Web TechnologiesabstractMultimedia content adaptation is gaining importance because of the growing amount of multimedia content on the one hand and the growing diversity in usage environments on the other hand. Furthermore, to deal with the growing amount of coding formats for multimedia content, format-independent adaptation systems are desired. These systems support the exploitation of scalability to meet the usage environment, as well as semantic adaptations to meet the user preferences. In this demonstration, we present NinSuna, a fully integrated multimedia content adaptation platform based on semantic web technologies. It aims at being deployable in streaming environments and relies on format-independent semantic-aware adaptation engines. Davy Van Deursen, Wim Van Lancker, Tom Paridaens, Wesley De Neve, Erik Mannens, Rik Van de Walle |
ISM | 4 |
| 2008 | gBFlavor: a new tool for fast and automatic generation of generic bitstream syntax descriptions
Davy Van Deursen, Wesley De Neve, Davy De Schrijver, Rik Van de Walle |
Multim. Tools Appl. | 2 |
| 2008 | A compressed-domain approach for shot boundary detection on H.264/AVC bit streams
Sarah De Bruyne, Davy Van Deursen, Jan De Cock, Wesley De Neve, Peter Lambert, Rik Van de Walle |
Signal Process. Image Commun. | 4 |
| 2007 | Exploitation of Combined Scalability in Scalable H.264/AVC Bitstreams by Using an MPEG-21 XML-Driven Framework
Davy De Schrijver, Wesley De Neve, Koen De Wolf, Davy Van Deursen, Rik Van de Walle |
ACIVS | 2 |
| 2007 | Analysis of Prediction Mode Decision in Spatial Enhancement Layers in H.264/AVC SVC
Koen De Wolf, Davy De Schrijver, Wesley De Neve, Saar De Zutter, Peter Lambert, Rik Van de Walle |
CAIP | 3 |
| 2007 | XML-driven Exploitation of Combined Scalability in Scalable H.264/AVC BitstreamsabstractThe heterogeneity in the contemporary multimedia environments requires a format-agnostic adaptation framework for the consumption of digital video content. Scalable bitstreams can be used in order to satisfy as many circumstances as possible. In this paper, the scalable extension on the H.264/AVC specification is used to obtain the parent bitstreams. The adaptation along the combined scalability axis of the bitstreams is done in a format-independent manner. Therefore, an abstraction layer of the bitstream is needed. In this paper, XML descriptions are used representing the high-level structure of the bitstreams by relying on the MPEG-21 bitstream syntax description language standard. The exploitation of the combined scalability is executed in the XML domain by implementing the adaptation process in a streaming transformation for XML (STX) stylesheet. The algorithm used in the transformation of the XML description is discussed in detail in this paper. From the performance measurements, one can conclude that the STX transformation in the XML domain and the generation of the corresponding adapted bitstream can be realized in real time. Davy De Schrijver, Wesley De Neve, Koen De Wolf, Peter Lambert, Davy Van Deursen, Rik Van de Walle |
ISCAS | 2 |
| 2007 | MuMiVA: A Multimedia Delivery Platform Using Format-Agnostic, XML-Driven Content AdaptationabstractDue to the increasing heterogeneity in the current multimedia landscape, the delivery of multimedia content has become an important issue today. This heterogeneity is not only reflected by a plethora of different usage environments, but also by the presence of multiple (scalable) coding formats. Therefore, format-independent adaptation engines have to be used within a multimedia delivery platform, which are able to adapt the multimedia content according to a certain usage environment, independent of the underlying coding format of the content. By relying on automatically created textual descriptions of the high-level syntax of binary media resources, a format-independent adaptation engine can be built. MPEG-21 generic bitstream syntax schema (gBS schema) is a tool that is part of the MPEG-21 multimedia framework. It enables the use of generic bitstream syntax descriptions (gBSDs), i.e., textual descriptions in XML, to steer the adaptation of a binary media resource, using format-independent adaptation logic. In this paper, we address the design and performance evaluation of a multimedia delivery platform that relies on gBS schema-driven adaptation engines. Our platform is called MuMiVA; it is a fully integrated, extensible platform for multimedia delivery in heterogeneous usage environments, using streaming technologies. To demonstrate the flexibility of our multimedia delivery platform, we discuss the functioning of two different applications (i.e., exploitation of temporal scalability and shot selection) applied to two different coding formats (i.e., MPEG-4 visual and H.264/AVC). Davy Van Deursen, Sarah De Bruyne, Wim Van Lancker, Wesley De Neve, Davy De Schrijver, Hermann Hellwagner, Rik Van de Walle |
ISM | 4 |
| 2007 | Temporal Video Segmentation on H.264/AVC Compressed Bitstreams
Sarah De Bruyne, Wesley De Neve, Koen De Wolf, Davy De Schrijver, Piet Verhoeve, Rik Van de Walle |
MMM (1) | 2 |
| 2007 | An optimized MPEG-21 BSDL framework for the adaptation of scalable bitstreams
Davy De Schrijver, Wesley De Neve, Koen De Wolf, Robbie De Sutter, Rik Van de Walle |
J. Vis. Commun. Image Represent. | 2 |
| 2006 | A Real-Time Content Adaptation Framework for Exploiting ROI Scalability in H.264/AVC
Peter Lambert, Davy De Schrijver, Davy Van Deursen, Wesley De Neve, Yves Dhondt, Rik Van de Walle |
ACIVS | 4 |
| 2006 | XML-based customization along the scalability axes of H.264/AVC scalable video codingabstractThe heterogeneity in the current and future multimedia environment requires an elegant adaptation framework for the production and consumption of different kinds of multimedia content. Such an architecture is preferably based on the usage of scalable bitstreams and a format-agnostic content adaptation engine. To obtain fully embedded scalable bitstreams, the joint scalable video model (JSVM) has been used in this paper. Hereby, JSVM defines a scalable extension on top of the H.264/AVC specification. This extension will make it possible to create bitstreams that are scalable along the temporal, spatial, and SNR axis. On the other hand, bitstream structure descriptions can be used to realize an elegant and format-agnostic adaptation engine. Such descriptions can be created by making use of the MPEG-21 bitstream syntax description language (BSDL) standard. The latter allows to describe the high-level structure of scalable bitstreams in XML. This paper explains how fully scalable bitstreams can be customized by transforming BSDL-based bitstream structure descriptions. From our performance analysis, one can conclude that the transformation of the XML description, as well as the generation of the adapted bitstream, can be done several times faster than real time Davy De Schrijver, Wesley De Neve, Koen De Wolf, Stijn Notebaert, Rik Van de Walle |
ISCAS | 2 |
| 2006 | Flexible macroblock ordering in H.264/AVC
Peter Lambert, Wesley De Neve, Yves Dhondt, Rik Van de Walle |
J. Vis. Commun. Image Represent. | 2 |
| 2006 | MPEG-21 bitstream syntax descriptions for scalable video codecs
Davy De Schrijver, Chris Poppe, Sam Lerouge, Wesley De Neve, Rik Van de Walle |
Multim. Syst. | 4 |
| 2006 | BFlavor: A harmonized approach to media resource adaptation, inspired by MPEG-21 BSDL and XFlavor
Wesley De Neve, Davy Van Deursen, Davy De Schrijver, Sam Lerouge, Koen De Wolf, Rik Van de Walle |
Signal Process. Image Commun. | 1 |
| 2006 | Rate-distortion performance of H.264/AVC compared to state-of-the-art video codecsabstractIn the domain of digital video coding, new technologies and solutions are emerging in a fast pace, targeting the needs of the evolving multimedia landscape. One of the questions that arises is how to assess these different video coding technologies in terms of compression efficiency. In this paper, several compression schemes are compared by means of peak signal-to-noise ratio (PSNR) and just noticeable difference (JND). The codecs examined are XviD 0.9.1 (conform to the MPEG-4 Visual Simple Profile), DivX 5.1 (implementing the MPEG-4 Visual Advanced Simple Profile), Windows Media Video 9, MC-EZBC and H.264/AVC AHM 2.0 (version JM 6.1 of the reference software, extended with rate control). The latter plays a key role in this comparison because the H.264/AVC standard can be considered as the de facto benchmark in the field of digital video coding. The obtained results show that H.264/AVC AHM 2.0 outperforms current proprietary and standards-based implementations in almost all cases. Another observation is that the choice of a particular quality metric can influence general statements about the relation between the different codecs. Peter Lambert, Wesley De Neve, Peter De Neve, Ingrid Moerman, Piet Demeester, Rik Van de Walle |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | Generating MPEG-21 BSDL Descriptions Using Context-Related AttributesabstractIn order to efficiently deal with the heterogeneity in the current and future multimedia ecosystem, it is necessary that content can be adapted in a format-agnostic manner. A first step toward a solution, able to fulfill the just mentioned requirement, is to rely on a scalable video codec and to describe the high-level structure of the resulting bitstreams in such a way that every terminal can understand it, in particular by using XML. This paper describes how such descriptions can be generated by making use of the media format independent BintoBSD tool of the MPEG-21 BSDL standard. However, regarding the current status of BSDL, it is impossible to create a description in real time and to keep the generation speed constant over the complete sequence. In this paper, we describe a number of extensions and algorithmic modifications that make it possible to generate a description of a bitstream in real time and at a constant speed. Our approach results in a significant reduction of the original execution times (up to 99% for the H.264/AVC coding format) and in a constant memory usage. Davy De Schrijver, Wesley De Neve, Koen De Wolf, Rik Van de Walle |
ISM | 2 |
| 2005 | GPU-assisted decoding of video samples represented in the YCoCg-R color spaceabstractAlthough pixel shaders were designed for the creation of programmable rendering effects, they can also be used as generic processing units for vector data. In this paper, attention is paid to an implementation of the YCoCg-R to RGB color space transform, as defined in the H.264/AVC Fidelity Range Extensions, by making use of pixel shaders. Our results show that a significant speedup can be achieved by relying on the processing power of the GPU, relative to the CPU. To be more specific, high definition video (1080p), represented in the YCoCg-R color space, could be decoded to RGB at 30 Hz on a PC with an AMD Athlon XP 2800+ CPU, an AGP bus and an NVIDIA GeForce 6800 graphics card, an effort that could not be realized in real-time by the CPU. Wesley De Neve, Dieter Van Rijsselbergen, Charles-Frederik Hollemeersch |
ACM Multimedia | 1 |
| 2004 | Assessment of the compression efficiency of the MPEG-4 AVC specificationabstractVideo coding is used under the hood of a lot of multimedia applications, such as video conferencing, digital storage media, television broadcasting, and internet streaming. Recently, new standards-based and proprietary technologies have emerged. An interesting problem is how to evaluate these different video coding solutions in terms of delivered quality. In this paper, a PSNR-based approach is applied in order to compare the coding potential of H.264/AVC AHM 2.0 with the compression efficiency of XviD 0.9.1, DivX 5.05, Windows Media Video 9, and MC-EZBC. Our results show that MPEG-4-based tools, and in particular H.264/AVC, can keep step with proprietary solutions. The rate-distortion performance of MC-EZBC, a wavelet-based video codec, looks very promising too. Wesley De Neve, Peter Lambert, Sam Lerouge, Rik Van de Walle |
VCIP | 1 |