EDBT 2026 Demo / reviewers in the wild / expert
Hanhe Lin
dblp:124/2752
· DBLP profile ↗
40ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0002-0297-3549ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Applications of Hypergraph Learning for Brain Disorder Diagnosis with Neuroimaging: A Survey
Meng-Shen He, Jun-Jian Li, Hai-Lin Yue, Hulin Kuang, Hanhe Lin, Zhen Qiu 0001, Jin Liu 0012 |
J. Comput. Sci. Technol. | 7 |
| 2026 | Trans-SURNet: A linear transformer approach to model picture-wise JND distribution for SUR curve prediction
Anni Zhang, Laifan Pei, Chunling Fan, Hanhe Lin |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | SGNet: Style-Guided Network With Temporal Compensation for Unpaired Low-Light Colonoscopy Video EnhancementabstractA low-light colonoscopy video enhancement method is needed as poor illumination in colonoscopy can hinder accurate disease diagnosis and adversely affect surgical procedures. Existing low-light video enhancement methods usually apply a frame-by-frame enhancement strategy without considering the temporal correlation between them, which often causes a flickering problem. In addition, most methods are designed for endoscopic devices with fixed imaging styles and cannot be easily adapted to different devices. In this paper, we propose a Style-Guided Network (SGNet) for unpaired Low-Light Colonoscopy Video Enhancement (LLCVE). Given that collecting content-consistent paired videos is difficult, SGNet adopts a CycleGAN-based framework to convert low-light videos to normal-light videos, in which a Temporal Compensation (TC) module and a Style Guidance (SG) module are proposed to alleviate the flickering problem and achieve flexible style transfer, respectively. The TC module compensates for a low-light frame by learning the correlated feature of its adjacent frames, thereby improving the temporal smoothness of the enhanced video. The SG module encodes the text of the imaging style and adaptively explores its intrinsic relationships with video features to obtain style representations, which are then used to guide the subsequent enhancement process. Extensive experiments on a curated database show that SGNet achieves promising performance on the LLCVE task, outperforming state-of-the-art methods in both quantitative metrics and visual quality. Guanghui Yue 0001, Wanqing Liu, Jingfeng Du, Tianwei Zhou, Hanhe Lin, Qiuping Jiang, Wenqi Ren |
IEEE Trans. Image Process. | 6 |
| 2025 | Radar-Based Cross-Environment Unsupervised Domain Adaptation for Human Activity RecognitionabstractMillimeter-wave radar, with its ease of deployment and around-the-clock operation, has become widely adopted for monitoring daily human activities in mobile health. Human activity recognition (HAR) models trained on collected radar data enable persistent detection of critical incidents such as falls, which is critical for reducing in-home care costs and improving healthcare resource utilization efficiency. However, radar signals are easily disrupted by environmental variations, causing models trained in one setting to generalize poorly when applied to new environment. To address this challenge, we propose a radar-based cross-environment unsupervised domain adaptation (UDA) method, which significantly enhances generalization performance in unlabeled target domains (new environment) by leveraging fully labeled source domain data (historical environment). Our approach comprises two key components: a feature fusion module and a knowledge distillation module. First, the feature fusion module employs an attention mechanism to perform weighted linear interpolation between source domain and target domain features and uses an alignment loss to learn a shared feature rep-resentation robust to environmental changes. Next, temperature-scaled knowledge distillation generates soft pseudo labels for unlabeled target domain samples, improving the model's dis-criminatory power without ground-truth annotations. Evaluated on twelve cross-environment HAR tasks, the proposed method achieves an average recognition accuracy of 94.19%, offering an efficient and reliable solution for radar-based mobile health monitoring. Ludi Li, Jin Liu 0012, Hanhe Lin, Mingzi Yuan, Jianchun Zhu |
BIBM | 3 |
| 2025 | MMP-2k: A Benchmark Multi-Labeled Macro Photography Image Quality Assessment DatabaseabstractMacro photography (MP) is a specialized field of photography that captures objects at an extremely close range, revealing tiny details. Although an accurate macro photography image quality assessment (MPIQA) metric can benefit macro photograph capturing, which is vital in some domains such as scientific research and medical applications, the lack of MPIQA data limits the development of MPIQA metrics. To address this limitation, we conducted a large-scale MPIQA study. Specifically, to ensure diversity both in content and quality, we sampled 2,000 MP images from 15,700 MP images, collected from three public image websites. For each MP image, 17 (out of 21 after outlier removal) quality ratings and a detailed quality report of distortion magnitudes, types, and positions are gathered by a lab study. The images, quality ratings, and quality reports form our novel multi-labeled MPIQA database, MMP-2k. Experimental results showed that the state-of-the-art generic IQA metrics underperform on MP images. The database and supplementary materials are available at https://github.com/Future-IQA/MMP-2k. Jiashuo Chang, Jianxun Lou, Zhen Qiu 0001, Hanhe Lin |
ICIP | 5 |
| 2025 | Improving Novel View Synthesis of 360° Scenes in Extremely Sparse Views by Jointly Training Hemisphere Sampled Synthetic ImagesabstractNovel view synthesis in 360° scenes from extremely sparse input views is essential for applications like virtual reality and augmented reality. This paper presents a novel framework for novel view synthesis in extremely sparse-view cases. As typical structure-from-motion methods are unable to estimate camera poses in extremely sparse-view cases, we apply DUSt3R to estimate camera poses and generate a dense point cloud. Using the poses of estimated cameras, we densely sample additional views from the upper hemisphere space of the scenes, from which we render synthetic images together with the point cloud. Training 3D Gaussian Splatting model on a combination of reference images from sparse views and densely sampled synthetic images allows a larger scene coverage in 3D space, addressing the overfitting challenge due to the limited input in sparse-view cases. Retraining a diffusion-based image enhancement model on our created dataset, we further improve the quality of the point-cloud-rendered images by removing artifacts. We compare our framework with benchmark methods in cases of only four input views, demonstrating significant improvement in novel view synthesis under extremely sparse-view conditions for 360° scenes. The source code is available at https: //github.com/angchen-dev/hemiSparseGS. Guangan Chen, Anh Minh Truong, Hanhe Lin, Michiel Vlaminck, Wilfried Philips, Hiêp Quang Luong |
ICIP | 3 |
| 2025 | Inspired by pathogenic mechanisms: A novel gradual multi-modal fusion framework for mild cognitive impairment diagnosis
Hong-Dong Li, Hanhe Lin, Chao Li 0031, Harrison X. Bai, Wei Lan 0001, Jin Liu 0012 |
Neural Networks | 3 |
| 2025 | Multi-Modal Multi-Kernel Graph Learning for Autism Prediction and Biomarker DiscoveryabstractGraph learning-based multi-modal integration and classification is one of the most challenging tasks for disease prediction. To effectively offset the negative impact among modalities in the process of multi-modal integration and heterogeneous information extractions from graphs, we propose a novel method called Multi-modal Multi-Kernel Graph Learning (MMKGL). To solve the problem of negative impact among modalities, we propose a multi-modal graph embedding module to construct a multi-modal graph. Different from conventional methods that manually construct static graphs for all modalities, each modality generates a separate graph by adaptive learning, where a function graph and a supervision graph are introduced for optimization during the multi-graph fusion embedding process. We then propose a multi-kernel graph learning module to extract heterogeneous information from the multi-modal graph. The information in the multi-modal graph at different levels is aggregated by convolutional kernels with different receptive field sizes, followed by generating a cross-kernel discovery tensor for disease prediction. Our method is evaluated on the benchmark Autism Brain Imaging Data Exchange (ABIDE) dataset and outperforms the state-of-the-art methods. In addition, discriminative brain regions associated with autism are identified by our model, providing guidance for the study of autism pathology. Jin Liu 0012, Junbin Mao, Hanhe Lin, Hulin Kuang, Shirui Pan, Xusheng Wu, Shan Xie, Fei Liu 0058, Yi Pan 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Benchmarking Laryngeal Neoplasm Segmentation: A Multicenter Dataset and an Effective MethodabstractWhile accurate and automatic Laryngeal Neoplasm Segmentation (LNS) can benefit the diagnosis and prevention of laryngeal cancers, existing LNS-related works are very limited due to the lack of public datasets. This paper conducts systematic research to take the research field a step further. Firstly, we create a multicenter LNS dataset, named as MLN-Seg. Collecting from four hospitals, it has 2,273 laryngeal images with a diversity in resolutions and modalities, where each image is pixel-wise annotated by experienced physicians. Secondly, considering the scarcity of LNS methods and similarity between LNS and Colorectal Polyp Segmentation (CPS) tasks, we collect 15 CPS methods and validate their performance on MLN-Seg. It shows that despite the similarity between the two tasks, existing CPS methods underperform on LNS, especially those with blurry boundaries and camouflaged characteristics. Lastly, considering the LNS challenges, we propose an effective segmentation method, termed Scale-Sensitive Network (S2Net). S2Net scales the feature at each layer of the network up and down and integrates all the scaled features to coarsely localize neoplasm regions. In addition, a Localization Calibration (LC) module is used to refine uncertain areas. By connecting the LC modules from top to down, S2Net can finally accurately segment the laryngeal neoplasms. Extensive tests on MLN-Seg shows that S2Net has better learning ability and generalizability than competing methods. In addition, evaluation on five public datasets shows that S2Net achieves comparable performance in the CPS task. Guanghui Yue 0001, Shangjie Wu, Ruxian Tian, Hanhe Lin, Huaiqing Lv, Zhenkun Yu, Xicheng Song |
IEEE Trans. Image Process. | 4 |
| 2025 | Toward Integrating Federated Learning With Split Learning via Spatio-Temporal Graph Framework for Brain Disease PredictionabstractFunctional Magnetic Resonance Imaging (fMRI) is used for extracting blood oxygen signals from brain regions to map brain functional connectivity for brain disease prediction. Despite its effectiveness, fMRI has not been widely used: on the one hand, collecting and labeling the data is time-consuming and costly, which limits the amount of valid data collected at a single healthcare site; on the other hand, integrating data from multiple sites is challenging due to data privacy restrictions. To address these issues, we propose a novel, integrated Federated learning and Split learning Spatio-temporal Graph framework (F G). Specifically, we introduce federated learning and split learning techniques to split a spatio-temporal model into a client temporal model and a server spatial model. In the client temporal model, we propose a time-aware mechanism to focus on changes in brain functional states and use an InceptionTime model to extract information about changes in the brain states of each subject. In the server spatial model, we propose a united graph convolutional network to integrate multiple graph convolutional networks. Integrating federated learning and split learning, F G can utilize multi-site fMRI data without violating data privacy protection and reduce the risk of overfitting as it is capable of learning from limited training data sets. Moreover, it boosts the extraction of spatio-temporal features of fMRI using spatio-temporal graph networks. Experiments on ABIDE and ADHD200 datasets demonstrate that our proposed method outperforms state-of-the-art methods. In addition, we explore biomarkers associated with brain disease prediction using community discovery algorithms using intermediate results of F G. The source code is available at https://github.com/yutian0315/FS2G. Junbin Mao, Jin Liu 0012, Yi Pan 0001, Emanuele Trucco, Hanhe Lin |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Source-Free Domain Adaptation for Millimeter Wave Radar Based Human Activity RecognitionabstractHuman activity recognition based on millimeter-wave radar is dedicated to monitor people’s daily activities and detect specific dangerous actions. Although existing methods achieve some improvement, they rarely consider the challenges of domain difference, such as ages and environments. To address this challenge, we propose a source-free domain adaptation method for millimeter wave radar based human activity recognition, which achieves knowledge transfer from the source domain to the target domain. Firstly, we propose balanced clustering to obtain cluster centers of source domain as the prior-knowledge through the pre-trained model. Then, in order to perform domain adaptation, the model is fine-tuned by the integration of domain adaptation and self-supervision of the target domain. Experiment results on several transfer tasks show that our proposed method is effective in human activity recognition and outperforms some other advanced transfer learning methods. Jin Liu 0012, Dejiao Zeng, Ludi Li, Hanhe Lin |
ICASSP | 4 |
| 2024 | Graph-Based Fusion of Imaging, Genetic and Clinical Data for Degenerative Disease DiagnosisabstractGraph learning methods have achieved noteworthy performance in disease diagnosis due to their ability to represent unstructured information such as inter-subject relationships. While it has been shown that imaging, genetic and clinical data are crucial for degenerative disease diagnosis, existing methods rarely consider how best to use their relationships. How best to utilize information from imaging, genetic and clinical data remains a challenging problem. This study proposes a novel graph-based fusion (GBF) approach to meet this challenge. To extract effective imaging-genetic features, we propose an imaging-genetic fusion module which uses an attention mechanism to obtain modality-specific and joint representations within and between imaging and genetic data. Then, considering the effectiveness of clinical information for diagnosing degenerative diseases, we propose a multi-graph fusion module to further fuse imaging-genetic and clinical features, which adopts a learnable graph construction strategy and a graph ensemble method. Experimental results on two benchmarks for degenerative disease diagnosis (Alzheimers Disease Neuroimaging Initiative and Parkinson's Progression Markers Initiative) demonstrate its effectiveness compared to state-of-the-art graph-based methods. Our findings should help guide further development of graph-based models for dealing with imaging, genetic and clinical data. Rui Guo 0009, Hanhe Lin, Stephen J. McKenna, Hong-Dong Li, Fei Guo 0001, Jin Liu 0012 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Predicting Radiologists' Gaze With Computational Saliency Models in Mammogram ReadingabstractPrevious studies have shown that there is a strong correlation between radiologists' diagnoses and their gaze when reading medical images. The extent to which gaze is attracted by content in a visual scene can be characterised as visual saliency. There is a potential for the use of visual saliency in computer-aided diagnosis in radiology. However, little is known about what methods are effective for diagnostic images, and how these methods could be adapted to address specific applications in diagnostic imaging. In this study, we investigate 20 state-of-the-art saliency models including 10 traditional models and 10 deep learning-based models in predicting radiologists' visual attention while reading 196 mammograms. We found that deep learning-based models represent the most effective type of methods for predicting radiologists' gaze in mammogram reading; and that the performance of these saliency models can be significantly improved by transfer learning. In particular, an enhanced model can be achieved by pre-training the model on a large-scale natural image saliency dataset and then fine-tuning it on the target medical image dataset. In addition, based on a systematic selection of backbone networks and network architectures, we proposed a parallel multi-stream encoded model which outperforms the state-of-the-art approaches for predicting saliency of mammograms. Jianxun Lou, Hanhe Lin, Philippa Young, Zelei Yang, Susan Cheng Shelmerdine, David Marshall 0001, Emiliano Spezi, Marco Palombo, Hantao Liu |
IEEE Trans. Multim. | 2 |
| 2024 | Going the Extra Mile in Face Image Quality Assessment: A Novel Database and ModelabstractAn accurate computational model for image quality assessment (IQA) benefits many vision applications, such as image filtering, image processing, and image generation. Although the study of face images is an important subfield in computer vision research, the lack of face IQA data and models limits the precision of current IQA metrics on face image processing tasks such as face superresolution, face enhancement, and face editing. To narrow this gap, in this article, we first introduce the largest annotated IQA database developed to date, which contains 20,000 human faces – an order of magnitude larger than all existing rated datasets of faces – of diverse individuals in highly varied circumstances. Based on the database, we further propose a novel deep learning model to accurately predict face image quality, which, for the first time, explores the use of generative priors for IQA. By taking advantage of rich statistics encoded in well pretrained off-the-shelf generative models, we obtain generative prior information and use it as latent references to facilitate blind IQA. The experimental results demonstrate both the value of the proposed dataset for face IQA and the superior performance of the proposed model. Shaolin Su, Hanhe Lin, Vlad Hosu, Oliver Wiedemann, Jinqiu Sun, Yu Zhu 0004, Hantao Liu, Yanning Zhang 0001, Dietmar Saupe |
IEEE Trans. Multim. | 2 |
| 2024 | Neural Inference Search for Multiloss Segmentation ModelsabstractSemantic segmentation is vital for many emerging surveillance applications, but current models cannot be relied upon to meet the required tolerance, particularly in complex tasks that involve multiple classes and varied environments. To improve performance, we propose a novel algorithm, neural inference search (NIS), for hyperparameter optimization pertaining to established deep learning segmentation models in conjunction with a new multiloss function. It incorporates three novel search behaviors, i.e., Maximized Standard Deviation Velocity Prediction, Local Best Velocity Prediction, and n -dimensional Whirlpool Search. The first two behaviors are exploratory, leveraging long short-term memory (LSTM)-convolutional neural network (CNN)-based velocity predictions, while the third employs n -dimensional matrix rotation for local exploitation. A scheduling mechanism is also introduced in NIS to manage the contributions of these three novel search behaviors in stages. NIS optimizes learning and multiloss parameters simultaneously. Compared with state-of-the-art segmentation methods and those optimized with other well-known search algorithms, NIS-optimized models show significant improvements across multiple performance metrics on five segmentation datasets. NIS also reliably yields better solutions as compared with a variety of search methods for solving numerical benchmark functions. Sam Slade, Li Zhang 0013, Haoqian Huang, Houshyar Asadi, Chee Peng Lim, Yonghong Yu, Dezong Zhao, Hanhe Lin, Rong Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | FedGST: Federated Graph Spatio-Temporal Framework for Brain Functional Disease PredictionabstractCurrently, most medical institutions face the challenge of training a unified model using fragmented and isolated data to address disease prediction problems. Although federated learning has become the recognized paradigm for privacy-preserving model training, how to integrate federated learning with fMRI temporal characteristics to enhance predictive performance remains an open question for functional disease prediction. To address this challenging task, we propose a novel Federated Graph Spatio-Temporal (FedGST) framework for brain functional disease prediction. Specifically, anchor sampling is used to process variable-length time series data on local clients. Then dynamic functional connectivity graphs are generated via sliding windows and Pearson correlation coefficients. Next, we propose an InceptionTime model to extract temporal information from the dynamic functional connectivity graphs on the local clients. Finally, the hidden activation variables are sent to a global server. We propose a UniteGCN model on the global server to receive and process the hidden activation variables from clients. Then, the global server returns gradient information to clients for backpropagation and model parameter updating. Client models aggregate model parameters on the local server and distribute them to clients for the next round of training. We demonstrate that FedGST outperforms other federated learning methods and baselines on ABIDE-1 and ADHD200 datasets. Junbin Mao, Hanhe Lin, Yi Pan 0001, Jin Liu 0012 |
BIBM | 2 |
| 2023 | Localization of Just Noticeable Difference for Image CompressionabstractThe just noticeable difference (JND) is the minimal difference between stimuli that can be detected by a person. The picture-wise just noticeable difference (PJND) for a given reference image and a compression algorithm represents the minimal level of compression that causes noticeable differences in the reconstruction. These differences can only be observed in some specific regions within the image, dubbed as JND-critical regions. Identifying these regions can improve the development of image compression algorithms. Due to the fact that visual perception varies among individuals, determining the PJND values and JND-critical regions for a target population of consumers requires subjective assessment experiments involving a sufficiently large number of observers. In this paper, we propose a novel framework for conducting such experiments using crowdsourcing. By applying this framework, we created a novel PJND dataset, KonJND++, consisting of 300 source images, compressed versions thereof under JPEG or BPG compression, and an average of 43 ratings of PJND and 129 self-reported locations of JND-critical regions for each source image. Our experiments demonstrate the effectiveness and reliability of our proposed framework, which is easy to be adapted for collecting a large-scale dataset. The source code and dataset are available at https://github.com/angchen-dev/LocJND. Guangan Chen, Hanhe Lin, Oliver Wiedemann, Dietmar Saupe |
QoMEX | 2 |
| 2022 | Predicting Radiologist Attention During Mammogram Reading with Deep and Shallow High-Resolution EncodingabstractRadiologists’ eye-movement during diagnostic image reading reflects their personal training and experience, which means that their diagnostic decisions are related to their perceptual processes. For training, monitoring, and performance evaluation of radiologists, it would be beneficial to be able to automatically predict the spatial distribution of the radiologist’s visual attention on the diagnostic images. The measurement of visual saliency is a well-studied area that allows for prediction of a person’s gaze attention. However, compared with the extensively studied natural image visual saliency (in free viewing tasks), the saliency for diagnostic images is less studied; there could be fundamental differences in eye-movement behaviours between these two domains. Most current saliency prediction models have been optimally developed for natural images, which could lead them to be less adept at predicting the visual attention of radiologists during the diagnosis. In this paper, we propose a method specifically for automatically capturing the visual attention of radiologists during mammogram reading. By adopting high-resolution image representations from both deep and shallow encoders, the proposed method avoids potential detail losses and achieves superior results on multiple evaluation metrics in a large mammogram eye-movement dataset. Jianxun Lou, Hanhe Lin, David Marshall 0001, Young Yang, Susan Cheng Shelmerdine, Hantao Liu |
ICIP | 2 |
| 2022 | Crowdsourced Quality Assessment of Enhanced Underwater Images - a Pilot StudyabstractUnderwater image enhancement (UIE) is essential for a high-quality underwater optical imaging system. While a number of UIE algorithms have been proposed in recent years, there is little study on image quality assessment (IQA) of enhanced underwater images. In this paper, we conduct the first crowdsourced subjective IQA study on enhanced underwater images. We chose ten state-of-the-art UIE algorithms and applied them to yield enhanced images from an underwater image benchmark. Their latent quality scales were reconstructed from pair comparison. We demonstrate that the existing IQA metrics are not suitable for assessing the perceived quality of enhanced underwater images. In addition, the overall performance of 10 UIE algorithms on the benchmark is ranked by the newly proposed simulated pair comparison of the methods. Hanhe Lin, Hui Men, Yijun Yan, Jinchang Ren, Dietmar Saupe |
QoMEX | 1 |
| 2022 | TranSalNet: Towards perceptually relevant visual saliency predictionabstractConvolutional neural networks (CNNs) have significantly advanced computational modelling for saliency prediction. However, accurately simulating the mechanisms of visual attention in the human cortex remains an academic challenge. It is critical to integrate properties of human vision into the design of CNN architectures, leading to perceptually more relevant saliency prediction. Due to the inherent inductive biases of CNN architectures, there is a lack of sufficient long-range contextual encoding capacity. This hinders CNN-based saliency models from capturing properties that emulate viewing behaviour of humans. Transformers have shown great potential in encoding long-range information by leveraging the self-attention mechanism. In this paper, we propose a novel saliency model that integrates transformer components to CNNs to capture the long-range contextual visual information. Experimental results show that the transformers provide added value to saliency prediction, enhancing its perceptual relevance in the performance. Our proposed saliency model using transformers has achieved superior results on public benchmarks and competitions for saliency prediction models. The source code of our proposed saliency model TranSalNet is available at: https://github.com/LJOVO/TranSalNet. Jianxun Lou, Hanhe Lin, David Marshall 0001, Dietmar Saupe, Hantao Liu |
Neurocomputing | 2 |
| 2022 | Large-Scale Crowdsourced Subjective Assessment of Picturewise Just Noticeable DifferenceabstractThe picturewise just noticeable difference (PJND) for a given image, compression scheme, and subject is the smallest distortion level that the subject can perceive when the image is compressed with this compression scheme. The PJND can be used to determine the compression level at which a given proportion of the population does not notice any distortion in the compressed image. To obtain accurate and diverse results, the PJND must be determined for a large number of subjects and images. This is particularly important when experimental PJND data are used to train deep learning models that can predict a probability distribution model of the PJND for a new image. To date, such subjective studies have been carried out in laboratory environments. However, the number of participants and images in all existing PJND studies is very small because of the challenges involved in setting up laboratory experiments. To address this limitation, we develop a framework to conduct PJND assessments via crowdsourcing. We use a new technique based on slider adjustment and a flicker test to determine the PJND. A pilot study demonstrated that our technique could decrease the study duration by 50% and double the perceptual sensitivity compared to the standard binary search approach that successively compares a test image side by side with its reference image. Our framework includes a robust and systematic scheme to ensure the reliability of the crowdsourced results. Using 1,008 source images and distorted versions obtained with JPEG and BPG compression, we apply our crowdsourcing framework to build the largest PJND dataset, KonJND-1k (Konstanz just noticeable difference 1k dataset). A total of 503 workers participated in the study, yielding 61,030 PJND samples that resulted in an average of 42 samples per source image. The KonJND-1k dataset is available athttp://database.mmsp-kn.de/konjnd-1k-database.html Hanhe Lin, Guangan Chen, Mohsen Jenadeleh, Vlad Hosu, Ulf-Dietrich Reips, Raouf Hamzaoui, Dietmar Saupe |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | KonIQ++: Boosting No-Reference Image Quality Assessment in the Wild by Jointly Predicting Image Quality and Defects
Shaolin Su, Vlad Hosu, Hanhe Lin, Yanning Zhang 0001, Dietmar Saupe |
BMVC | 3 |
| 2021 | Positional Encoding: Improving Class-Imbalanced Motorcycle Helmet use ClassificationabstractRecent advances in the automated detection of motorcycle riders’ helmet use have enabled road safety actors to process large scale video data efficiently and with high accuracy. To distinguish drivers from passengers in helmet use, the most straightforward way is to train a multi-class classifier, where each class corresponds to a specific combination of rider position and individual riders’ helmet use. However, such strategy results in long-tailed data distribution, with critically low class samples for a number of uncommon classes. In this paper, we propose a novel approach to address this limitation. Let n be the maximum number of riders a motorcycle can hold, we encode the helmet use on a motorcycle as a vector with 2n bits, where the first n bits denote if the encoded positions have riders, and the latter n bits denote if the rider in the corresponding position wears a helmet. With the novel helmet use positional encoding, we propose a deep learning model that stands on existing image classification architecture. The model simultaneously trains 2n binary classifiers, which allows more balanced samples for training. This method is simple to implement and requires no hyperparameter tuning. Experimental results demonstrate our approach outperforms the state-of-the-art approaches by 1.9% accuracy. Hanhe Lin, Guangan Chen, Felix W. Siebert |
ICIP | 1 |
| 2020 | EvolGAN: Evolutionary Generative Adversarial Networks
Baptiste Rozière, Fabien Teytaud, Vlad Hosu, Hanhe Lin, Jérémy Rapin, Mariia Zameshina, Olivier Teytaud |
ACCV (4) | 4 |
| 2020 | Deep Learning VS. Traditional Algorithms for Saliency Prediction of Distorted ImagesabstractSaliency has been widely studied in relation to image quality assessment (IQA). The optimal use of saliency in IQA metrics, however, is nontrivial and largely depends on whether saliency can be accurately predicted for images containing various distortions. Although tremendous progress has been made in saliency modelling, very little is known about whether and to what extent state-of-the-art methods are beneficial for saliency prediction of distorted images. In this paper, we analyse the ability of deep learning versus traditional algorithms in predicting saliency, based on an IQA-aware saliency benchmark, the SIQ288 database. Building off the variations in model performance, we make recommendations for model selections for IQA applications. Hanhe Lin, Dietmar Saupe, Hantao Liu |
ICIP | 2 |
| 2020 | Tarsier: Evolving Noise Injection in Super-Resolution GANsabstractSuper-resolution aims at increasing the resolution and level of detail within an image. The current state of the art in general single-image super-resolution is held by NESRGAN+, which injects a Gaussian noise after each residual layer at training time. In this paper, we harness evolutionary methods to improve NESRGAN+ by optimizing the noise injection at inference time. More precisely, we use Diagonal CMA to optimize the injected noise according to a novel criterion combining quality assessment and realism. Our results are validated by the PIRM perceptual score and a human study. Our method outperforms NESRGAN+ on several standard super-resolution datasets. More generally, our approach can be used to optimize any method based on noise injection. Baptiste Rozière, Nathanaël Carraz Rakotonirina, Vlad Hosu, Andry Rasoanaivo, Hanhe Lin, Camille Couprie, Olivier Teytaud |
ICPR | 5 |
| 2020 | Visual Quality Assessment for Interpolated Slow-Motion Videos Based on a Novel DatabaseabstractProfessional video editing tools can generate slow-motion video by interpolating frames from video recorded at a standard frame rate. Thereby the perceptual quality of such interpolated slow-motion videos strongly depends on the underlying interpolation techniques. We built a novel benchmark database that is specifically tailored for interpolated slow-motion videos (KoSMo-1k). It consists of 1,350 interpolated video sequences, from 30 different content sources, along with their subjective quality ratings from up to ten subjective comparisons per video pair. Moreover, we evaluated the performance of twelve existing full-reference (FR) image/video quality assessment (I/VQA) methods on the benchmark. In this way, we are able to show that specifically tailored quality assessment methods for interpolated slow-motion videos are needed, since the evaluated methods — despite their good performance on real-time video databases — do not give satisfying results when it comes to frame interpolation. Hui Men, Vlad Hosu, Hanhe Lin, Andrés Bruhn, Dietmar Saupe |
QoMEX | 3 |
| 2020 | Foveated Video Coding for Real-Time Streaming ApplicationsabstractVideo streaming under real-time constraints is an increasingly widespread application. Many recent video encoders are unsuitable for this scenario due to theoretical limitations or run time requirements. In this paper, we present a framework for the perceptual evaluation of foveated video coding schemes. Foveation describes the process of adapting a visual stimulus according to the acuity of the human eye. In contrast to traditional region-of-interest coding, where certain areas are statically encoded at a higher quality, we utilize feedback from an eye-tracker to spatially steer the bit allocation scheme in real-time. We evaluate the performance of an H.264 based foveated coding scheme in a lab environment by comparing the bitrates at the point of just noticeable distortion (JND). Furthermore, we identify perceptually optimal codec parameterizations. In our trials, we achieve an average bitrate savings of 63.24% at the JND in comparison to the unfoveated baseline. Oliver Wiedemann, Vlad Hosu, Hanhe Lin, Dietmar Saupe |
QoMEX | 3 |
| 2020 | KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality AssessmentabstractDeep learning methods for image quality assessment (IQA) are limited due to the small size of existing datasets. Extensive datasets require substantial resources both for generating publishable content and annotating it accurately. We present a systematic and scalable approach to creating KonIQ-10k, the largest IQA dataset to date, consisting of 10,073 quality scored images. It is the first in-the-wild database aiming for ecological validity, concerning the authenticity of distortions, the diversity of content, and quality-related indicators. Through the use of crowdsourcing, we obtained 1.2 million reliable quality ratings from 1,459 crowd workers, paving the way for more general IQA models. We propose a novel, deep learning model (KonCept512), to show an excellent generalization beyond the test set (0.921 SROCC), to the current state-of-the-art database LIVE-in-the-Wild (0.825 SROCC). The model derives its core performance from the InceptionResNet architecture, being trained at a higher resolution than previous models (512 × 384). Correlation analysis shows that KonCept512 performs similar to having 9 subjective scores for each test image. Vlad Hosu, Hanhe Lin, Tamás Szirányi, Dietmar Saupe |
IEEE Trans. Image Process. | 2 |
| 2019 | SUR-Net: Predicting the Satisfied User Ratio Curve for Image Compression with Deep LearningabstractThe Satisfied User Ratio (SUR) curve for a lossy image compression scheme, e.g., JPEG, characterizes the probability distribution of the Just Noticeable Difference (JND) level, the smallest distortion level that can be perceived by a subject. We propose the first deep learning approach to predict such SUR curves. Instead of the direct approach of regressing the SUR curve itself for a given reference image, our model is trained on pairs of images, original and compressed. Relying on a Siamese Convolutional Neural Network (CNN), feature pooling, a fully connected regression-head, and transfer learning, we achieved a good prediction performance. Experiments on the MCL-JCI dataset showed a mean Bhattacharyya distance between the predicted and the original JND distributions of only 0.072. Chunling Fan, Hanhe Lin, Vlad Hosu, Yun Zhang 0002, Qingshan Jiang, Raouf Hamzaoui, Dietmar Saupe |
QoMEX | 2 |
| 2019 | KADID-10k: A Large-scale Artificially Distorted IQA DatabaseabstractCurrent artificially distorted image quality assessment (IQA) databases are small in size and limited in content. Larger IQA databases that are diverse in content could benefit the development of deep learning for IQA. We create two datasets, the Konstanz Artificially Distorted Image quality Database (KADID-10k) and the Konstanz Artificially Distorted Image quality Set (KADIS-700k). The former contains 81 pristine images, each degraded by 25 distortions in 5 levels. The latter has 140,000 pristine images, with 5 degraded versions each, where the distortions are chosen randomly. We conduct a subjective IQA crowdsourcing study on KADID-10k to yield 30 degradation category ratings (DCRs) per image. We believe that the annotated set KADID-10k, together with the unlabelled set KADIS-700k, can enable the full potential of deep learning based IQA methods by means of weakly-supervised learning. Hanhe Lin, Vlad Hosu, Dietmar Saupe |
QoMEX | 1 |
| 2019 | Visual Quality Assessment for Motion Compensated Frame InterpolationabstractCurrent benchmarks for optical flow algorithms evaluate the estimation quality by comparing their predicted flow field with the ground truth, and additionally may compare interpolated frames, based on these predictions, with the correct frames from the actual image sequences. For the latter comparisons, objective measures such as mean square errors are applied. However, for applications like image interpolation, the expected user's quality of experience cannot be fully deduced from such simple quality measures. Therefore, we conducted a subjective quality assessment study by crowdsourcing for the interpolated images provided in one of the optical flow benchmarks, the Middlebury benchmark. We used paired comparisons with forced choice and reconstructed absolute quality scale values according to Thurstone's model using the classical least squares method. The results give rise to a re-ranking of 141 participating algorithms w.r.t. visual quality of interpolated frames mostly based on optical flow estimation. Our re-ranking result shows the necessity of visual quality assessment as another evaluation metric for optical flow and frame interpolation benchmarks. Hui Men, Hanhe Lin, Vlad Hosu, Daniel Maurer 0002, Andrés Bruhn, Dietmar Saupe |
QoMEX | 2 |
| 2018 | Expertise screening in crowdsourcing image qualityabstractWe propose a screening approach to find reliable and effectively expert crowd workers in image quality assessment (IQA). Our method measures the users' ability to identify image degradations by using test questions, together with several relaxed reliability checks. We conduct multiple experiments, obtaining reproducible results with a high agreement between the expertise-screened crowd and the freelance experts of 0.95 Spearman rank order correlation (SROCC), with one restriction on the image type. Our contributions include a reliability screening method for uninformative users, a new type of test questions that rely on our proposed database1of pristine and artificially distorted images, a group agreement extrapolation method and an analysis of the crowdsourcing experiments. Vlad Hosu, Hanhe Lin, Dietmar Saupe |
QoMEX | 2 |
| 2018 | Spatiotemporal Feature Combination Model for No-Reference Video Quality AssessmentabstractOne of the main challenges in no-reference video quality assessment is temporal variation in a video. Methods typically were designed and tested on videos with artificial distortions, without considering spatial and temporal variations simultaneously. We propose a no-reference spatiotemporal feature combination model which extracts spatiotemporal information from a video, and tested it on a database with authentic distortions. Comparing with other methods, our model gave satisfying performance for assessing the quality of natural videos. Hui Men, Hanhe Lin, Dietmar Saupe |
QoMEX | 2 |
| 2018 | Disregarding the Big Picture: Towards Local Image Quality AssessmentabstractImage quality has been studied almost exclusively as a global image property. It is common practice for IQA databases and metrics to quantify this abstract concept with a single number per image. We propose an approach to blind IQA based on a convolutional neural network (patchnet) that was trained on a novel set of 32,000 individually annotated patches of 64×64 pixel. We use this model to generate spatially small local quality maps of images taken from KonIQ-10k, a large and diverse in-the-wild database of authentically distorted images. We show that our local quality indicator correlates well with global MOS, going beyond the predictive ability of quality related attributes such as sharpness. Averaging of patchnet predictions already outperforms classical approaches to global MOS prediction that were trained to include global image features. We additionally experiment with a generic second-stage aggregation CNN to estimate mean opinion scores. Our latter model performs comparable to the state of the art with a PLCC of 0.81 on KonIQ-10k. Oliver Wiedemann, Vlad Hosu, Hanhe Lin, Dietmar Saupe |
QoMEX | 3 |
| 2017 | The Konstanz natural video database (KoNViD-1k)abstractSubjective video quality assessment (VQA) strongly depends on semantics, context, and the types of visual distortions. Currently, all existing VQA databases include only a small number of video sequences with artificial distortions. The development and evaluation of objective quality assessment methods would benefit from having larger datasets of real-world video sequences with corresponding subjective mean opinion scores (MOS), in particular for deep learning purposes. In addition, the training and validation of any VQA method intended to be ‘general purpose’ requires a large dataset of video sequences that are representative of the whole spectrum of available video content and all types of distortions. We report our work on KoNViD-1k, a subjectively annotated VQA database consisting of 1,200 public-domain video sequences, fairly sampled from a large public video dataset, YFCC100m. We present the challenges and choices we have made in creating such a database aimed at ‘in the wild’ authentic distortions, depicting a wide variety of content. Vlad Hosu, Franz Götz-Hahn, Mohsen Jenadeleh, Hanhe Lin, Hui Men, Tamás Szirányi, Shujun Li 0001, Dietmar Saupe |
QoMEX | 4 |
| 2017 | Empirical evaluation of no-reference VQA methods on a natural video quality databaseabstractNo-Reference (NR) Video Quality Assessment (VQA) is a challenging task since it predicts the visual quality of a video sequence without comparison to some original reference video. Several NR-VQA methods have been proposed. However, all of them were designed and tested on databases with artificially distorted videos. Therefore, it remained an open question how well these NR-VQA methods perform for natural videos. We evaluated two popular VQA methods on our newly built natural VQA database KoNViD-1k. In addition, we found that merely combining five simple VQA-related features, i.e., contrast, colorfulness, blurriness, spatial information, and temporal information, already gave a performance about as well as those of the established NR-VQA methods. However, for all methods we found that they are unsatisfying when assessing natural videos (correlation coefficients below 0.6). These findings show that NR-VQA is not yet matured and in need of further substantial improvement. Hui Men, Hanhe Lin, Dietmar Saupe |
QoMEX | 2 |
| 2016 | Online Weighted Clustering for Real-time Abnormal Event Detection in Video SurveillanceabstractDetecting abnormal events in video surveillance is a challenging problem due to the large scale, stream fashion video data as well as the real-time constraint. In this paper, we present an online, adaptive, and real-time framework to address this problem. The spatial locations in a frame is partitioned into grids, in each grid the proposed Adaptive Multi-scale Histogram Optical Flow (AMHOF) features are extracted and modelled by an Online Weighted Clustering (OWC) algorithm. The AMHOFs which cannot be fit to a cluster with large weight are regarded as abnormal events. The OWC algorithm is simple to implement and computational efficient. In addition, we improve the detection performance by a Multiple Target Tracking (MTT) algorithm. Experimental results demonstrate our approach outperforms the state-of-the-art approaches in pixel-level rate of detection at a processing speed of 30 FPS. Hanhe Lin, Jeremiah D. Deng, Brendon J. Woodford, Ahmad Shahi |
ACM Multimedia | 1 |
| 2016 | Shot Boundary Detection Using Multi-instance Incremental and Decremental One-Class Support Vector Machine
Hanhe Lin, Jeremiah D. Deng, Brendon J. Woodford |
PAKDD (1) | 1 |
| 2015 | Anomaly detection in crowd scenes via online adaptive one-class support vector machinesabstractWe propose a novel, online adaptive one-class support vector machines algorithm for anomaly detection in crowd scenes. Integrating incremental and decremental one-class support vector machines with a sliding buffer offers an efficient and effective scheme, which not only updates the model in an online fashion with low computational cost, but also discards obsolete patterns. Our method provides a unified framework to detect both global and local anomalies. Extensive experiments have been carried out on two benchmark datasets and the comparison to the state-of-the-art methods validates the advantages of our approach. Hanhe Lin, Jeremiah D. Deng, Brendon J. Woodford |
ICIP | 1 |