VLDB 2026 Research / reviewers in the wild / expert
Qianni Zhang
dblp:50/930
· DBLP profile ↗
50ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0001-7685-2187ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 8 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Intrusion Feature Selection Method Based on Feature Distribution and Gini ImpurityabstractIntrusion detection systems (IDS) can effectively monitor network traffic and accurately detect malicious behaviors. In Internet of Things (IoT) environments, the massive influx of heterogeneous, resource-constrained devices introduces more complex security challenges, making IDS even more crucial for maintaining network security. However, the presence of redundant or irrelevant features in network traffic can significantly degrade the detection performance of IDS. To address this issue, this paper proposes a Feature distribution and Gini Impurity Filter-based intrusion feature selection method (FGIF). It combines the cardinality and Gini impurity distributions of features within a dataset to construct a multi-parameter evaluation framework, which is used to define efficient feature filtering rules that eliminate redundant and irrelevant features. Theoretical analysis demonstrates that, compared to entropy-based methods, FGIF mitigates selection bias during the feature selection process and significantly reduces computational overhead. Experiments conducted on six widely used IDS benchmark datasets and five commonly adopted classification models further confirm its effectiveness. FGIF significantly reduces feature dimensionality while maintaining detection performance comparable to that of the full feature set. Moreover, compared to existing state-of-the-art methods, FGIF achieves a superior balance between dimensionality reduction and model performance. Ying Xie 0008, Qianni Zhang, Xuyang Ding, Yongzhao Zhang, Jie Yang 0003 |
IEEE Internet Things J. | 2 |
| 2026 | MICCAI STS 2024 challenge: Semi-supervised instance-level tooth segmentation in panoramic X-ray and CBCT imagesabstractOrthopantomogram (OPGs) and Cone-Beam Computed Tomography (CBCT) are vital for dentistry, but creating large datasets for automated tooth segmentation is hindered by the labor-intensive process of manual instance-level annotation. This research aimed to benchmark and advance semi-supervised learning (SSL) as a solution for this data scarcity problem. We organized the 2nd Semi-supervised Teeth Segmentation (STS 2024) Challenge at MICCAI 2024. We provided a large-scale dataset comprising over 90,000 2D images and 3D axial slices, which includes 2380 OPG images and 330 CBCT scans, all featuring detailed instance-level FDI annotations on part of the data. The challenge attracted 114 (OPG) and 106 (CBCT) registered teams. To ensure algorithmic excellence and full transparency, we rigorously evaluated the valid, open-source submissions from the top 10 (OPG) and top 5 (CBCT) teams, respectively. All successful submissions were deep learning-based SSL methods. The winning semi-supervised models demonstrated impressive performance gains over a fully-supervised nnU-Net baseline trained only on the labeled data. For the 2D OPG track, the top method improved the Instance Affinity (IA) score by over 44 percentage points. For the 3D CBCT track, the winning approach boosted the Instance Dice score by 61 percentage points. This challenge demonstrates the potential benefit benefit of SSL for complex, instance-level medical image segmentation tasks where labeled data is scarce. The most effective approaches consistently leveraged hybrid semi-supervised frameworks that combined knowledge from foundational models like SAM with multi-stage, coarse-to-fine refinement pipelines. Both the challenge dataset and the participants' submitted code have been made publicly available on GitHub (https://github.com/ricoleehduu/STS-Challenge-2024), ensuring transparency and reproducibility. Yaqi Wang 0002, Jun Liu 0027, Jiaxue Ni, Hongyuan Zhang 0002, Jin Liu 0025, Can Han, Kaiwen Fu, Changkai Ji, Xinxu Cai, Junqiang Chen, Qianni Zhang, Dahong Qian, Shuai Wang 0003, Huiyu Zhou 0001 |
Medical Image Anal. | 20 |
| 2026 | MICCAI 2023 STS Challenge: A retrospective study of semi-supervised approaches for teeth segmentationabstractComputer-aided diagnosis greatly enhances personalized treatment planning and diagnostic efficiency by providing accurate dental anatomy through teeth segmentation. However, it still constrained by the scarcity of high-quality annotated dental datasets. To address this issue, this paper presents a dataset combining both 2D panoramic X-rays with over 6,500 images and 3D CBCT with over 580 volumes (88,500+ slices) to support the Semi-supervised Teeth Segmentation (STS) Challenge, which includes partially meticulous annotations and covers all age groups. Moreover, multi-phase semi-supervised teeth segmentation algorithms and high-confidence pseudo-labels refinement strategies were proposed by competitors during this challenge. Algorithms were verified on this proposed dataset and good segmentation performance were achieved, over 93+ and 80+ Dice score were obtained for top three 2D and 3D participants, demonstrating the high quality of this proposed dataset. This paper also summarizes the diverse methods employed by the top-ranking teams in the MICCAI 2023 STS Challenge. Our dataset is publicly accessible through Zenodo ( https://zenodo.org/records/10597292 ), and the participants’ code is hosted on GitHub ( https://github.com/ricoleehduu/STS-Challenge ). Yaqi Wang 0002, Shuai Wang 0003, Dahong Qian, Hongyuan Zhang 0002, Ruilong Dan, Qianni Zhang, Xingru Huang, Jun Liu 0027, Zhean Ma, Weiwei Cui 0003, Shan Luo 0003, Chengkai Wang, Jiaxue Ni, Dongyun Liu, Zhouhao Lin, Chunshi Wang, Qiupu Chen, Mingqian Li, Huiyu Zhou 0001, Qun Jin |
Pattern Recognit. | 10 |
| 2026 | GarmentRec: Towards Individual Garment Reconstruction From a Monocular Human ImageabstractReconstructing high-quality garment models from monocular images is important as it provides a practical and effective solution for human digitization and virtual try-on etc. Recent implicit function-based garment reconstruction methods recover free-form geometry but struggle to reconstruct individual garment meshes from human images and tend to produce disembodied limbs or degenerate shapes for novel views. In contrast, explicit parametric garment template models can be utilised to construct separate meshes and constrain the shape reconstruction robustly. However, this limits the reconstruction of garment details and shape variations, such as the wrinkles and pockets etc. To address this problem, in this paper, we introduce a novel explicit garment template that is designed for both closed and open garment topology. Powered by our new garment template, we further propose a detailed garment reconstruction method based on a monocular view that can process both the closed and open types for shape recovery. To capture those challenging parts with unknown geometry and topology, we predict displacement maps on the parameterization domain for the target garment from the monocular image and elaborate it to the 3D garment surface via the UV coordinates, achieving realistic details on the 3D garment shape. Extensive experiments demonstrate the accuracy and robustness of our method and show that realistic details like garment wrinkles and pockets can be faithfully recovered in an explicit way. The code and dataset are available at https://github.com/worryDes/GarmentRec. Zongyi Xu, Shiyang Cheng 0001, Wang Fei, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | ALCReg: Active Label Correction for Partial Point Cloud RegistrationabstractDeep point cloud registration methods encounter challenges due to partial overlaps and are heavily reliant on labeled data. In this paper, we propose ALCReg, an active label correction method for partial point cloud registration learning. ALCReg utilises a multimodal approach to generate pseudo labels, mitigating the cold-start issue in active learning. To ensure the diversity and representativeness of selected samples, we propose an inlier ratio based query strategy for manual correction. Furthermore, an innovative self-correction mechanism based on consistency is introduced, allowing the model to refine pseudo labels autonomously and further improve model performance. Experimental results on the 3DMatch and 3DLoMatch datasets demonstrate that ALCReg achieves comparable performance with the fully-supervised registration methods, even with only 5% of labeled samples, making it the first active learning method tailored for partial point cloud registration. Code is available at https://github.com/Jiang0903/ALCReg. Zongyi Xu, Xinqi Jiang, Shanshan Zhao 0001, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 5 |
| 2025 | DEFN: Dual-Encoder Fourier Group Harmonics Network for three-dimensional indistinct-boundary object segmentation
Xiaohua Jiang, Jian Huang 0015, Meiyi Luo, Zhaoyang Xu, Qianni Zhang, Xingru Huang, Shaowei Jiang, Mang Xiao |
Expert Syst. Appl. | 7 |
| 2025 | PricoMS: Prior-coordinated multiscale synthesis network for self-supervised-aided vessel segmentation in intravascular ultrasound image amidst label scarcity
Xingru Huang, Shuaibin Chen, Shaowei Jiang, Retesh Bajaj, Nathan Angelo Lecaros Yap, Murat Çap, Xiaoshuai Zhang, Xingwei He 0007, Anantharaman Ramasamy, Ryo Torii, Jouke Dijkstra, Huiyu Zhou 0001, Christos V. Bourantas, Qianni Zhang |
Knowl. Based Syst. | 15 |
| 2025 | S2Reg: Structure-semantics collaborative point cloud registration
Zongyi Xu, Xinqi Jiang, Shiyang Cheng 0001, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 5 |
| 2024 | SANGRE: a Shallow Attention Network Guided by Resolution Expansion for MR Image Segmentation
Marc E. Miquel, Qianni Zhang |
MICCAI (9) | 3 |
| 2024 | Retrieval-and-alignment based large-scale indoor point cloud semantic segmentation
Zongyi Xu, Xiaoshui Huang, Yangfu Wang, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
Sci. China Inf. Sci. | 5 |
| 2024 | Sketch-Supervised Histopathology Tumour Segmentation: Dual CNN-Transformer With Global Normalised CAMabstractDeep learning methods are frequently used in segmenting histopathology images with high-quality annotations nowadays. Compared with well-annotated data, coarse, scribbling-like labelling is more cost-effective and easier to obtain in clinical practice. The coarse annotations provide limited supervision, so employing them directly for segmentation network training remains challenging. We present a sketch-supervised method, called DCTGN-CAM, based on a dual CNN-Transformer network and a modified global normalised class activation map. By modelling global and local tumour features simultaneously, the dual CNN-Transformer network produces accurate patch-based tumour classification probabilities by training only on lightly annotated data. With the global normalised class activation map, more descriptive gradient-based representations of the histopathology images can be obtained, and inference of tumour segmentation can be performed with high accuracy. Additionally, we collect a private skin cancer dataset named BSS, which contains fine and coarse annotations for three types of cancer. To facilitate reproducible performance comparison, experts are also invited to label coarse annotations on the public liver cancer dataset PAIP2019. On the BSS dataset, our DCTGN-CAM segmentation outperforms the state-of-the-art methods and achieves 76.68 % IOU and 86.69 % Dice scores on the sketch-based tumour segmentation task. On the PAIP2019 dataset, our method achieves a Dice gain of 8.37 % compared with U-Net as the baseline network. Yilong Li 0002, Linyan Wang, Xingru Huang, Yaqi Wang 0002, Ruiquan Ge, Huiyu Zhou 0001, Juan Ye, Qianni Zhang |
IEEE J. Biomed. Health Informatics | 9 |
| 2024 | SASAN: Spectrum-Axial Spatial Approach Networks for Medical Image SegmentationabstractOphthalmic diseases such as central serous chorioretinopathy (CSC) significantly impair the vision of millions of people globally. Precise segmentation of choroid and macular edema is critical for diagnosing and treating these conditions. However, existing 3D medical image segmentation methods often fall short due to the heterogeneous nature and blurry features of these conditions, compounded by medical image clarity issues and noise interference arising from equipment and environmental limitations. To address these challenges, we propose the Spectrum Analysis Synergy Axial-Spatial Network (SASAN), an approach that innovatively integrates spectrum features using the Fast Fourier Transform (FFT). SASAN incorporates two key modules: the Frequency Integrated Neural Enhancer (FINE), which mitigates noise interference, and the Axial-Spatial Elementum Multiplier (ASEM), which enhances feature extraction. Additionally, we introduce the Self-Adaptive Multi-Aspect Loss (LSM), which balances image regions, distribution, and boundaries, adaptively updating weights during training. We compiled and meticulously annotated the Choroid and Macular Edema OCT Mega Dataset (CMED-18k), currently the world’s largest dataset of its kind. Comparative analysis against 13 baselines shows our method surpasses these benchmarks, achieving the highest Dice scores and lowest HD95 in the CMED and OIMHS datasets. Our code is publicly available at https://github.com/IMOP-lab/SASAN-Pytorch. Xingru Huang, Jian Huang 0015, Tianyun Zhang, Changpeng Yue, Xuanbin Chen, Qianni Zhang, Ying Fu 0001, Yangyundou Wang |
IEEE Trans. Medical Imaging | 10 |
| 2024 | IGReg: Image-Geometry-Assisted Point Cloud Registration via Selective Correlation FusionabstractPoint cloud registration suffers from repeated patterns and low geometric structures in indoor scenes. The recent transformer utilises attention mechanism to capture the global correlations in feature space and improves the registration performance. However, for indoor scenarios, global correlation loses its advantages as it cannot distinguish real useful features and noise. To address this problem, we propose an image-geometry-assisted point cloud registration method by integrating image information into point features and selectively fusing the geometric consistency with respect to reliable salient areas. Firstly, an Intra-Image-Geometry fusion module is proposed to integrate the texture and structure information into the point feature space by the cross-attention mechanism. Initial corresponding superpoints are acquired as salient anchors in the source and target. Then, a selective correlation fusion module is designed to embed the correlations between the salient anchors and points. During training, the saliency location and selective correlation fusion modules exchange information iteratively to identify the most reliable salient anchors and achieve effective feature fusion. The obtained distinctive point cloud features allow for accurate correspondence matching, leading to the success of indoor point cloud registration. Extensive experiments are conducted on 3DMatch and 3DLoMatch datasets to demonstrate the outstanding performance of the proposed approach compared to the state-of-the-art, particularly in those geometrically challenging cases such as repetitive patterns and low-geometry regions. Zongyi Xu, Xinqi Jiang, Changjun Gu, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | Hierarchical Point-based Active Learning for Semi-supervised Point Cloud Semantic SegmentationabstractImpressive performance on point cloud semantic segmentation has been achieved by fully-supervised methods with large amounts of labelled data. As it is labour-intensive to acquire large-scale point cloud data with point-wise labels, many attempts have been made to explore learning 3D point cloud segmentation with limited annotations. Active learning is one of the effective strategies to achieve this purpose but is still under-explored. The most recent methods of this kind measure the uncertainty of each pre-divided region for manual labelling but they suffer from redundant information and require additional efforts for region division. This paper aims at addressing this issue by developing a hierarchical point-based active learning strategy. Specifically, we measure the uncertainty for each point by a hierarchical minimum margin uncertainty module which considers the contextual information at multiple levels. Then, a feature-distance suppression strategy is designed to select important and representative points for manual labelling. Besides, to better exploit the unlabelled data, we build a semi-supervised segmentation framework based on our active strategy. Extensive experiments on the S3DIS and ScanNetV2 datasets demonstrate that the proposed framework achieves 96.5% and 100% performance of fully-supervised baseline with only 0.07% and 0.1% training data, respectively, outperforming the state-of-the-art weakly-supervised and active learning methods. The code will be available at https://github.com/SmiletoE/HPAL. Zongyi Xu, Shanshan Zhao 0001, Qianni Zhang, Xinbo Gao 0001 |
ICCV | 4 |
| 2023 | Joint Dense-Point Representation for Contour-Aware Graph Segmentation
Kit Mills Bransby, Gregory Slabaugh, Christos V. Bourantas, Qianni Zhang |
MICCAI (3) | 4 |
| 2023 | POST-IVUS: A perceptual organisation-aware selective transformer framework for intravascular ultrasound segmentationabstractIntravascular ultrasound (IVUS) is recommended in guiding coronary intervention. The segmentation of coronary lumen and external elastic membrane (EEM) borders in IVUS images is a key step, but the manual process is time-consuming and error-prone, and suffers from inter-observer variability. In this paper, we propose a novel perceptual oganisation-aware selective transformer framework that can achieve accurate and robust segmentation of the vessel walls in IVUS images. In this framework, temporal context-based feature encoders extract efficient motion features of vessels. Then, a perceptual oganisation-aware selective transformer module is proposed to extract accurate boundary information, supervised by a dedicated boundary loss. The obtained EEM and lumen segmentation results will be fused in a temporal constraining and fusion module, to determine the most likely correct boundaries with robustness to morphology. Our proposed methods are extensively evaluated in non-selected IVUS sequences, including normal, bifurcated, and calcified vessels with shadow artifacts. The results show that the proposed methods outperform the state-of-the-art, with a Jaccard measure of 0.92 for lumen and 0.94 for EEM on the IVUS 2011 open challenge dataset. This work has been integrated into a software QCU-CMS2 to automatically segment IVUS images in a user-friendly environment. Xingru Huang, Retesh Bajaj, Yilong Li 0002, Xin Ye 0006, Ji Lin 0004, Francesca Pugliese, Anantharaman Ramasamy, Yaqi Wang 0002, Ryo Torii, Jouke Dijkstra, Huiyu Zhou 0001, Christos V. Bourantas, Qianni Zhang |
Medical Image Anal. | 14 |
| 2023 | Stimulus-guided adaptive transformer network for retinal blood vessel segmentation in fundus imagesabstractAutomated retinal blood vessel segmentation in fundus images provides important evidence to ophthalmologists in coping with prevalent ocular diseases in an efficient and non-invasive way. However, segmenting blood vessels in fundus images is a challenging task, due to the high variety in scale and appearance of blood vessels and the high similarity in visual features between the lesions and retinal vascular. Inspired by the way that the visual cortex adaptively responds to the type of stimulus, we propose a Stimulus-Guided Adaptive Transformer Network (SGAT-Net) for accurate retinal blood vessel segmentation. It entails a Stimulus-Guided Adaptive Module (SGA-Module) that can extract local-global compound features based on inductive bias and self-attention mechanism. Alongside a light-weight residual encoder (ResEncoder) structure capturing the relevant details of appearance, a Stimulus-Guided Adaptive Pooling Transformer (SGAP-Former) is introduced to reweight the maximum and average pooling to enrich the contextual embedding representation while suppressing the redundant information. Moreover, a Stimulus-Guided Adaptive Feature Fusion (SGAFF) module is designed to adaptively emphasize the local details and global context and fuse them in the latent space to adjust the receptive field (RF) based on the task. The evaluation is implemented on the largest fundus image dataset (FIVES) and three popular retinal image datasets (DRIVE, STARE, CHASEDB1). Experimental results show that the proposed method achieves a competitive performance over the other existing method, with a clear advantage in avoiding errors that commonly happen in areas with highly similar visual features. The sourcecode is publicly available at: https://github.com/Gins-07/SGAT. Ji Lin 0004, Xingru Huang, Huiyu Zhou 0001, Yaqi Wang 0002, Qianni Zhang |
Medical Image Anal. | 5 |
| 2023 | MSCA-Net: Multi-scale contextual attention network for skin lesion segmentationabstractLesion segmentation algorithms automatically outline lesion areas in medical images, facilitating more effective identification and assessment of the clinically relevant features, and improving the efficacy and diagnosis accuracy. However, most fully convolutional network based segmentation methods suffer from spatial and contextual information loss when decreasing image resolution. To overcome this shortcoming, this paper proposes a skin lesion segmentation model , namely, the Multi-Scale Contextual Attention Network (MSCA-Net), which can exploit the multi-scale contextual information in images. Inspired by the skip connection of U-Net, we design a multi-scale bridge (MSB) module which interacts with multi-scale features to effectively fuse the multi-scale contextual information of the encoder and decoder path features. We further propose a global-local channel spatial attention module (GL-CSAM), aiming at capturing global contextual information. In addition, to take full advantage of the multi-scale features of the decoder, we propose a scale-aware deep supervision (SADS) module to achieve hierarchical iterative deep supervision. Comprehensive experimental results on the public dataset of ISIC 2017, ISIC 2018, and PH 2 show that our proposed method outperforms other state-of-the-art methods, demonstrating the efficacy of our method in skin lesion segmentation. Our code is available at https://github.com/YonghengSun1997/MSCA-Net . Yongheng Sun, Duwei Dai, Qianni Zhang, Yaqi Wang 0002, Songhua Xu, Chunfeng Lian |
Pattern Recognit. | 3 |
| 2023 | Cost-Sensitive Boosting Pruning Trees for Depression Detection on TwitterabstractDepression is one of the most common mental health disorders, and a large number of depressed people commit suicide each year. Potential depression sufferers usually do not consult psychological doctors because they feel ashamed or are unaware of any depression, which may result in severe delay of diagnosis and treatment. In the meantime, evidence shows that social media data provides valuable clues about physical and mental health conditions. In this paper, we argue that it is feasible to identify depression at an early stage by mining online social behaviours. Our approach, which is innovative to the practice of depression detection, does not rely on the extraction of numerous or complicated features to achieve accurate depression detection. Instead, we propose a novel classifier, namely, Cost-sensitive Boosting Pruning Trees (CBPT), which demonstrates a strong classification ability on two publicly accessible Twitter depression detection datasets. To comprehensively evaluate the classification capability of CBPT, we use additional three datasets from the UCI machine learning repository and CBPT obtains appealing classification results against several state of the arts boosting algorithms. Finally, we comprehensively explore the influence factors for the model prediction, and the results manifest that our proposed framework is promising for identifying Twitter users with depression. Zheheng Jiang, Feixiang Zhou, Long Chen 0019, Jialin Lyu, Xiangrong Zhang, Qianni Zhang, Abdul Hamid Sadka, Yinhai Wang, Ling Li 0010, Huiyu Zhou 0001 |
IEEE Trans. Affect. Comput. | 8 |
| 2022 | Complexity Reduction of Learned In-Loop Filtering in Video CodingabstractIn video coding, in-loop filters are applied on reconstructed video frames to enhance their perceptual quality, before storing the frames for output. Conventional in-loop filters are obtained by hand-crafted methods. Recently, learned filters based on convolutional neural networks that utilize attention mechanisms have been shown to improve upon traditional techniques. However, these solutions are typically significantly more computationally expensive, limiting their potential for practical applications. The proposed method uses a novel combination of sparsity and structured pruning for complexity reduction of learned in-loop filters. This is done through a three-step training process of magnitude-guided weight pruning, insignificant neuron identification and removal, and fine-tuning. Through initial tests we find that network parameters can be significantly reduced with a minimal impact on network performance. Woody Bayliss, Luka Murn, Ebroul Izquierdo, Qianni Zhang, Marta Mrak |
ISCAS | 4 |
| 2022 | Multimodal Gait Recognition for Neurodegenerative DiseasesabstractIn recent years, single modality-based gait recognition has been extensively explored in the analysis of medical images or other sensory data, and it is recognized that each of the established approaches has different strengths and weaknesses. As an important motor symptom, gait disturbance is usually used for diagnosis and evaluation of diseases; moreover, the use of multimodality analysis of the patient's walking pattern compensates for the one-sidedness of single modality gait recognition methods that only learn gait changes in a single measurement dimension. The fusion of multiple measurement resources has demonstrated promising performance in the identification of gait patterns associated with individual diseases. In this article, as a useful tool, we propose a novel hybrid model to learn the gait differences between three neurodegenerative diseases, between patients with different severity levels of Parkinson's disease, and between healthy individuals and patients, by fusing and aggregating data from multiple sensors. A spatial feature extractor (SFE) is applied to generating representative features of images or signals. In order to capture temporal information from the two modality data, a new correlative memory neural network (CorrMNN) architecture is designed for extracting temporal features. Afterward, we embed a multiswitch discriminator to associate the observations with individual state estimations. Compared with several state-of-the-art techniques, our proposed framework shows more accurate classification results. Aite Zhao, Junyu Dong, Lin Qi 0004, Qianni Zhang, Ning Li 0012, Xin Wang 0068, Huiyu Zhou 0001 |
IEEE Trans. Cybern. | 5 |
| 2022 | AGMB-Transformer: Anatomy-Guided Multi-Branch Transformer Network for Automated Evaluation of Root Canal TherapyabstractAccurate evaluation of the treatment result on X-ray images is a significant and challenging step in root canal therapy since the incorrect interpretation of the therapy results will hamper timely follow-up which is crucial to the patients’ treatment outcome. Nowadays, the evaluation is performed in a manual manner, which is time-consuming, subjective, and error-prone. In this article, we aim to automate this process by leveraging the advances in computer vision and artificial intelligence, to provide an objective and accurate method for root canal therapy result assessment. A novel anatomy-guided multi-branch Transformer (AGMB-Transformer) network is proposed, which first extracts a set of anatomy features and then uses them to guide a multi-branch Transformer network for evaluation. Specifically, we design a polynomial curve fitting segmentation strategy with the help of landmark detection to extract the anatomy features. Moreover, a branch fusion module and a multi-branch structure including our progressive Transformer and Group Multi-Head Self-Attention (GMHSA) are designed to focus on both global and local features for an accurate diagnosis. To facilitate the research, we have collected a large-scale root canal therapy evaluation dataset with 245 root canal therapy X-ray images, and the experiment results show that our AGMB-Transformer can improve the diagnosis accuracy from 57.96% to 90.20% compared with the baseline network. The proposed AGMB-Transformer can achieve a highly accurate evaluation of root canal therapy. To our best knowledge, our work is the first to perform automatic root canal therapy evaluation and has important clinical value to reduce the workload of endodontists. Guodong Zeng, Jun Wang 0041, Qun Jin, Lingling Sun, Qianni Zhang, Qisi Lian, Guiping Qian, Neng Xia, Ruizi Peng, Shuai Wang 0003, Yaqi Wang 0002 |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | SRPN: similarity-based region proposal networks for nuclei and cells detection in histology images
Yibao Sun, Xingru Huang, Huiyu Zhou 0001, Qianni Zhang |
Medical Image Anal. | 4 |
| 2021 | Perceptual Underwater Image Enhancement With Deep Learning and Physical PriorsabstractUnderwater image enhancement, as a pre-processing step to support the following object detection task, has drawn considerable attention in the field of underwater navigation and ocean exploration. However, most of the existing underwater image enhancement strategies tend to consider enhancement and detection as two fully independent modules with no interaction, and the practice of separate optimisation does not always help the following object detection task. In this article, we propose two perceptual enhancement models, each of which uses a deep enhancement model with a detection perceptor. The detection perceptor provides feedback information in the form of gradients to guide the enhancement model to generate patch level visually pleasing or detection favourable images. In addition, due to the lack of training data, a hybrid underwater image synthesis model, which fuses physical priors and data-driven cues, is proposed to synthesise training data and generalise our enhancement model for real-world underwater images. Experimental results show the superiority of our proposed method over several state-of-the-art methods on both real-world and synthetic underwater datasets. Long Chen 0019, Zheheng Jiang, Aite Zhao, Qianni Zhang, Junyu Dong, Huiyu Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Deep Learning Methods for Lung Cancer Segmentation in Whole-Slide Histopathology Images - The ACDC@LungHP Challenge 2019abstractAccurate segmentation of lung cancer in pathology slides is a critical step in improving patient care. We proposed the ACDC@LungHP (Automatic Cancer Detection and Classification in Whole-slide Lung Histopathology) challenge for evaluating different computer-aided diagnosis (CADs) methods on the automatic diagnosis of lung cancer. The ACDC@LungHP 2019 focused on segmentation (pixel-wise detection) of cancer tissue in whole slide imaging (WSI), using an annotated dataset of 150 training images and 50 test images from 200 patients. This paper reviews this challenge and summarizes the top 10 submitted methods for lung cancer segmentation. All methods were evaluated using metrics using the precision, accuracy, sensitivity, specificity, and DICE coefficient (DC). The DC ranged from 0.7354 ±0.1149 to 0.8372 ±0.0858. The DC of the best method was close to the inter-observer agreement (0.8398 ±0.0890). All methods were based on deep learning and categorized into two groups: multi-model method and single model method. In general, multi-model methods were significantly better (p 0.01) than single model methods, with mean DC of 0.7966 and 0.7544, respectively. Deep learning based methods could potentially help pathologists find suspicious regions for further analysis of lung cancer in WSI. Tao Tan 0002, Xichao Teng, Xiaoliang Sun, Lihong Liu, Byungjae Lee, Yilong Li 0002, Qianni Zhang, Shujiao Sun, Yushan Zheng, Junyu Yan, Yiyu Hong, Junsu Ko, Hyun Jung, Ching-Wei Wang, Vladimir Yurovskiy, Pavel Maevskikh, Vahid Khanagha, Daiqiang Li, Peter J. Schüffler, Hui Chen 0020, Yuling Tang, Geert Litjens 0001 |
IEEE J. Biomed. Health Informatics | 11 |
| 2021 | CANet: Context Aware Network for Brain Glioma SegmentationabstractAutomated segmentation of brain glioma plays an active role in diagnosis decision, progression monitoring and surgery planning. Based on deep neural networks, previous studies have shown promising technologies for brain glioma segmentation. However, these approaches lack powerful strategies to incorporate contextual information of tumor cells and their surrounding, which has been proven as a fundamental cue to deal with local ambiguity. In this work, we propose a novel approach named Context-Aware Network (CANet) for brain glioma segmentation. CANet captures high dimensional and discriminative features with contexts from both the convolutional space and feature interaction graphs. We further propose context guided attentive conditional random fields which can selectively aggregate features. We evaluate our method using publicly accessible brain glioma segmentation datasets BRATS2017, BRATS2018 and BRATS2019. The experimental results show that the proposed algorithm has better or competitive performance against several State-of-The-Art approaches under different segmentation metrics on the training and validation sets. Long Chen 0019, Feixiang Zhou, Zheheng Jiang, Qianni Zhang, Yinhai Wang, Caifeng Shan, Ling Li 0010, Huiyu Zhou 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Building High-Fidelity Human Body Models From User-Generated DataabstractWe propose a key point-based approach, refers to asKPhub-PC, to estimate high-fidelity human body models from low-quality point clouds acquired with an affordable 3D scanner and a variationKPhub-Ithat can achieve the same purpose based on low-resolution single images taken by smartphones. In KPhub-PC, a sparse set of key points is annotated to guide the deformation of a parametric 3D human body model SMPL and then a high-fidelity human body model that can explain the target point clouds is built. Besides building 3D human body models from point clouds, KPhub-I is designed to estimate accurate 3D human body models from single 2D images. The SMPL model is fitted to 2D joints and the boundary of the human body which are detected using CNN based methods automatically. Considering that people are in stable poses most of the time, a stable pose prior is defined from CMU motion capture dataset for further improving accuracy. Extensive experiments demonstrate that in both types of user-generated data, the proposed approaches can build believable and animatable human body models robustly. Our approach outperforms the state-of-the-arts in the accuracy of both human body shape and pose estimation. Zongyi Xu, Yindi Zhu, Huiyu Zhou 0001, Qianni Zhang |
IEEE Trans. Multim. | 6 |
| 2020 | One Shot Logo Recognition Based on Siamese Neural NetworksabstractThis work presents an approach for one-shot logo recognition that relies on a Siamese neural network (SNN) embedded with a pre-trained model that is fine-tuned on a challenging logo dataset. Although the model is fine-tuned using logo images, the training and testing datasets do not have overlapped categories; meaning that, all the classes used for testing the one-shot recognition framework remain unseen during the fine-tuning process. The recognition process follows the standard SNN approach in which a pair of input images are encoded by each sister network. The encoded outputs for each image are afterwards compared using a trained metric and thresholded to define matches and mismatches. The proposed approach achieves an accuracy of 77.07% under the one-shot constraints in the QMUL-OpenLogo dataset. Code is available at https://github.com/cjvargasc/oneshot_siamese/. Camilo Vargas, Qianni Zhang, Ebroul Izquierdo |
ICMR | 2 |
| 2018 | Region Based User-Generated Human Body Scan RegistrationabstractWe present a region-based registration method to robustly register low-quality human body scans that are acquired with cost-effective devices accessible to general users. When traditional closest point based registration approaches are performed on these noisy scan data, it is easy to fall into local minimum. To address this problem, we learn prior knowledge of body shape from publicly available dataset and combine it with the Iterative Closest Point (ICP) algorithm. Firstly, sparse markers are used to change pose of template, making it perform in the same way as target scans do. In the registration stage, the holistic shape model for the basic figure of human and a set of local shape models for describing the details of each human body part are trained. We fit the holistic model roughly to the target mesh. To capture more body details, we combine local shape models with the non-rigid ICP method to deform the template part-by-part. Extensive experiments over data scanned using devices from professional to low-cost types verify that our approach is both accurate and robust to incomplete and noisy data. Zongyi Xu, Qianni Zhang |
ICME | 2 |
| 2018 | Detection of Breast Tumour Tissue Regions in Histopathological Images using Convolutional Neural NetworksabstractDuctal carcinoma in situ (DCIS) is considered a pre-invasive breast cancer and sometimes it can develop into an invasive ductal carcinoma. The analysis of histopathological images to detect tumour border of DCIS could provide important information for better diagnosis of patients. We present a deep learning based system to automatically identify DCIS in histopathological images. Specifically, a convolutional neural network (CNN) is first trained to predict labels of small patches cropped out of a histopathological whole slide image. Next, a sliding window method is used to produce a probability map of DCIS. Finally, given the probability map, a tumor border of DCIS is produced and delineated with the method of Marching Cubes to facilitate pathologists' review and assessment. Evaluation of cross validation demonstrates that the CNN model of GoogleNet performs well in histology image patch classification with an overall accuracy of (98.46±0.40)% and identifies the DCIS tissue patches with a F1-score of (97.40±1.18)% (mean±variance). Moreover, around 95.6% tumour tissue within the enclosed tumour regions can be identified by our developed method. Finally, the goal of tumor border detection can be well achieved with a few post-processing steps. Yibao Sun, Zhaoyang Xu, Carina Strell, Carlos Fernández Moro, Fredrik Warnberg, Qianni Zhang |
IPAS | 7 |
| 2018 | Multilevel active registration for kinect human body scans: from low quality to high qualityabstractRegistration of 3D human body has been a challenging research topic for over decades. Most of the traditional human body registration methods require manual assistance, or other auxiliary information such as texture and markers. The majority of these methods are tailored for high-quality scans from expensive scanners. Following the introduction of the low-quality scans from cost-effective devices such as Kinect, the 3D data capturing of human body becomes more convenient and easier. However, due to the inevitable holes, noises and outliers in the low-quality scan, the registration of human body becomes even more challenging. To address this problem, we propose a fully automatic active registration method which deforms a high-resolution template mesh to match the low-quality human body scans. Our registration method operates on two levels of statistical shape models: (1) the first level is a holistic body shape model that defines the basic figure of human; (2) the second level includes a set of shape models for every body part, aiming at capturing more body details. Our fitting procedure follows a coarse-to-fine approach that is robust and efficient. Experiments show that our method is comparable with the state-of-the-art methods for high-quality meshes in terms of accuracy and it outperforms them in the case of low-quality scans where noises, holes and obscure parts are prevalent. Zongyi Xu, Qianni Zhang, Shiyang Cheng 0001 |
Multim. Syst. | 2 |
| 2018 | ADORE: An Adaptive Holons Representation Framework for Human Pose EstimationabstractIn this paper, the problem of human pose estimation in a 2D still image is addressed. A framework called adaptive holons representation (ADORE) that takes advantage of local and global cues is proposed to improve the pose estimation accuracy. In particular, ADORE is made up of two components: 1) the holons part, independent losses pose nets (ILPNs) is designed to first infer joints location on the global level; and 2) the adaptive part, convolutional local detectors (CLDs) is proposed to subsequently detect the joints in the potential regions generated by ILPN. Pose estimation is formulated as a classification problem toward body joints in ILPN, which consists of two independent loss layers that, respectively, instruct the learning of x and y coordinates of a joint. Experimental results on two challenging benchmark tasks demonstrate that our proposed framework is more efficient than other deep models and has desirable performance. Xiuyuan Chen, Qianni Zhang, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | CUNet: A Compact Unsupervised Network For Image ClassificationabstractIn this paper, we propose a compact network called compact unsupervised network (CUNet) to address the image classification challenge. Contrasting the usual learning approach of convolutional neural networks, learning is achieved by the simple K-means on diverse image patches. This approach performs well even with scarcely labeled training images, greatly reducing the computational cost, while maintaining high discriminative power. Furthermore, we propose a new weighted pooling method in which different weighting values of adjacent neurons are considered. This strategy leads to improved classification since the network becomes more robust against small image distortions. In the output layer, CUNet integrates feature maps obtained in the last hidden layer, and straightforwardly computes histograms in nonoverlapped blocks. To reduce feature redundancy, we also implement the max-pooling operation on adjacent blocks to select the most competitive features. Comprehensive experiments on well-established databases are conducted to validate the classification performances of the introduced CUNet approach. Mengdie Mao, Gaipeng Kong, Xi Wu 0004, Qianni Zhang, Xiaochun Cao, Ebroul Izquierdo |
IEEE Trans. Multim. | 6 |
| 2016 | Formal representation of events in a surveillance domain ontologyabstractFollowing the exponential deployment of surveillance systems across a wide-spread region of geographic locations, detection and representation of events has become a critical element in automated surveillance systems. In this paper, we present an extensive ontology framework for representing complex semantic events. The proposed ontology builds on DOLCE ontology and relies on the linguistic and cognitive modelling of philosophical knowledge to achieve interoperability between proprietary surveillance systems. The explicit definition of event vocabulary presented in the paper is aimed at aiding forensic analysts to objectively identify and represent complex events. The expressiveness of the proposed ontology framework is described in the context of London Riots which took place in 2011. Faranak Sobhani, Krishna Chandramouli, Qianni Zhang, Ebroul Izquierdo |
ICIP | 3 |
| 2016 | Symmetry-Aware Human Shape Correspondence Using Skeleton
Zongyi Xu, Qianni Zhang |
MMM (1) | 2 |
| 2016 | Decomposition and matching: Towards efficient automatic Chinese character stroke extractionabstractIn this paper, given images of Chinese characters, we present an automatic stroke extraction system, which consists of character decomposition, shape matching, cross area extraction and stroke segment combination. First, a character is decomposed into isolated stroke structures according to the connection. Then, we extract shape contexts of stroke structures and find the matched counterparts in a standard database by shape matching. Cross points are computed, and cross areas are extracted according to a proposed adaptive cross area extraction method based on point-to-boundary orientation distance. For a shape structure with the matched structure, we optimize cross point set and combine its stroke segments according to the correct cross point set and combination way of its matched counterpart. For those without matched structures, we propose an angle based stroke segment combination method to combine the segments into a complete stroke. Experimental results indicate that the proposed system achieves high accuracy and demonstrates prominent augmentation on efficiency. Zongyi Xu, Qianni Zhang, Ebroul Izquierdo |
VCIP | 3 |
| 2016 | Optimal kernel choice for domain adaption learning
Ning Feng, Pinjie Quan, Gaipeng Kong, Xiuyuan Chen, Qianni Zhang |
Eng. Appl. Artif. Intell. | 6 |
| 2016 | LSI: Latent semantic inference for natural image segmentation
Ning Feng, Qianni Zhang |
Pattern Recognit. | 3 |
| 2016 | Holons Visual Representation for Image RetrievalabstractAlong with the enlargement of image scale, convolutional local features, such as SIFT, are ineffective for representing or indexing and more compact visual representations are required. Due to the intrinsic mechanism, the state-of-the-art vector of locally aggregated descriptors (VLAD) has a few limits. Based on this, we propose a new descriptor named holons visual representation (HVR). The proposed HVR is a derivative mutational self-contained combination of global and local information. It exploits both global characteristics and the statistic information of local descriptors in the image dataset. It also takes advantages of local features of each image and computes their distribution with respect to the entire local descriptor space. Accordingly, the HVR is computed by a two-layer hierarchical scheme, which splits the local feature space and obtains raw partitions, as well as the corresponding refined partitions. Then, according to the distances from the centroids of partition spaces to local features and their spatial correlation, we assign the local features into their nearest raw partitions and refined partitions to obtain the global description of an image. Compared with VLAD, HVR holds critical structure information and enhances the discriminative power of individual representation with a small amount of computation cost, while using the same memory overhead. Extensive experiments on several benchmark datasets demonstrate that the proposed HVR outperforms conventional approaches in terms of scalability as well as retrieval accuracy for images with similar intra local information. Gaipeng Kong, Qianni Zhang, Xiaochun Cao, Ebroul Izquierdo |
IEEE Trans. Multim. | 4 |
| 2015 | Action Recognition based on Subdivision-Fusion ModelabstractThis paper proposes a novel Subdivision-Fusion Model (SFM) to recognize human actions. In most action recognition tasks, overlapping feature distribution is a common problem leading to overfitting. In the subdivision stage of the proposed SFM, samples in each category are clustered. Then, such samples are grouped into multiple more concentrated subcategories. Boundaries for the subcategories are easier to find and as consequence overfitting is avoided. In the subsequent fusion stage, the multi-subcategories classification results are converted back to the original category recognition problem. Two methods to determine the number of clusters are provided. The proposed model has been thoroughly tested with four popular datasets. In the Hollywood2 dataset, an accuracy of 79.4% is achieved, outperforming the state-of-the-art accuracy of 64.3%. The performance on the YouTube Action dataset has been improved from 75.8% to 82.5%, while considerably improvements are also observed on the KTH and UCF50 datasets. Zong Bo Hao, Linlin Lu, Qianni Zhang, Ebroul Izquierdo, Juanyu Yang |
BMVC | 3 |
| 2015 | Discriminative Light Unsupervised Learning Network for Image Representation and ClassificationabstractThis paper proposes a discriminative light unsupervised learning network (DLUN) to counter the image classification challenge. Compared with the traditional convolutional networks learning filters by the time-consuming stochastic gradient descent, DLUN learns the filter bank from diverse image patches with the classical K-means, which significantly reduces the training complexity while maintains the high discriminative ability. Besides, we design a new pooling strategy named voting pooling which considers the contribution difference of the adjacent activations. In the output layer, DLUN computes histograms in the size-changed dense sliding windows, followed by a max pooling operation on histogram bins at different scales to obtain the most competitive features. The classification performance on two widely used benchmarks verifies that DLUN is competitive among some state-of-the-arts. Qianni Zhang |
ACM Multimedia | 3 |
| 2014 | Accurate stereo 3D point cloud generation suitable for multi-view stereo reconstructionabstractThis paper proposes a novel methodology for generating 3D point clouds of good accuracy from stereo pairs. Initially, the methodology defines some conditions for the proper selection of image pairs. Then, the selected stereo images are used to estimate dense correspondences using the Daisy descriptor. An efficient two-phase strategy to remove outliers is then introduced. Finally, the 3D point cloud is refined by combining sub-pixel accuracy correspondences estimation and the moving least squares algorithm. The proposed methodology can be exploited by multi-view stereo algorithms due to its good accuracy and its fast computation. Georgios Kordelas, Petros Daras, Patrycia Klavdianos, Ebroul Izquierdo, Qianni Zhang |
VCIP | 5 |
| 2013 | Human action recognition by fast dense trajectoriesabstractIn this paper, we propose the fast dense trajectories algorithm for human action recognition. Dense trajectories are robust to fast irregular motions and outperform other state-of-the-art descriptors such as KLT tracker or SIFT descriptors. However, the use of dense trajectories is time consuming. To improve the efficiency, we extract feature trajectories in the ROI rather than in the whole frames, and we use the temporal pyramids to achieve adaptable mechanism for different action speed. We evaluate the method on the dataset of Huawei/3DLife -- 3D human reconstruction and action recognition Grand Challenge in ACM Multimedia 2013. Experimental results show a significant improvement over the dense trajectories descriptor in real-time, and adaptable to different speed. Zong Bo Hao, Qianni Zhang, Ebroul Izquierdo, Nan Sang |
ACM Multimedia | 2 |
| 2013 | Histology Image Retrieval in Optimized Multifeature SpacesabstractContent based histology image retrieval systems have shown great potential in supporting decision making in clinical activities, teaching, and biological research. In content based image retrieval, feature combination plays a key role. It aims at enhancing the descriptive power of visual features corresponding to semantically meaningful queries. It is particularly valuable in histology image analysis where intelligent mechanisms are needed for interpreting varying tissue composition and architecture into histological concepts. This paper presents an approach to automatically combine heterogeneous visual features for histology image retrieval. The aim is to obtain the most representative fusion model for a particular keyword that is associated to multiple query images. The core of this approach is a multi-objective learning method, which aims to understand an optimal visual-semantic matching function by jointly considering the different preferences of the group of query images. The task is posed as an optimisation problem, and a multi-objective optimisation strategy is employed in order to handle potential contradictions in the query images associated to the same keyword. Experiments were performed on two different collections of histology images. The results show that it is possible to improve a system for content based histology image retrieval by using an appropriately defined multi-feature fusion model, which takes careful consideration of the structure and distribution of visual features. Qianni Zhang, Ebroul Izquierdo |
IEEE J. Biomed. Health Informatics | 1 |
| 2013 | Multifeature analysis and semantic context learning for image classificationabstractThis article introduces an image classification approach in which the semantic context of images and multiple low-level visual features are jointly exploited. The context consists of a set of semantic terms defining the classes to be associated to unclassified images. Initially, a multiobjective optimization technique is used to define a multifeature fusion model for each semantic class. Then, a Bayesian learning procedure is applied to derive a context model representing relationships among semantic classes. Finally, this context model is used to infer object classes within images. Selected results from a comprehensive experimental evaluation are reported to show the effectiveness of the proposed approaches. Qianni Zhang, Ebroul Izquierdo |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | Enhanced visualisation of dance performance from automatically synchronised multimodal recordingsabstractThe Huawei/3DLife Grand Challenge Dataset provides multimodal recordings of Salsa dancing, consisting of audiovisual streams along with depth maps and inertial measurements. In this paper, we propose a system for augmented reality-based evaluations of Salsa dancer performances. An essential step for such a system is the automatic temporal synchronisation of the multiple modalities captured from different sensors, for which we propose efficient solutions. Furthermore, we contribute modules for the automatic analysis of dance performances and present an original software application, specifically designed for the evaluation scenario considered, which enables an enhanced dance visualisation experience, through the augmentation of the original media with the results of our automatic analyses. Marc Gowing, Philip Kelly, Noel E. O'Connor, Cyril Concolato, Slim Essid, Jean Le Feuvre, Robin Tournemenne, Ebroul Izquierdo, Vlado Kitanovski, Qianni Zhang |
ACM Multimedia | 11 |
| 2010 | 3DLife: Bringing the Media Internet to Life
Ebroul Izquierdo, Tomas Piatrik, Qianni Zhang |
ISoLA (2) | 3 |
| 2010 | ACM workshop on surreal media and virtual cloningabstractThis paper gives an overview of ACM Multimedia 2010 Workshop on Surreal Media and Virtual Cloning, including research work towards the creation of surreal media and realistic 3D virtual environments where virtual humans and objects can interact remotely. The primary objective is to discuss key research issues related to the generation of surreal media and 3D cooperative virtual worlds. We expect that the one-day program will bring together research groups from related fields and explore research problems, potential applications and collaborative opportunities. Ebroul Izquierdo, Yang Cai 0002, Qianni Zhang, Manuel García-Herranz |
ACM Multimedia | 3 |
| 2007 | Adaptive salient block-based image retrieval in multi-feature space
Qianni Zhang, Ebroul Izquierdo |
Signal Process. Image Commun. | 1 |
| 2006 | Optimizing Metrics Combining Low-Level Visual Descriptors for Image Annotation and RetrievalabstractAn object oriented approach for key-word based image annotation and classification is presented. It considers combinations of low-level descriptors and suitable metrics to represent and measure similarity between semantically meaningful objects. The objective is to obtain "optimal" metrics based on a linear combination of single metrics and descriptors in a multi-feature space. The proposed approach estimates an optimal linear combination of predefined metrics by applying a Multi-Objective Optimization technique based on a Pareto Archived Evolution Strategy. The proposed approach has been evaluated and tested for annotation of objects in images. Qianni Zhang, Ebroul Izquierdo |
ICASSP (2) | 1 |