Jing Zhang 0051

dblp:05/3499-51 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
12since 2021 · last 2024
0000-0003-2663-053XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021
YearPublicationVenuePosition
2024 An automatic detection method for schizophrenia based on abnormal eye movements in reading tasks
Ling He 0003, Xiujuan Zheng, Jing Zhang 0051
Expert Syst. Appl.7
2023 CLC-Net: Contextual and local collaborative network for lesion segmentation in diabetic retinopathy images
Yuqi Fang, Sen Yang 0006, Delong Zhu 0001, Jing Zhang 0051, Jun Zhang 0018, Jun Cheng 0003, Raymond Kai-Yu Tong, Xiao Han 0011
Neurocomputing6
2023 RetCCL: Clustering-guided contrastive learning for whole-slide image retrieval
Yuexi Du, Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
Medical Image Anal.6
2023 A generalizable and robust deep learning algorithm for mitosis detection in multicenter breast histopathological images
Jun Zhang 0018, Sen Yang 0006, Jingxi Xiang, Feng Luo 0003, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
Medical Image Anal.7
2023 WNSA-Net: An Axial-Attention-Based Network for Schizophrenia Detection Using Wideband and Narrowband Spectrograms
abstract
Schizophrenia is a severe mental disease that affects patients' thoughts, feelings, and behaviors. Speech signal has proven to be a biomarker in the early diagnosis of schizophrenia. Previous studies on schizophrenic speech detection are mainly based on manual feature extraction engineering, which requires domain knowledge for researchers and has difficulties extracting effective features. This work proposes an end-to-end architecture, called Axial-attention-based Network using Wideband and Narrowband Spectrograms (WNSA-Net), to detect schizophrenia. Specifically, we adopt both wideband and narrowband spectrograms as inputs to represent speech signals using fine time and frequency structures. Then dilated convolution blocks are employed to capture detailed and long-range information in spectrograms. Axial-attention blocks are introduced to augment the information in feature maps along the time and frequency axes. In addition, we employ a gate mechanism to fuse the output feature maps from all channels. Experimental results on the Schizophrenia dataset and its subdatasets show that schizophrenic patients have difficulties in expressing emotions. To validate the performance of our WNSA-Net, experiments are conducted on Schizophrenia dataset and open-access TORGO database, achieving 97.37% and 98.16% accuracy in detecting schizophrenia and dysarthria, respectively. The results show promise for the proposed method in the diagnosis of disordered speech.
Ling He 0003, Jing Zhang 0051
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 SCL-WC: Cross-Slide Contrastive Learning for Weakly-Supervised Whole-Slide Image Classification
abstract
Weakly-supervised whole-slide image (WSI) classification (WSWC) is a challenging task where a large number of unlabeled patches (instances) exist within each WSI (bag) while only a slide label is given. Despite recent progress for the multiple instance learning (MIL)-based WSI analysis, the major limitation is that it usually focuses on the easy-to-distinguish diagnosis-positive regions while ignoring positives that occupy a small ratio in the entire WSI. To obtain more discriminative features, we propose a novel weakly-supervised classification method based on cross-slide contrastive learning (called SCL-WC), which depends on task-agnostic self-supervised feature pre-extraction and task-specific weakly-supervised feature refinement and aggregation for WSI-level prediction. To enable both intra-WSI and inter-WSI information interaction, we propose a positive-negative-aware module (PNM) and a weakly-supervised cross-slide contrastive learning (WSCL) module, respectively. The WSCL aims to pull WSIs with the same disease types closer and push different WSIs away. The PNM aims to facilitate the separation of tumor-like patches and normal ones within each WSI. Extensive experiments demonstrate state-of-the-art performance of our method in three different classification tasks (e.g., over 2% of AUC in Camelyon16, 5% of F1 score in BRACS, and 3% of AUC in DiagSet). Our method also shows superior flexibility and scalability in weakly-supervised localization and semi-supervised classification experiments (e.g., first place in the BRIGHT challenge). Our code will be available at https://github.com/Xiyue-Wang/SCL-WC.
Jinxi Xiang, Jun Zhang 0018, Sen Yang 0006, Zhongyi Yang, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
NeurIPS7
2022 Transformer-based unsupervised contrastive learning for histopathological image classification
Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
Medical Image Anal.5
2022 A Robust Movement Quantification Algorithm of Hyperactivity Detection for ADHD Children Based on 3D Depth Images
abstract
Attention deficit hyperactivity disorder (ADHD) is one of the most common childhood mental disorders. Hyperactivity is a typical symptom of ADHD in children. Clinicians diagnose this symptom by evaluating the children's activities based on subjective rating scales and clinical experience. In this work, an objective system is proposed to quantify the movements of children with ADHD automatically. This system presents a new movement detection and quantification method based on depth images. A novel salient object extraction method is proposed to segment body regions. In movement detection, we explore a new local search algorithm to detect any potential motions of children based on three newly designed evaluation metrics. In the movement quantification, two parameters are investigated to quantify the participation degree and the displacements of each body part in the movements. This system is tested by a depth dataset of children with ADHD. The movement detection results of this dataset mainly range from 91.0% to 95.0%. The movement quantification results of children are consistent with the clinical observations. The public MSR Action 3D dataset is tested to validate the performance of this system.
Ling He 0003, Jing Zhang 0051
IEEE Trans. Image Process.5
2021 TransPath: Transformer-Based Self-supervised Learning for Histopathological Image Classification
Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Junzhou Huang, Wei Yang 0032, Xiao Han 0011
MICCAI (8)5
2021 A hybrid network for automatic hepatocellular carcinoma segmentation in H&E-stained whole slide images
Yuqi Fang, Sen Yang 0006, Delong Zhu 0001, Jing Zhang 0051, Raymond Kai-Yu Tong, Xiao Han 0011
Medical Image Anal.6
2021 Automatic Detection of Affective Flattening in Schizophrenia: Acoustic Correlates to Sound Waves and Auditory Perception
abstract
Affective flattening is a typical negative symptom in schizophrenia that causes a diminution of normal behaviors and functions in schizophrenic patients. In this work, an automatic algorithm of schizophrenia detection is proposed. This algorithm includes three newly proposed features. These features establish a representation for the production and perception of schizophrenic speech, called the K-Sf (kurtosis skewness fraction), SPI (sound perception indicator), and EBMC (enhanced bilateral matching coefficient). The K-Sf is proposed to reflect the overall distribution of speech segments. The SPI evaluates the relation between the speech composition and the sound perception. The EBMC feature is proposed based on the theory of air pressure oscillations which could reflect the details of air modulation in the vocal tract. Experiments evaluating the discriminative capabilities of the three features are conducted using a speech dataset that is collected from 56 participants (28 schizophrenic patients and 28 healthy controls) and an ensemble classifier. Comparative experiments with SVM classifiers are also conducted. The discrimination accuracies of patients and control subjects using the K-Sf, SPI, EBMC, and the ensemble classifier are in the range of 76.8-92.9%, 82.1-92.8%, and 80.4-91.1%, respectively. When the three features are combined, the discrimination results range from 89.3% to 94.6%. The experimental results indicate that the three features have stronger robustness and better discrimination capability than those previous features relating to the detection of flat affect.
Ling He 0003, Jing Zhang 0051
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Trajectory and image-based detection and identification of UAV
Luchuan Liao, Harry Qin, Ling He 0003, Han Zhang 0034, Jing Zhang 0051
Vis. Comput.8
2020 Unsupervised Learning for CT Image Segmentation via Adversarial Redrawing
Youyi Song, Teng Zhou, Jeremy Yuen-Chun Teoh, Jing Zhang 0051, Harry Qin
MICCAI (4)4
2014 Automatic Evaluation of Hypernasality and Consonant Misarticulation in Cleft Palate Speech
abstract
The automatic evaluation of CP (Cleft Plate) speech has various clinical applications. In this work, automatic classification of hypernasality levels and detection of consonant omission methods are proposed. Considering that the data collection is a major bottleneck in the field of CP speech signal processing, an extensive CP speech database is used. This database is collected over 10 years by the Hospital of Stomatology, Sichuan University, which has the largest number of CLP (Cleft Lip and Palate) patients in China. The vocabulary used for this database covers all the initial consonants and the most widely used vowels in Mandarin. Based on the production process of hypernasality, a two-step classification algorithm is proposed to identify four levels of hypernasality: normal, mild, moderate and severe. In order to locate consonant omission, an effective classification algorithm is proposed through analyzing the acoustic characteristics of initial consonants and finals. The classification accuracy for four hypernasality levels reaches up to 83% and the correct identification of consonant omission is over 94%.
Ling He 0003, Jing Zhang 0051, Qi Liu 0046, Margaret Lech
IEEE Signal Process. Lett.2