Zengchang Qin

dblp:05/1860 · DBLP profile ↗
← Back
80ranked-venue papers
14as first author
16since 2021 · last 2025
0000-0002-8084-6721ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 11 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Radgen: A Cross-Modal Fusion System for Automated Radiology Report Generation
abstract
Automating radiology report generation can significantly reduce the workload of radiologists while improving the accuracy and consistency of clinical documentation. However, achieving optimal alignment between visual and textual representations in medical imaging remains a challenge. To address this, we demonstrate RadGen, a cross-modal fusion based system for automated medical report generation. RadGen uses MedCLIP as both a vision extractor and a retrieval mechanism to enhance the integration of imaging and textual data. By extracting features from retrieved reports and medical images through an attentionbased extraction module and integrating them with a fusion module, our system improves the coherence, accuracy, and clinical relevance of generated reports.
Qianhao Han, Daniel Ding, Zengchang Qin, Zheng Zheng 0005
CBMS4
2025 Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation
abstract
Integrating information from various reference databases is a major challenge for Retrieval-Augmented Generation (RAG) systems because each knowledge source adopts a unique data structure and follows different conventions. Retrieving from multiple knowledge sources with one fixed strategy usually leads to under-exploitation of information. To mitigate this drawback, inspired by Mix-of-Expert, we introduce Mix-of-Granularity (MoG), a method that dynamically determines the optimal granularity of a knowledge source based on input queries using a router. The router is efficiently trained with a newly proposed loss function employing soft labels. We further extend MoG to MoG-Graph (MoGG), where reference documents are pre-processed as graphs, enabling the retrieval of distantly situated snippets. Experiments demonstrate that MoG and MoGG effectively predict optimal granularity levels, significantly enhancing the performance of the RAG system in downstream tasks. The code of both MoG and MoGG will be made public.
Zijie Zhong, Xiaoya Cui, Xiaofan Zhang 0012, Zengchang Qin
COLING5
2025 SyntheT2C: Generating Synthetic Data for Fine-Tuning Large Language Models on the Text2Cypher Task
abstract
Integrating Large Language Models (LLMs) with existing Knowledge Graph (KG) databases presents a promising avenue for enhancing LLMs’ efficacy and mitigating their “hallucinations”. Given that most KGs reside in graph databases accessible solely through specialized query languages (e.g., Cypher), it is critical to connect LLMs with KG databases by automating the translation of natural language into Cypher queries (termed as “Text2Cypher” task). Prior efforts tried to bolster LLMs’ proficiency in Cypher generation through Supervised Fine-Tuning (SFT). However, these explorations are hindered by the lack of annotated datasets of Query-Cypher pairs, resulting from the labor-intensive and domain-specific nature of such annotation. In this study, we propose SyntheT2C, a methodology for constructing a synthetic Query-Cypher pair dataset, comprising two distinct pipelines: (1) LLM-based prompting and (2) template-filling. SyntheT2C is applied to two medical KG databases, culminating in the creation of a synthetic dataset, MedT2C. Comprehensive experiments demonstrate that the MedT2C dataset effectively enhances the performance of backbone LLMs on Text2Cypher task via SFT. Both the SyntheT2C codebase and the MedT2C dataset will be released.
Zijie Zhong, Linqing Zhong, Zhaoze Sun, Qingyun Jin, Zengchang Qin, Xiaofan Zhang 0012
COLING5
2025 GraIS: Graph-Based Interpretable Network for Sepsis Mortality Prediction Using Multi-modal Electronic Health Records
Zengchang Qin
ICONIP (3)5
2025 Contrastive Instruction Fine-Tuning Large Multimodal Model for Hateful Meme Classification
abstract
Detecting hateful memes requires a model that possesses extensive background knowledge and robust reasoning abilities, especially when the memes contain ambiguous descriptions. Previous research has used large language models (LLMs) and large multimodal models (LMMs) to interpret and categorize these memes. However, distinguishing subtly different hateful and non-hateful memes is still challenging. In recognition of this, our study introduces a unique contrastive instruction fine-tuning approach, InstructMemeCL. This method improves an LMM's ability to discern between memes that have similar visual or textual elements by intensifying its focus on semantic subtleties that separate hateful from non-hateful content. We evaluated our model using AUROC and accuracy metrics on three publicly available hateful meme datasets. The results indicate that our improved LMM more accurately identifies hateful and non-hateful memes, demonstrating superior performance compared to conventional LLMs and LMMs used in similar tasks.
Ming Shan Hee, Xiangxiang Chu, Roy Ka-Wei Lee, Zengchang Qin
ICWSM6
2025 MCA: Multimodal Contrastive Augmentation for Medical Report Generation
Zengchang Qin
PAKDD (7)3
2025 CADReN: Contextual Anchor-Driven Relational Network for Controllable Cross-Graphs Node Importance Estimation
Zijie Zhong, Yunhui Zhang, Ziyi Chang, Zengchang Qin
PAKDD (1)4
2024 Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation
abstract
Automating radiology report generation can significantly reduce the workload of radiologists and enhance the accuracy, consistency, and efficiency of clinical documentation. We propose a novel cross-modal framework that uses MedCLIP as both a vision extractor and a retrieval mechanism to improve the process of medical report generation. By extracting retrieved report features and image features through an attention-based extract module, and integrating them with a fusion module, our method improves the coherence and clinical relevance of generated reports. Experimental results on the widely used IU-Xray dataset demonstrate the effectiveness of our approach, showing improvements over commonly used methods in both report quality and relevance. Additionally, ablation studies provide further validation of the framework, highlighting the importance of accurate report retrieval and feature integration in generating comprehensive medical reports.
Qianhao Han, Zengchang Qin, Zheng Zheng 0005
IEEE Big Data3
2024 Robust Lightweight Depth Estimation Model via Data-Free Distillation
abstract
Existing Monocular Depth Estimation (MDE) methods often use large and complex neural networks. Despite the advanced performance of these methods, we consider the efficiency and generalization for practical applications with limited resources. In our paper, we present an efficient transformer-based monocular relative depth estimation network and train it with a diverse depth dataset to obtain good generalization performance. Knowledge distillation (KD) is employed to transfer the general knowledge from a pre-trained teacher network to the compact student network, demonstrating that KD can improve the generalization ability as well as the accuracy. Moreover, we propose a geometric label-free distillation method to improve the lightweight model in specific domains utilizing 3D geometric cues with unlabeled data. We show that our method outperforms other KD methods with or without ground truth supervision. Finally, we propose an application of the lightweight network to a two-stage depth completion task. Our method shows on par or even superior cross-domain generalization ability compared to large networks.
Wei Yin 0006, Yifan Liu 0001, Zengchang Qin
ICASSP5
2024 Caseg: Clip-Based Action Segmentation With Learnable Text Prompt
abstract
Video action segmentation aims to identify and localize actions. Existing models have achieved impressive performance with pre-extracted frame-level features, but this may limit zero-shot learning and cross-dataset inference, especially for new actions or scenes. To overcome this problem, we propose a novel end-to-end network designed for robust performance across both familiar and novel action segmentation scenarios. Our approach combines a plug-and-play visual prompt module enhancing CLIP features’ temporal understanding, and a learnable text prompt that enriches label semantics and refines the model’s focus, significantly boosting performance. Our results demonstrate that CLIP features can assist in action segmentation tasks, and prompts can improve task effectiveness. Furthermore, our findings show that CLIP features contain information that i3d features do not. We evaluate the proposed method on several video datasets, including Georgia Tech Egocentric Activities (GTEA), 50Salads, and Breakfast, and the results show that the proposed model outperforms existing SOTA models.
Suyuan Huang 0001, Yan Gao 0017, Yao Hu 0002, Zengchang Qin
ICIP6
2024 S3GCN: Sport Scoring Siamese Graph Convolution Network
abstract
Temporal sequences of human body key points provide detailed motion information, serving as a crucial foundation for human action analysis. Existing public methods and datasets predominantly focus on action category estimation, lacking a comprehensive evaluation of sport scoring. In this work, we propose a novel model of sport scoring called Sport Scoring Siamese Graph Convolution Network S3GCN)1, which surpasses the constraints inherent in prior methods by implicitly capturing nuanced differences between teacher pose and student pose. In a Few-shot dataset, Taichi, it achieves a benchmark level of performance through spacial and temporal augmentation with comprehensive ablation experiments. Furthermore, our approach outperforms the original model on classification, including NTU-RGB-D and Taichi classification datasets.1https://github.com/divided7/SSSGCN
Zhuming Zhang, Shiming Lin, Dengpan Zhang, Haibin Ma, Zengchang Qin
ICIP6
2023 Boosting Semantic Segmentation from the Perspective of Explicit Class Embeddings
abstract
Semantic segmentation is a computer vision task that associates a label with each pixel in an image. Modern approaches tend to introduce class embeddings into semantic segmentation for deeply utilizing category semantics, and regard supervised class masks as final predictions. In this paper, we explore the mechanism of class embeddings and have an insight that more explicit and meaningful class embeddings can be generated based on class masks purposely. Following this observation, we propose ECENet, a new segmentation paradigm, in which class embeddings are obtained and enhanced explicitly during interacting with multi-stage image features. Based on this, we revisit the traditional decoding process and explore inverted information flow between segmentation masks and class embeddings. Furthermore, to ensure the discriminability and informativity of features from backbone, we propose a Feature Reconstruction module, which combines intrinsic and diverse branches together to ensure the concurrence of diversity and redundancy in features. Experiments show that our ECENet outperforms its counterparts on the ADE20K dataset with much less computational cost and achieves new state-of-the-art results on PASCALContext dataset. The code will be released at https://gitee.com/mindspore/models and https://github.com/Carol-lyh/ECENet.
Yuhe Liu, Chuanjian Liu, Kai Han 0002, Quan Tang 0001, Zengchang Qin
ICCV5
2022 Sparse Double Descent: Where Network Pruning Aggravates Overfitting
abstract
People usually believe that network pruning not only reduces the computational cost of deep networks, but also prevents overfitting by decreasing model capacity. However, our work surprisingly discovers that network pruning sometimes even aggravates overfitting. We report an unexpected sparse double descent phenomenon that, as we increase model sparsity via network pruning, test performance first gets worse (due to overfitting), then gets better (due to relieved overfitting), and gets worse at last (due to forgetting useful information). While recent studies focused on the deep double descent with respect to model overparameterization, they failed to recognize that sparsity may also cause double descent. In this paper, we have three main contributions. First, we report the novel sparse double descent phenomenon through extensive experiments. Second, for this phenomenon, we propose a novel learning distance interpretation that the curve of l2 learning distance of sparse models (from initialized parameters to final parameters) may correlate with the sparse double descent curve well and reflect generalization better than minima flatness. Third, in the context of sparse double descent, a winning ticket in the lottery ticket hypothesis surprisingly may not always win.
Zeke Xie, Quanzhi Zhu, Zengchang Qin
ICML4
2022 A deep learning method for automatic evaluation of diagnostic information from multi-stained histopathological images
Junyu Ji, Tao Wan 0001, Hao Wang 0131, Menghan Zheng, Zengchang Qin
Knowl. Based Syst.6
2021 Random Neural Graph Generation with Structure Evolution
Yuguang Zhou, Tao Wan 0001, Zengchang Qin
ICONIP (2)4
2021 Learning Dual Encoding Model for Adaptive Visual Understanding in Visual Dialogue
abstract
Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue task involves multiple rounds of dialogues which cover a broad range of visual content that could be related to any objects, relationships or high-level semantics. Thus one of the key challenges in Visual Dialogue task is to learn a more comprehensive and semantic-rich image representation that can adaptively attend to the visual content referred by variant questions. In this paper, we first propose a novel scheme to depict an image from both visual and semantic views. Specifically, the visual view aims to capture the appearance-level information in an image, including objects and their visual relationships, while the semantic view enables the agent to understand high-level visual semantics from the whole image to the local regions. Furthermore, on top of such dual-view image representations, we propose a Dual Encoding Visual Dialogue (DualVD) module, which is able to adaptively select question-relevant information from the visual and semantic views in a hierarchical mode. To demonstrate the effectiveness of DualVD, we propose two novel visual dialogue models by applying it to the Late Fusion framework and Memory Network framework. The proposed models achieve state-of-the-art results on three benchmark datasets. A critical advantage of the DualVD module lies in its interpretability. We can analyze which modality (visual or semantic) has more contribution in answering the current question by explicitly visualizing the gate values. It gives us insights in understanding of information selection mode in the Visual Dialogue task. The code is available at https://github.com/JXZe/Learning_DualVD.
Jing Yu 0007, Xiaoze Jiang, Zengchang Qin, Weifeng Zhang 0002, Yue Hu 0002, Qi Wu 0001
IEEE Trans. Image Process.3
2020 DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
abstract
Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of visual content that could be related to any objects, relationships or semantics. The key challenge in Visual Dialogue task is thus to learn a more comprehensive and semantic-rich image representation which may have adaptive attentions on the image for variant questions. In this research, we propose a novel model to depict an image from both visual and semantic perspectives. Specifically, the visual view helps capture the appearance-level information, including objects and their relationships, while the semantic view enables the agent to understand high-level visual semantics from the whole image to the local regions. Futhermore, on top of such multi-view image features, we propose a feature selection framework which is able to adaptively capture question-relevant information hierarchically in fine-grained level. The proposed method achieved state-of-the-art results on benchmark Visual Dialogue datasets. More importantly, we can tell which modality (visual or semantic) has more contribution in answering the current question by visualizing the gate values. It gives us insights in understanding of human cognition in Visual Dialogue.
Xiaoze Jiang, Jing Yu 0007, Zengchang Qin, Yingying Zhuang, Yue Hu 0002, Qi Wu 0001
AAAI3
2020 Prior Visual Relationship Reasoning For Visual Question Answering
abstract
Visual Question Answering (VQA) is a representative task of cross-modal reasoning where an image and a free-form question in natural language are presented and the correct answer needs to be determined using both visual and textual information. One of the key issues of VQA is to reason with semantic clues in the visual content under the guidance of the question. In this paper, we propose Scene Graph Convolutional Network (SceneGCN) to jointly reason the object properties and their semantic relations for the correct answer. The visual relationship is projected into a deep learned semantic space constrained by visual context and language priors. Based on comprehensive experiments on two challenging datasets: GQA and VQA 2.0, we demonstrate the effectiveness and interpretability of the new model.
Zhuoqian Yang, Zengchang Qin, Jing Yu 0007, Tao Wan 0001
ICIP2
2020 A Lightweight Network Model For Video Frame Interpolation Using Spatial Pyramids
abstract
In recent years, deep learning based video frame interpolation methods have shown impressive results in handling occlusion, blur and large motion. However, they are usually very heavy in terms of model size, and they hardly to be employed in i.e. mobile phones or other portable devices with limited computing power. To address the problem, we propose light-weighted Spatial Pyramid Frame Interpolation Network (SPFIN), a hierarchical network in a coarse-to-fine approach to reconstruct frames. At each pyramid level, we apply two light sub-networks to model optical flow and visibility mask instead of commonly used U-Net architecture. The flow and mask are up-sampled and optimized progressively. Finally, the intermediate frame is formed by linearly blending warped frames and masks. Experimental results on two benchmark problems show that our model has the smallest size, but better or comparable performance comparing to existing state-of-the art models.
Jiankai Zhuang, Zengchang Qin, Tao Wan 0001
ICIP2
2020 A Deep Learning Model for Early Prediction of Sepsis from Intensive Care Unit Records
Rui Zhao 0019, Tao Wan 0001, Zhengbo Zhang, Zengchang Qin
ICONIP (4)5
2020 Many-to-One Stable Matching for Prediction in Social Networks
Zengchang Qin, Tao Wan 0001
IEA/AIE2
2020 DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue
abstract
Visual Dialogue task requires an agent to be engaged in a conversation with human about an image. The ability of generating detailed and non-repetitive responses is crucial for the agent to achieve human-like conversation. In this paper, we propose a novel generative decoding architecture to generate high-quality responses, which moves away from decoding the whole encoded semantics towards the design that advocates both transparency and flexibility. In this architecture, word generation is decomposed into a series of attention-based information selection steps, performed by the novel recurrent Deliberation, Abandon and Memory (DAM) module. Each DAM module performs an adaptive combination of the response-level semantics captured from the encoder and the word-level semantics specifically selected for generating each word. Therefore, the responses contain more detailed and non-repetitive descriptions while maintaining the semantic accuracy. Furthermore, DAM is flexible to cooperate with existing visual dialogue encoders and adaptive to the encoder structures by constraining the information selection mode in DAM. We apply DAM to three typical encoders and verify the performance on the VisDial v1.0 dataset. Experimental results show that the proposed models achieve new state-of-the-art performance with high-quality responses. The code is available at https://github.com/JXZe/DAM.
Xiaoze Jiang, Jing Yu 0007, Yajing Sun, Zengchang Qin, Yue Hu 0002, Qi Wu 0001
IJCAI4
2020 KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual Dialogue
abstract
Visual dialogue is a challenging task that needs to extract implicit information from both visual (image) and textual (dialogue history) contexts. Classical approaches pay more attention to the integration of the current question, vision knowledge and text knowledge, despising the heterogeneous semantic gaps between the cross-modal information. In the meantime, the concatenation operation has become de-facto standard to the cross-modal information fusion, which has a limited ability in information retrieval. In this paper, we propose a novel Knowledge-Bridge Graph Network (KBGN) model by using graph to bridge the cross-modal semantic relations between vision and text knowledge in fine granularity, as well as retrieving required knowledge via an adaptive information selection mode. Moreover, the reasoning clues for visual dialogue can be clearly drawn from intra-modal entities and inter-modal bridges. Experimental results on VisDial v1.0 and VisDial-Q datasets demonstrate that our model outperforms existing models with state-of-the-art results.
Xiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun, Jing Yu 0007
ACM Multimedia3
2020 Robust nuclei segmentation in histopathology using ASPPU-Net and boundary refinement
Tao Wan 0001, Hongxiang Feng, Chao Tong 0001, Zengchang Qin
Neurocomputing6
2020 Cross-modal learning with prior visual relation knowledge
Jing Yu 0007, Weifeng Zhang 0002, Zhuoqian Yang, Zengchang Qin, Yue Hu 0002
Knowl. Based Syst.4
2020 Learning cross-modal correlations by exploring inter-word semantics and stacked co-attention
Jing Yu 0007, Weifeng Zhang 0002, Zengchang Qin, Yanbing Liu 0007, Yue Hu 0002
Pattern Recognit. Lett.4
2020 Reasoning on the Relation: Enhancing Visual Representation for Visual Question Answering and Cross-Modal Retrieval
abstract
Cross-modal analysis has become a promising direction for artificial intelligence. Visual representation is crucial for various cross-modal analysis tasks that require visual content understanding. Visual features which contain semantical information can disentangle the underlying correlation between different modalities, thus benefiting the downstream tasks. In this paper, we propose a Visual Reasoning and Attention Network (VRANet) as a plug-and-play module to capture rich visual semantics and help to enhance the visual representation for improving cross-modal analysis. Our proposed VRANet is built based on the bilinear visual attention module which identifies the critical objects. We propose a novel Visual Relational Reasoning (VRR) module to reason about pair-wise and inner-group visual relationships among objects guided by the textual information. The two modules enhance the visual features at both relation level and object level. We demonstrate the effectiveness of the proposed VRANet by applying it to both Visual Question Answering (VQA) and Cross-Modal Information Retrieval (CMIR) tasks. Extensive experiments conducted on VQA 2.0, CLEVR, CMPlaces, and MS-COCO datasets indicate superior performance comparing with state-of-the-art work.
Jing Yu 0007, Weifeng Zhang 0002, Zengchang Qin, Yue Hu 0002, Jianlong Tan, Qi Wu 0001
IEEE Trans. Multim.4
2019 Stock Volatility Prediction Based on Self-attention Networks with Social Information
abstract
Stock volatility prediction is a challenging task in time-series prediction according to the Efficient Market Hypothesis which supposes all the investors are rational. However, many theories have showed that stock markets are not efficient due to the effects of psychological and social factors. In this paper, we constructed self-attention networks (SAN) to quantify the impact on the volatility of Chinese stock market of social information, such as social opinion and social concern. Our SAN model can explore the relationships among features at different time steps more flexibly, and thus, explore stock historical information more effectively. Empirical results show the superiority of our model compared to other existing models on given stock data.
Andi Xia, Tao Wan 0001, Zengchang Qin
CIFEr5
2019 Structured Knowledge Distillation for Semantic Segmentation
abstract
In this paper, we investigate the issue of knowledge distillation for training compact semantic segmentation networks by making use of cumbersome networks. We start from the straightforward scheme, pixel-wise distillation, which applies the distillation scheme originally introduced for image classification and performs knowledge distillation for each pixel separately. We further propose to distill the structured knowledge from cumbersome networks into compact networks, which is motivated by the fact that semantic segmentation is a structured prediction problem. We study two such structured distillation schemes: (i) pair-wise distillation that distills the pairwise similarities, and (ii) holistic distillation that uses adversarial training to distill holistic knowledge. The effectiveness of our knowledge distillation approaches is demonstrated by extensive experiments on three scene parsing datasets: Cityscapes, Camvid and ADE20K.
Yifan Liu 0001, Chris Liu, Zengchang Qin, Zhenbo Luo, Jingdong Wang 0001
CVPR4
2019 Pixel Level Data Augmentation for Semantic Image Segmentation Using Generative Adversarial Networks
abstract
Semantic segmentation is one of the basic topics in computer vision, it aims to assign semantic labels to every pixel of an image. Unbalanced semantic label distribution could have a negative influence on segmentation accuracy. In this paper, we investigate using data augmentation approach to balance the semantic label distribution in order to improve segmentation performance. We propose using generative adversarial networks (GANs) to generate realistic images for improving the performance of semantic segmentation networks. Experimental results show that the proposed method can not only improve segmentation performance on those classes with low accuracy, but also obtain 1.3% to 2.1% increase in average segmentation accuracy. It shows that this augmentation method can boost the accuracy and be easily applicable to any other segmentation models.
Shuangting Liu, Yifan Liu 0001, Zengchang Qin, Tao Wan 0001
ICASSP5
2019 A Sequential Guiding Network with Attention for Image Captioning
abstract
The recent advances of deep learning in both computer vision (CV) and natural language processing (NLP) provide us a new way of understanding semantics, by which we can deal with more challenging tasks such as automatic description generation from natural images. In this challenge, the encoder-decoder framework has achieved promising performance when a convolutional neural network (CNN) is used as image encoder and a recurrent neural network (RNN) as decoder. In this paper, we introduce a sequential guiding network that guides the decoder during word generation. The new model is an extension of the encoder-decoder framework with attention that has an additional guiding long short-term memory (LSTM) and can be trained in an end-to-end manner by using image/descriptions pairs. We validate our approach by conducting extensive experiments on a benchmark dataset, i.e., MS COCO Captions. The proposed model achieves significant improvement comparing to the other state-of-the-art deep learning models.
Daouda Sow, Zengchang Qin, Mouhamed Niasse, Tao Wan 0001
ICASSP2
2019 Multi-Level Network for High-Speed Multi-Person Pose Estimation
abstract
In multi-person pose estimation, the left/right joint type discrimination is always a hard problem because of the similar appearance. Traditionally, we solve this problem by stacking multiple refinement modules to increase network's receptive fields and capture more global context, which can also increase a great amount of computation. In this paper, we propose a Multi-level Network (MLN) that learns to aggregate features from lower-level (left/right information), upper-level (localization information), joint-limb level (complementary information) and global-level (context) information for discrimination of joint type. Through feature reuse and its intra-relation, MLN can attain comparable performance to other conventional methods while runtime speed retains at 42 FPS.
Ying Huang 0003, Jiankai Zhuang, Zengchang Qin
ICIP3
2019 Semantic Modeling of Textual Relationships in Cross-modal Retrieval
Jing Yu 0007, Zengchang Qin, Zhuoqian Yang, Yue Hu 0002
KSEM (1)3
2019 FollowMeUp Sports: New Benchmark for 2D Human Keypoint Recognition
Ying Huang 0003, Haipeng Kan, Jiankai Zhuang, Zengchang Qin
PRCV (3)5
2019 Accurate segmentation of overlapping cells in cervical cytology with deep convolutional neural networks
Tao Wan 0001, Shusong Xu, Chen Sang, Yulan Jin, Zengchang Qin
Neurocomputing5
2018 Improved Nuclear Segmentation on Histopathology Images Using a Combination of Deep Learning and Active Contour Model
Tao Wan 0001, Hongxiang Feng, Zengchang Qin
ICONIP (6)4
2018 Text Generation Based on Generative Adversarial Nets with Latent Variables
Zengchang Qin, Tao Wan 0001
PAKDD (2)2
2018 Emotion Classification with Data Augmentation Using Generative Adversarial Networks
Xinyue Zhu, Yifan Liu 0001, Tao Wan 0001, Zengchang Qin
PAKDD (3)5
2018 Auto-painter: Cartoon image generation from sketch by using conditional Wasserstein generative adversarial networks
Yifan Liu 0001, Zengchang Qin, Tao Wan 0001, Zhenbo Luo
Neurocomputing2
2017 Kinetic measures for distinguishing vulnerable from stable atherosclerotic plaque with dynamic contrast-enhanced MRI
abstract
Carotid atherosclerosis is a primary cause of stroke, which is responsible for a majority of disabilities and deaths worldwide. Plaque inflammation and abundant microvasculature have been identified as important aspects contributing to plaque vulnerability that can be studied non-invasively with dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI). Due to the asymptomatic nature of vulnerable plaque, there is an unmet clinical need to identify and characterize these lesions before they rupture. We presented an automated computerized method based on kinetic measures to distinguish vulnerable from stable atherosclerotic plaques on DCE-MRI. Four classes of kinectic features, including pharmacokinetic, intensity kinetic, histogram kinetic, and textual kinetic features, were extracted for capturing the pathophysiologic changes in various aspects of plaque vascular structure and functionality in atherosclerosis. These features can reflect the local inflammatory processes and microvasculature changes appearing in plaque destabilization. Our method was evaluated on real clinical data and achieved the area under the curve of 0.95 using a combined feature set, suggesting a potential of this method applied to a computer-aided diagnosis system for an early detection of vulnerable plaques.
Zengchang Qin, Wanshu Zhang, Tao Wan 0001
ICIP1
2017 Motif Iteration Model for Network Representation
Lintao Lv, Zengchang Qin, Tao Wan 0001
ICONIP (5)2
2017 A Radiomics Approach for Automated Identification of Aggressive Tumors on Combined PET and Multi-parametric MRI
Tao Wan 0001, Bixiao Cui, Zengchang Qin, Jie Lu 0010
ICONIP (6)4
2017 Stock Volatility Prediction Using Recurrent Neural Networks with Sentiment Analysis
Yifan Liu 0001, Zengchang Qin, Tao Wan 0001
IEA/AIE (1)2
2017 A Bayesian Model of Game Decomposition
Zengchang Qin, Tao Wan 0001
IEA/AIE (1)2
2017 Automated grading of breast cancer histopathology using cascaded ensemble with combination of multi-level image features
Tao Wan 0001, Jiajia Cao, Zengchang Qin
Neurocomputing4
2017 Automated mitosis detection in histopathology based on non-gaussian modeling of complex wavelet coefficients
Tao Wan 0001, Wanshu Zhang, Alin Achim, Zengchang Qin
Neurocomputing6
2016 Stable Matching in Structured Networks
Ying Ling, Tao Wan 0001, Zengchang Qin
PKAW3
2016 Learning Sentimental Weights of Mixed-gram Terms for Classification and Visualization
Tszhang Guo, Tao Wan 0001, Zengchang Qin
PRICAI5
2016 Topic modeling of Chinese language beyond a bag-of-words
Zengchang Qin, Yonghui Cong, Tao Wan 0001
Comput. Speech Lang.1
2016 Collective game behavior learning with probabilistic graphical models
Zengchang Qin, Farhan Khawar, Tao Wan 0001
Neurocomputing1
2016 Topic correlation model for cross-modal multimedia information retrieval
Zengchang Qin, Jing Yu 0007, Yonghui Cong, Tao Wan 0001
Pattern Anal. Appl.1
2014 A confidence growing model for super-resolution
abstract
Single image super-resolution (SR) aims at generating a high-resolution (HR) image from one low-resolution (LR) input. In this paper, we focus on single image SR by using a confidence growing model based on an example-based super resolution approach. Compared to previous works that reconstruct high-resolution image in a raster scan order, the new proposed method reconstructs the patches using a new confidence measure. More confident reconstructions are propagated to neighboring areas by enforcing a smoothness constraint in selecting patches. We also adopt hierarchical clustering to construct a training set to speed up processing. Experimental results demonstrate that this simple method outperforms existing state-of-the-art algorithms on a the given benchmark SR test images.
Sina Lin, Zengchang Qin, Renjie Liao 0001, Tao Wan 0001
ICIP2
2014 Wavelet-based statistical features for distinguishing mitotic and non-mitotic cells in breast cancer histopathology
abstract
To diagnose breast cancer (BCa), the number of mitotic cells present in tissue sections is an important parameter to examine and grade breast biopsy specimen. The differentiation of mitotic from non-mitotic cells in breast histopathological images is a crucial step for automatical mitosis detection. This work aims at improving the accuracy of mitosis classification by characterizing objects of interest (tissue cells) in wavelet based multi-resolution representations that better capture the statistical features having mitosis discrimination. A dual-tree complex wavelet transform (DT-CWT) is performed to decompose the image patches into multi-scale forms. Five commonly-used statistical features are extracted on each wavelet subband. Since both mitotic and non-mitotic cells appear as small objects with a large variety of shapes in the images, characterization of mitosis is a challenging problem. The inter-scale dependencies of wavelet coefficients allow extraction of important texture features within the cells that are more likely to appear at all different scales. The wavelet-based statistical features were evaluated on a dataset containing 327 mitotic and 406 non-mitotic cells via a support vector machine classifier in iterative cross-validation. The quantitative results showed that our DT-CWT based approach achieved superior classification performance with the accuracy of 87.94%, sensitivity of 86.80%, specificity of 89.89%, and the area under the curve (AUC) value of 0.94.
Tao Wan 0001, Zengchang Qin
ICIP4
2014 A Graphical Model for Collective Behavior Learning Using Minority Games
Farhan Khawar, Zengchang Qin
PAKDD (2)2
2014 Nonparametric bayesian upstream supervised multi-modal topic models
abstract
Learning with multi-modal data is at the core of many multimedia applications, such as cross-modal retrieval and image annotation. In this paper, we present a nonparametric Bayesian approach to learning upstream supervised topic models for analyzing multi-modal data. Our model develops a compound nonparametric Bayesian multi-modal prior to describe the correlation structure of data both within each individual modality and between different modalities. It extends the hierarchical Dirichlet process (HDP) through incorporating upstream supervised response variables and values of latent functions under Gaussian process (GP). Upstream responses shared by data from multiple modalities are beneficial for discriminatively training and GP allows flexible structure learning of correlations. Hence, our model inherits the automatic determination of the number of topics from HDP, structure learning from GP and enhanced predictive capacity from upstream supervision. We also provide efficient variational inference and prediction algorithms. Empirical studies demonstrate superior performances on several benchmark datasets compared with previous competitors.
Renjie Liao 0001, Jun Zhu 0001, Zengchang Qin
WSDM3
2013 A Bag-of-Tones Model with MFCC Features for Musical Genre Classification
Zengchang Qin, Tao Wan 0001
ADMA (1)1
2013 Color saliency model based on mean shift segmentation
abstract
Saliency detection is one of the extraordinary capabilities of the human visual system (HVS). In this paper, we present a novel saliency detection model to capture visual selective attention of images. The new model does not require prior knowledge of salient regions as well as manual labeling. The mean shift segmentation algorithm and quaternion discrete cosine transform (QDCT) are used to generate a rough saliency map by integrating low-level features and spatial saliency information. In each segmented region, the color saliency is measured based on the probability of its occurrences in foreground and background defined by the rough saliency map. The experimental results on a widely used benchmark database demonstrated that the presented model achieves the best performance in terms of visual and quantitative evaluations compared to existing state-of-the-art saliency detection models.
Zengchang Qin, Xiaofan Zhang 0008, Tao Wan 0001
ICASSP2
2013 A robust fusion scheme for multifocus images using sparse features
abstract
Multifocus image fusion is an important research topic in the computer vision and image processing field. The optical lenses that are commonly used by imaging devices, such as auto-focus cameras, have a limiting focus range. Thus, only objects within the range of distances from the devices can be captured and recorded sharply while out-of-range objects become blur. In this paper, we present a novel image fusion scheme for combining two or multiple images with different focus points to generate an all-in-focus image. We formulate the problem of fusing multifocus images as choosing most significant features from a sparse matrix produced by a newly developed robust principal component analysis (RPCA) decomposition method to form a composite feature space. Thus, the salient features presented in sharp regions can be captured and integrated into a single representation. The sparse matrix is first divided into small blocks, and standard deviation is then calculated on each block as a selection criterion. To reduce blocking artifacts, a sliding window technique is utilized to smooth the transitions between blocks. The proposed fusion scheme has been demonstrated to successfully improve fusion quality in terms of visual and quantitative evaluations. The method is also able to effectively handle both grayscale and color images.
Tao Wan 0001, Zengchang Qin, Chenchen Zhu, Renjie Liao 0001
ICASSP2
2013 Salient object detection in image sequences via spatial-temporal cue
abstract
Contemporary video search and categorization are non-trivial tasks due to the massively increasing amount and content variety of videos. We put forward the study of visual saliency models in video. Such a model is employed to identify salient objects from the image background. Starting from the observation that motion information in video often attracts more human attention compared to static images, we devise a region contrast based saliency detection model using spatial-temporal cues (RCST). We introduce and study four saliency principles to realize the RCST. This generalizes the previous static image for saliency computational model to video. We conduct experiments on a publicly available video segmentation database where our method significantly outperforms seven state-of-the-art methods with respect to PR curve, ROC curve and visual comparison.
Chuang Gan 0001, Zengchang Qin, Jia Xu 0004, Tao Wan 0001
VCIP2
2013 What color is an object?
abstract
Color perception is one of the major cognitive abilities of human being. Color information is also one of the most important features in various computer vision tasks including object recognition, tracking, scene classification and so on. In this paper, we proposed a simple and effective method for learning color composition of objects from large annotated datasets. The new proposed model is based on a region-based bag-of-colors model and saliency detection. The effectiveness of the model is empirically verified on manually labelled datasets with single or multiple tags. The significance of this research is that the color information of an object can provide useful prior knowledge to help improving the existing computer vision models in image segmentation, object recognition and tracking.
Xiaofan Zhang 0008, Zengchang Qin, Tao Wan 0001
VCIP2
2013 Hybrid Bayesian estimation tree learning with discrete and fuzzy labels
Zengchang Qin, Tao Wan 0001
Frontiers Comput. Sci.1
2013 Feature integration analysis of bag-of-features model for image retrieval
Jing Yu 0007, Zengchang Qin, Tao Wan 0001
Neurocomputing2
2013 Multifocus image fusion based on robust principal component analysis
Tao Wan 0001, Chenchen Zhu, Zengchang Qin
Pattern Recognit. Lett.3
2012 Image Super-Resolution Using Local Learnable Kernel Regression
Renjie Liao 0001, Zengchang Qin
ACCV (3)2
2012 Cross-Modal Information Retrieval - A Case Study on Chinese Wikipedia
Yonghui Cong, Zengchang Qin, Jing Yu 0007, Tao Wan 0001
ADMA2
2012 Cross-modal topic correlations for multimedia retrieval
Jing Yu 0007, Yonghui Cong, Zengchang Qin, Tao Wan 0001
ICPR3
2012 An Efficient Minimum Vocabulary Construction Algorithm for Language Modeling
Sina Lin, Zengchang Qin, Zehua Huang, Tao Wan 0001
IEA/AIE2
2011 Clustering data and imprecise concepts
abstract
Cluster analysis is the assignment of grouping a set of observations into clusters so that observations in the same cluster are similar in some sense. One of the key features for clustering is how to define a sensible similarity measure. However, classical clustering algorithms have no ability to cluster data instances and imprecise concepts using traditional distance measures. In this paper, we proposed a (dis)similarity measure based on a new knowledge representation framework called label semantics. Based on this new measure, we can automatically cluster data instance and descriptive concepts represented by logical expressions of linguistic labels. Experimental results on a toy problem in image classification demonstrate the effectiveness of the new proposed clustering algorithm. Since the new proposed measure can be extended to measuring distance between any two granularities, the new clustering algorithms can also be extended to clustering data instance and imprecise concepts represented by other granularities.
Weifeng Zhang 0008, Zengchang Qin
FUZZ-IEEE2
2011 Exploring Market Behaviors with Evolutionary Mixed-Games Learning Model
Yingsai Dong, Zengchang Qin, Tao Wan 0001
ICCCI (1)3
2011 Topic Modeling of Chinese Language Using Character-Word Relations
Zengchang Qin, Tao Wan 0001
ICONIP (3)2
2009 Ranking Answers by Hierarchical Topic Models
Zengchang Qin, Marcus Thint, Zhiheng Huang
IEA/AIE1
2008 Question Classification using Head Words and their Hypernyms
Zhiheng Huang, Marcus Thint, Zengchang Qin
EMNLP3
2008 LFOIL: Linguistic rule induction in the label semantics framework
Zengchang Qin, Jonathan Lawry
Fuzzy Sets Syst.1
2007 PNL-Enhanced Restricted Domain Question Answering System
abstract
The concept of PNL (Precisiated Natural Language) has been proposed by Zadeh for computation with perceptions and some problems described in natural language. We describe a design for restricted domain question answering systems enhanced by PNL-based reasoning. For a subset of a knowledge corpus (e.g. critical or frequently-asked topics) where fuzzy set definitions of vague terms are provided, more precise answers can be computed via protoformal deduction. Nested structure in the system design also enables processing of natural language statements that are not PNL protoforms using phrase-based deduction and concept matching to generate the most relevant facts for a query. If deduction results yield low confidence factor, standard search engine provides a baseline response (relevant paragraphs based on keyword matches). Our design principles aim for flexible, domain independent capability and minimize human input to provision of semantic clues and background knowledge during design or application set-up.
Mirza Mohd. Sufyan Beg, Marcus Thint, Zengchang Qin
FUZZ-IEEE3
2007 Fuzziness and Performance: An Empirical Study with Linguistic Decision Trees
Zengchang Qin, Jonathan Lawry
IFSA (1)1
2007 Deduction Engine Design for PNL-Based Question Answering System
Zengchang Qin, Marcus Thint, Mirza Mohd. Sufyan Beg
IFSA (1)1
2006 Naive Bayes Classification Given Probability Estimation Trees
abstract
Tree induction is one of the most effective and widely used models in classification. Unfortunately, decision trees such as C4.5 have been found to provide poor probability estimates. By the empirical studies, Provost and Domingos found that probability estimation trees (PETs) give a fairly good probability estimation. However, different from normal decision trees, pruning reduces the performances of PETs. In order to get a good probability estimation, we usually need large trees which are not good in terms of the model transparency. In this paper, two hybrid models by combining the naive Bayes classifier and PETs are proposed in order to build a model with good performance without losing too much transparency. The first model use naive Bayes estimation given a PET and the second model use a group of small-sized PETs as naive Bayes estimators. Empirical studies show that the first model outperforms the PET model at shallow depth and the second model is equivalent to naive Bayes and PET
Zengchang Qin
ICMLA1
2006 Market Mechanism Designs with Heterogeneous Trading Agents
abstract
Market mechanism design research is playing an important role in computational economics for resolving multi-agent allocation problems. A genetic algorithm was used to design auction mechanisms in order to automatically generate a desired market mechanism in agent based E-markets. In previous research, a hybrid market was studied, in which the probability that buyers rather than sellers are able to quote on a given time step, this probability was adapted by the GA which attempted to minimise Smith's coefficient of convergence. However, in previous experiments, all trading agents involved are of the same type or have identical preferences. This assumption does not hold in real-world markets which are always populated with heterogeneous agents. In this paper, the research of using evolutionary computing methods for auction designs is extended by using heterogeneous trading agents
Zengchang Qin
ICMLA1
2005 Hybrid Bayesian Estimation Trees Based on Label Semantics
Zengchang Qin, Jonathan Lawry
ECSQARU1
2005 Decision tree learning with fuzzy labels
Zengchang Qin, Jonathan Lawry
Inf. Sci.1