EDBT 2026 Demo / reviewers in the wild / expert
Yang Zhang 0002
dblp:06/6785-2
· DBLP profile ↗
42ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0002-7260-5098ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-shot skeleton-based action recognition with dual visual-text alignment
Jidong Kuang, Hongsong Wang 0001, Chaolei Han 0001, Yang Zhang 0002, Jie Gui |
Pattern Recognit. | 4 |
| 2026 | Efficient Neural Architecture Search for brain-inspired spiking neural network
Peilin Lai, Yang Zhang 0002, Weizhao He, Hongsong Wang 0001 |
Pattern Recognit. | 2 |
| 2026 | A Bootstrap Pipeline for Chat-Based Image Retrieval with Effective Question GenerationabstractChat-based image retrieval uses Large Language Models (LLM) to guide user input to enable more specific and precise search results, where LLM can enhance this process by asking user retrieval-oriented questions eliciting additional details about the target image. Despite the potential of this approach, no specialized Questioner model has been developed for this task due to the following significant challenges: (a) the difficulty of determining the optimal questions to ask; (b) the lack of a suitable protocol for fair model comparison; and (c) the notable scarcity of dialog-to-image retrieval data. To address these challenges, two fundamental principles are developed in this article to ensure the simplicity and effectiveness of the generated questions while enabling a fair comparison and accurate estimation of data quality and model performance. A bootstrap training methodology is introduced to collect retrieval-oriented dialog data and concurrently train the Questioner and the image Retriever. Under a fair comparison protocol, our extensive experiments have demonstrated that our proposed method can not only address the critical data gap but also achieve state-of-the-art results, which substantially surpass GPT-4o and GPT-4-Turbo through the fine-tuning of an 8B model. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with AttributesabstractRecent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capacity for fine-grained linguistic comprehension and leads to a significant decline in performance when faced with detailed descriptions or contextual information. To tackle these problems, we develop DoGA: Detect objects with Grouped Attributes, which employs commonly apparent attributes to bridge different granular semantics and uses specific attributes to identify the object discrepancy. Our DoGA incorporates three principle components: 1) Generation of attribute-based prompts, consisting of linguistic definitions enriched with common-sense visible attributes and hard negative notations deriving from the image-specific attribute features; 2) Paralleled entity fusion and optimization, designed to manage long attribute-based descriptions and negative concepts efficiently; and 3) Prompt-wise grouped training to accommodate model to perform many-to-many assignments, facilitating simultaneous training and inferring with multiple attribute-based synonyms. Extensive experiments demonstrate that training with synonymous attribute-based prompts allows DoGA to generalize multi-granular prompts and surpass previous state-of-the-art approaches, yielding 50.2 on the COCO and 38.0 on the LVIS benchmarks under the zero-short setting. We will make our code publicly available upon acceptance. Yang Liu 0250, Feng Hou, Yunjie Peng, Gangjian Zhang, Yao Zhang 0010, Peng Wang 0095, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
AAAI | 8 |
| 2025 | LOTA: Bit-Planes Guided AI-Generated Image Detection
Hongsong Wang 0001, Renxi Cheng, Yang Zhang 0002, Chaolei Han 0001, Jie Gui |
ICCV | 3 |
| 2025 | Learning Preference Distributions: A Label-Side Paradigm for Explainable Reward ModelsabstractThe reward model is a critical component in training powerful large language models. However, current methods largely overlook the inherent subjectivity and variability in human preferences. Typically, this issue is indirectly addressed through model-side approaches, such as ensemble methods to estimate uncertainty from multiple predictions or uncertainty-aware regression to predict mean and variance. These indirect approaches fail to capture the intrinsic distributional characteristics and inter-rater disagreements present in human judgments. In contrast, we propose a direct, label-side solution by explicitly modeling human preference distributions. We recover missing information from scalar ratings to construct meaningful distribution labels. By employing Label Distribution Learning (LDL), each dimension of the resulting multidimensional discrete distribution explicitly corresponds to a specific preference score, naturally reflecting the subjective and multidimensional nature of human evaluations. Our approach improves explainability and confidence estimation, while also enabling more effective data selection and sample-efficient test-time adaptation. Empirical results demonstrate that our method not only achieves state-of-the-art performance but also provides a robust framework for uncertainty quantification and nuanced preference modeling. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
SMC | 3 |
| 2025 | Collective domain adversarial learning for unsupervised domain adaptation
Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
Frontiers Comput. Sci. | 3 |
| 2025 | Achieving Procedure-Aware Instructional Video Correlation Learning Under Weak Supervision from a Collaborative Perspective
Tianyao He, Huabin Liu 0001, Zelin Ni, Yuxi Li 0009, Yang Zhang 0002, Weiyao Lin |
Int. J. Comput. Vis. | 7 |
| 2025 | Epistemic graph: A plug-and-play module for hybrid representation learningabstractIn recent years, deep models have achieved remarkable success in various vision tasks. However, their performance heavily relies on large training datasets. In contrast, humans exhibit hybrid learning, seamlessly integrating structured knowledge for cross-domain recognition or relying on a smaller amount of data samples for few-shot learning. Motivated by this human-like epistemic process, we aim to extend hybrid learning to computer vision tasks by integrating structured knowledge with data samples for more effective representation learning. Nevertheless, this extension faces significant challenges due to the substantial gap between structured knowledge and deep features learned from data samples, encompassing both dimensions and knowledge granularity. In this paper, a novel Epistemic Graph Layer (EGLayer) is introduced to enable hybrid learning, enhancing the exchange of information between deep features and a structured knowledge graph. Our EGLayer is composed of three major parts, including a local graph module, a query aggregation model, and a novel correlation alignment loss function to emulate human epistemic ability. Serving as a plug-and-play module that can replace the standard linear classifier, EGLayer significantly improves the performance of deep models. Extensive experiments demonstrate that EGLayer can greatly enhance representation learning for the tasks of cross-domain recognition and few-shot learning, and the visualization of knowledge graphs can aid in model interpretation. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Yong Rui |
Neurocomputing | 4 |
| 2025 | Gradient-aware domain-invariant learning for domain generalization
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
Multim. Syst. | 6 |
| 2025 | DomainVerse: A Benchmark Towards Real-World Distribution Shifts for Training-Free Adaptive Domain GeneralizationabstractTraditional cross-domain tasks, including unsupervised domain adaptation (UDA), domain generalization (DG) and test-time adaptation (TTA), rely heavily on the training model by source domain data whether for specific or arbitrary target domains. With the recent advance of vision-language models (VLMs), recognized as natural source models that can be transferred to various downstream tasks without any parameter training, we propose a novel cross-domain task directly combining the strengths of both UDA and DG, named Training-Free Adaptive Domain Generalization (TF-ADG). However, current cross-domain datasets have many limitations, such as unrealistic domains, unclear domain definitions, and the inability to fine-grained domain decomposition, which hinder the real-world application of current cross-domain models due to the lack of accurate and fair evaluation of fine-grained realistic domains. These insights motivate us to establish a novel realistic benchmark for TF-ADG. Benefiting from the introduced hierarchical definition of domain shifts, our proposed dataset DomainVerse addresses these issues by providing about 0.5 million images from 390 realistic, hierarchical, and balanced domains, allowing for decomposition across multiple domains within each image. With the help of the constructed DomainVerse and VLMs, we further propose two algorithms called Domain CLIP and Domain++ CLIP for training-free adaptive domain generalization. Extensive and comprehensive experiments demonstrate the significance of the dataset and the effectiveness of the proposed methods. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 6 |
| 2024 | Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video AnalysisabstractVideo Correlation Learning (VCL), which aims to analyze the relationships between videos, has been widely studied and applied in various general video tasks. However, applying VCL to instructional videos is still quite challenging due to their intrinsic procedural temporal structure. Specifically, procedural knowledge is critical for accurate correlation analyses on instructional videos. Nevertheless, current procedure-learning methods heavily rely on step-level annotations, which are costly and not scalable. To address this problem, we introduce a weakly supervised framework called Collaborative Procedure Alignment (CPA) for procedure-aware correlation learning on instructional videos. Our framework comprises two core modules: collaborative step mining and frame-to-step alignment. The collaborative step mining module enables simultaneous and consistent step segmentation for paired videos, leveraging the semantic and temporal similarity between frames. Based on the identified steps, the frame-to-step alignment module performs alignment between the frames and steps across videos. The alignment result serves as a measurement of the correlation distance between two videos. We instantiate our framework in two distinct instructional video tasks: sequence verification and action quality assessment. Extensive experiments validate the effectiveness of our approach in providing accurate and interpretable correlation analyses for instructional videos. Tianyao He, Huabin Liu 0001, Yuxi Li 0009, Yang Zhang 0002, Weiyao Lin |
AAAI | 6 |
| 2024 | TimeCraft: Navigate Weakly-Supervised Temporal Grounded Video Question Answering via Bi-directional Reasoning
Huabin Liu 0001, Yang Zhang 0002, Weiyao Lin |
ECCV (5) | 4 |
| 2024 | MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningabstractVideo causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily executed in a question-answering paradigm and focusing on short videos containing only a single event and simple causal relationships, lacking comprehensive and structured causality analysis for videos with multiple events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relationships between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD requires identifying the causal associations between these events to derive a comprehensive, structured event-level video causal diagram explaining why and how the final result event occurred. To address MECD, we devise a novel framework inspired by the Granger Causality method, using an efficient mask-based event prediction model to perform an Event Granger Test, which estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to address challenges in MECD like causality confounding and illusory causality. Experiments validate the effectiveness of our framework in providing causal relationships in multi-event videos, outperforming GPT-4o and VideoLLaVA by 5.7% and 4.1%, respectively. Tieyuan Chen, Huabin Liu 0001, Tianyao He, Yihang Chen 0002, Chaofan Gan, Yang Zhang 0002, Weiyao Lin |
NeurIPS | 8 |
| 2024 | Trust it or not: Confidence-guided automatic radiology report generation
Yixin Wang 0003, Zihao Lin 0003, Zhe Xu 0012, Jie Luo 0003, Jiang Tian, Zhongchao Shi, Lifu Huang, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002 |
Neurocomputing | 9 |
| 2024 | Learning rich features for gait recognition by integrating skeletons and silhouettes
Yunjie Peng, Yang Zhang 0002, Zhiqiang He 0002 |
Multim. Tools Appl. | 3 |
| 2024 | Domain-Aware Graph Network for Bridging Multi-Source Domain AdaptationabstractDomain adaptation (DA) addresses the challenge of distribution discrepancy between the training and test data, while multi-source domain adaptation (MSDA) is particularly appealing for realistic scenarios. With the emergence of extensive unlabeled datasets, self-supervised learning has gained significant popularity in deep learning. It is noteworthy that multi-source domain adaptation and self-supervised learning share a common objective: leveraging unlabeled data to acquire more informative representations. However, conventional self-supervised learning encounters two main limitations. Firstly, the traditional pretext task falls to transfer fine-grained knowledge to downstream task with general representation learning. Secondly, the scheme of the same feature extractor with distinct prediction heads makes the cross-task knowledge exchange and information sharing ineffective. In order to tackle these challenges, we introduce a novel approach called Domain-Aware Graph Network (DAGNet). DAGNet utilizes a graph neural network as a bridge to facilitate efficient cross-task knowledge exchange. By employing a mask token strategy, we enhance the robustness of representations by selectively masking certain domain or self-supervised information. In terms of datasets, the uneven and style-based domain shifts in current datasets make it challenging to measure the model's domain adaptation performance in real-world applications. To address this issue, we introduce a benchmark dataset DomainVerse with continuous spatio-temporal domain shifts encountered in the real world. Our extensive experiments demonstrate that DAGNet achieves state-of-the-art performance not only on mainstream multi-source domain adaptation datasets but also on different settings within DomainVerse. Code is available athttps://github.com/a791702141/SSG. Feng Hou, Yang Zhang 0002, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 4 |
| 2024 | A Survey of Visual TransformersabstractTransformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP). Inspired by such significant achievements, some pioneering works have recently been done on employing Transformer-liked architectures in the computer vision (CV) field, which have demonstrated their effectiveness on three fundamental CV tasks (classification, detection, and segmentation) as well as multiple sensory data stream (images, point clouds, and vision-language data). Because of their competitive modeling capabilities, the visual Transformers have achieved impressive performance improvements over multiple benchmarks as compared with modern convolution neural networks (CNNs). In this survey, we have reviewed over 100 of different visual Transformers comprehensively according to three fundamental CV tasks and different data stream types, where taxonomy is proposed to organize the representative methods according to their motivations, structures, and application scenarios. Because of their differences on training settings and dedicated vision tasks, we have also evaluated and compared all these existing visual Transformers under different configurations. Furthermore, we have revealed a series of essential but unexploited aspects that may empower such visual Transformers to stand out from numerous architectures, e.g., slack high-level semantic embeddings to bridge the gap between the visual Transformers and the sequential ones. Finally, two promising research directions are suggested for future investment. We will continue to update the latest articles and their released source codes at https://github.com/liuyang-ict/awesome-visual-transformers. Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | SAP-DETR: Bridging the Gap Between Salient Points and Queries-Based Transformer Detector for Fast Model ConvergencyabstractRecently, the dominant DETR-based approaches apply central-concept spatial prior to accelerating Transformer detector convergency. These methods gradually refine the reference points to the center of target objects and imbue object queries with the updated central reference information for spatially conditional attention. However, centralizing reference points may severely deteriorate queries' saliency and confuse detectors due to the indiscriminative spatial prior. To bridge the gap between the reference points of salient queries and Transformer detectors, we propose SAlient Point-based DETR (SAP-DETR) by treating object detection as a transformation from salient points to instance objects. Concretely, we explicitly initialize a query-specific reference point for each object query, gradually aggregate them into an instance object, and then predict the distance from each side of the bounding box to these points. By rapidly attending to query-specific reference regions and the conditional box edges, SAP-DETR can effectively bridge the gap between the salient point and the query-based Transformer detector with a significant convergency speed. Experimentally, SAP-DETR achieves 1.4× convergency speed with competitive performance and stably promotes the SoTA approaches by ∼1.0 AP. Based on ResNet-DC-101, SAP-DETR achieves 46.9 AP. The code will be released at https://github.com/liuyang-ict/SAP-DETR. Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
CVPR | 4 |
| 2023 | Learning How to Learn Domain-Invariant Parameters for Domain GeneralizationabstractDue to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs are optimized to extract domain-invariant representations, we expect a general model that is capable of well perceiving and emphatically updating such domain-invariant parameters. In this paper, we propose two modules of Domain Decoupling and Combination (DDC) and Domain-invariance-guided Backpropagation (DIGB), which can encourage such general model to focus on the parameters that have a unified optimization direction between pairs of contrastive samples. Our extensive experiments on two benchmarks have demonstrated that our proposed method has achieved state-of-the-art performance with strong generalization capability. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
ICASSP | 6 |
| 2022 | Cross-Domain Few-Shot Learning for Rare-Disease Skin Lesion SegmentationabstractRecently, deep learning (DL)-based skin lesion segmentation in dermoscopic images has advanced the efficient diagnosis of skin diseases. Commonly, most of the DL-based methods require a large amount of training data and can only perform accurate predictions on pre-defined classes. However, there exist some rare skin diseases with very limited labeled samples, which poses great challenges to typical DL-based methods. Few-shot learning (FSL) technique, which aims to train models with abundant seen classes and then generalizes to related unseen classes, is promising in addressing a similar problem. Unfortunately, simply borrowing the typical FSL is infeasible since collecting such abundant seen-class data (common skin diseases), is also difficult. In this paper, we propose a cross-domain few-shot segmentation (CD-FSS) framework, which enables the model to leverage the learning ability obtained from the natural domain, to facilitate rare-disease skin lesion segmentation with limited data of common diseases. Specifically, the framework consists of two processes, i.e., specific learning and generic learning, which are alternately optimized in a meta-training manner. A specific learner and a generic learner are tailored to build relationships between both processes. Experimental results demonstrate that our framework significantly improves the generalization ability from natural domain to unseen medical domain. Yixin Wang 0003, Zhe Xu 0012, Jiang Tian, Jie Luo 0003, Zhongchao Shi, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002 |
ICASSP | 6 |
| 2022 | A Novel Distilled Generative Essay Polish System via Hierarchical Pre-TrainingabstractIn language processing tasks, the most important process in automated text polishment always consists of text correction and text supplementation. Finding that text polishment is a necessary step in the field of English essay reviewing, we are motivated to be the first of building an end-to-end automated English essay polish system, to support writing instruction. There were independent methods for text correction tasks and text supplementation tasks, but when combining them for essay polishment tasks, conflicts arise from their interplay. In this paper, we propose a polish system that elegantly performs text correction and text supplementation at the same time, achieving an improved revision quality. Furthermore, we design a closed-loop essay polishing process, made up of a Rewriting Model and a Scoring Model, which refers to modified GPT2 and ensembled Bert respectively. The rewriting process targets the deficiencies of the essays by a threshold controlled mechanism. Lastly, the performance of our proposed system is further enhanced by an optimization method. Intensive experiments on both real data and simulated data have shown score improvements on full essays by Scoring Model, as well as higher text correction accuracy and longer text supplementation length. Qichuan Yang, Liuxin Zhang, Yang Zhang 0002, Jinghua Gao |
IJCNN | 3 |
| 2022 | mmFormer: Multimodal Medical Transformer for Incomplete Multimodal Learning of Brain Tumor Segmentation
Yao Zhang 0010, Nanjun He, Jiawei Yang 0002, Yuexiang Li, Dong Wei 0004, Yawen Huang, Yang Zhang 0002, Zhiqiang He 0002, Yefeng Zheng 0001 |
MICCAI (5) | 7 |
| 2021 | A Bi-modal Automated Essay Scoring System for Handwritten EssaysabstractDuring the past few decades, Automated Essay Scoring (AES) technology has been widely used to alleviate the workload of teachers and improve the feedback cycle in educational systems. However, the scoring of handwritten essays poses great challenges for existing systems, since most of them only take text as input without consideration of errors or bias which may be introduced by Optical Character Recognition (OCR) as a necessary pre-processing step. This paper proposes VisualAES, a bimodal automated essay scoring system that utilizes both textual and visual features for handwritten essay scoring. Specifically, we first employ three powerful pre-trained transformer-based models as the backbone, and extend them to take both textual features and visual features extracted by Faster R-CNN. Then, a stacking ensemble model is subsequently adopted to robustly map their outputs to a final score. We evaluate VisualAES with the public Automated Student Assessment Prize (ASAP) dataset and our proposed handwritten Chinese Students' handwritten essay dataset (ChnStd). Results show that the proposed VisualAES outperforms all state-of-the-art methods on both datasets. More importantly, by incorporating handwritten image information, we also achieve a further performance improvement on ChnStd and reduce the side effects of OCR. Jinghua Gao, Qichuan Yang, Yang Zhang 0002, Liuxin Zhang |
IJCNN | 3 |
| 2021 | VisDG: Towards Spontaneous Visual Dialogue GenerationabstractCurrent vision and language understanding tasks can be generally divided into two categories: image (video) description and visual question answering. Both of them aim to train a model that either straightforwardly generates a description or answers the predefined questions based on an image or a sequence of images from a spectators perspective. However, a large proportion of real-world human interactions also involve spontaneous dialogue exchanges among multiple speakers as well as dynamic visual-textual context, which requires an AI agent to hold a natural and open-ended dialog with humans in a first-person manner based on both visual and textual context. To move closer towards achieving such a spontaneous multimodal conversation, we introduce a new visual dialogue generation dataset (VisDG) based on keyframes and corresponding subtitles extracted from Friends - an American sitcom television series. Specifically, given a start frame and its corresponding dialogue text, the agent has to generate both a meaningful textual response as well as a correct image candidate for the latter part of the dialogue turn. Furthermore, we also propose an end-to-end image-text synergistic network (ITSN) for the task, which outperforms several sophisticated baselines on the proposed VisDG. Qichuan Yang, Liuxin Zhang, Yang Zhang 0002, Jinghua Gao |
IJCNN | 3 |
| 2021 | ACN: Adversarial Co-training Network for Brain Tumor Segmentation with Missing Modalities
Yixin Wang 0003, Yang Zhang 0002, Yang Liu 0250, Zihao Lin 0003, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
MICCAI (7) | 2 |
| 2021 | TumorCP: A Simple but Effective Object-Level Data Augmentation for Tumor Segmentation
Jiawei Yang 0002, Yao Zhang 0010, Yuan Liang 0001, Yang Zhang 0002, Lei He 0001, Zhiqiang He 0002 |
MICCAI (1) | 4 |
| 2021 | Modality-Aware Mutual Learning for Multi-modal Medical Image Segmentation
Yao Zhang 0010, Jiawei Yang 0002, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 6 |
| 2021 | Grabbing the Long Tail: A data normalization method for diverse and informative dialogue generation
Zhiqiang Zhan, Yang Zhang 0002, Jiangtao Gong, Qianying Wang 0002, Liuxin Zhang |
Neurocomputing | 3 |
| 2021 | Generative adversarial network for Table-to-Text generation
Zhiqiang Zhan, Rang Li, Changjian Hu, Yang Zhang 0002 |
Neurocomputing | 7 |
| 2020 | DARN: Deep Attentive Refinement Network for Liver Tumor Segmentation from 3D CT volumeabstractAutomatic liver tumor segmentation from 3D Computed Tomography (CT) is a necessary prerequisite in the interventions of hepatic abnormalities and surgery planning. However, accurate liver tumor segmentation remains challenging due to the large variability of tumor sizes and inhomogeneous texture. Recent advances based on Fully Convolutional Network (FCN) in liver tumor segmentation draw on success of learning discriminative multi-level features. In this paper, we propose a Deep Attentive Refinement Network (DARN) for improved liver tumor segmentation from CT volumes by fully exploiting both low and high level features embedded in different layers of FCN. Different from existing works, we exploit attention mechanism to leverage the relation of different levels of features encoded in different layers of FCN. Specifically, we introduce a Semantic Attention Refinement (SemRef) module to selectively emphasize global semantic information in low level features with the guidance of high level ones, and a Spatial Attention Refinement (SpaRef) module to adaptively enhance spatial details in high level features with the guidance of low level ones. We evaluate our network on the public MICCAI 2017 Liver Tumor Segmentation Challenge dataset (LiTS dataset) and it achieves state-of-the-art performance. The proposed refinement modules are an effective strategy to exploit multi-level features and has great potential to generalize to other medical image segmentation tasks. Yao Zhang 0010, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Zhiqiang He 0002 |
ICPR | 4 |
| 2020 | Double-Uncertainty Weighted Method for Semi-supervised Learning
Yixin Wang 0003, Yao Zhang 0010, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 6 |
| 2020 | Introspection unit in memory network: Learning to generalize inference in OOV scenarios
Qichuan Yang, Zhiqiang He 0002, Zhiqiang Zhan, Yang Zhang 0002, Rang Li, Changjian Hu |
Neurocomputing | 4 |
| 2020 | Knowledge attention sandwich neural network for text classification
Zhiqiang Zhan, Zifeng Hou, Qichuan Yang, Yang Zhang 0002, Changjian Hu |
Neurocomputing | 5 |
| 2019 | MIDS: End-to-End Personalized Response Generation in Untrimmed Multi-Role Dialogue*abstractMulti-role dialogue is a challenging issue in nature language process (NLP), which needs not only to understand the sentences, but also to simulate the interaction among roles. However, existing methods treat all roles’ speeches as one sequence and assume that only two speakers take turn to talk, which blurs the characteristics of roles and rarely happen in daily life. To address these issues, we propose a Multi-role Interposition Dialogue System (MIDS) which generates reasonable responses based on dialogue context and next speaker prediction. MIDS employs multiple role-defined encoders to understand each speaker, and an independent sequence model to predict the next speaker. The independent sequence model also works as a scheduler to integrate encoders with weights. Then, an attention-enhanced decoder generates responses based on dialogue context, speaker prediction and integrated encoders. Moreover, with the help of the unique speaker prediction, MIDS is able to generate diverse responses and join conversation actively when appropriate. Experimental results demonstrate that MIDS significantly improves the accuracy of speaker prediction and reduces the perplexity of generation over baselines. Furthermore, MIDS is able to interact with users without cue during real-life online conversations. This work marks a first step towards simulating multi-role dialogue generation. Qichuan Yang, Zhiqiang He 0002, Zhiqiang Zhan, Yang Zhang 0002, Changjian Hu |
IJCNN | 5 |
| 2019 | SSA: A More Humanized Automatic Evaluation Method for Open Dialogue GenerationabstractDialogue generation has been gaining ever-increasing attention, and various models have been proposed and adopted in many fields in recent years. How to evaluate their performance is critical. However, current evaluation metrics tend to be insufficient because of their simplicity and crudeness, resulting in weak correlation with human judgements. To solve this issue, we propose an automatic and comprehensive evaluation metric, which consists of three assessment criteria: Semantic Coherence, Syntactic Validity and Ability of Expression (SSA). The first two criteria are used to evaluate the generations from semantic and syntactic aspects respectively at the sentence level and the last one is to evaluate the overall performance at the model level. With two generative models, we conduct experiments on three datasets, including Twitter, Subtitle and Lenovo. Comparing with the previous metrics such as BLEU, METEOR and ROUGE, the correlation coefficient between SSA and human judgements is increased by 0.23-0.35, i.e. 324%-864% relative improvements. The experimental results demonstrate that SSA correlates more strongly with human judgements on the evaluation for open dialogue generation. Additionally, SSA is able to evaluate the semantic coherence and syntactic validity of generations exactly. More importantly, the evaluation models can be trained without human annotations. Thus, SSA is flexible and extensible to different datasets. Zhiqiang Zhan, Zifeng Hou, Qichuan Yang, Yang Zhang 0002, Changjian Hu |
IJCNN | 5 |
| 2019 | Discrimination Assessment for Saliency Maps
Ruiyi Li, Yangzhou Du, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
NLPCC (2) | 4 |
| 2018 | Adaptive Learning of Local Semantic and Global Structure Representations for Text ClassificationabstractRepresentation learning is a key issue for most Natural Language Processing (NLP) tasks. Most existing representation models either learn little structure information or just rely on pre-defined structures, leading to degradation of performance and generalization capability. This paper focuses on learning both local semantic and global structure representations for text classification. In detail, we propose a novel Sandwich Neural Network (SNN) to learn semantic and structure representations automatically without relying on parsers. More importantly, semantic and structure information contribute unequally to the text representation at corpus and instance level. To solve the fusion problem, we propose two strategies: Adaptive Learning Sandwich Neural Network (AL-SNN) and Self-Attention Sandwich Neural Network (SA-SNN). The former learns the weights at corpus level, and the latter further combines attention mechanism to assign the weights at instance level. Experimental results demonstrate that our approach achieves competitive performance on several text classification tasks, including sentiment analysis, question type classification and subjectivity classification. Specifically, the accuracies are MR (82.1%), SST-5 (50.4%), TREC (96%) and SUBJ (93.9%). Zhiqiang Zhan, Qichuan Yang, Yang Zhang 0002, Changjian Hu, Zhensheng Li, Liuxin Zhang, Zhiqiang He 0002 |
COLING | 4 |
| 2018 | SequentialSegNet: Combination with Sequential Feature for Multi-Organ SegmentationabstractMulti-organ segmentation from computed tomography (CT) images is essential for computer aided diagnosis (CAD), and recent advances in fully convolutional networks (FCNs) for volumetric image segmentation have demonstrated the importance of leveraging spatial information. In this paper, we propose a novel framework called SequentialSegNet, which efficiently combines features within a single CT image (intra-slice) and among multiple adjacent images (inter-slice) for a multi-organ segmentation. Experimental results show that our approach can effectively improve the segmentation performance on both large-size and small-size abdominal organs including liver, spleen and gallbladder. Yao Zhang 0010, Yang Zhang 0002, Zhongchao Shi, Zhensheng Li, Zhiqiang He 0002 |
ICPR | 4 |
| 2017 | Sequence-to-sequence prediction of personal computer software by recurrent neural networkabstractSequence to sequence (seq2seq) prediction is a key to many tasks of machine learning. Personal computer software sequence, as one of these tasks, was regarded as stochastic and unpredictable in the past. However, the deep neural networks (DNNs) have achieved excellent performance recently in sequence to sequence tasks, especially in the field of natural language process (NLP) such as language model, machine translation and dialogue systems. This paper examines the most popular DNNs approaches: LSTM, Encoder-Decoder network and Memory network in sequence prediction field to handle the software sequence learning and prediction task. Then three modified approaches based on these state of the art models are proposed to deal with additional information in sequence. These approaches focus on three aspects: adding information to enrich embedding input of Long-Short Term Memory, adding classifier to encoder-decoder neural network as an assistive model and processing data to be structured for memory unit in memory network. Experimental results based in real user data sets show that these proposed approaches outperform their corresponding standard DNNs and additional information can benefit the sequence neural network in different phases while constructing models. Experiments in different users shown that these modified strategies are robust and can be applied widely. Qichuan Yang, Zhiqiang He 0002, Fujiang Ge, Yang Zhang 0002 |
IJCNN | 4 |
| 2015 | Adaptive 3D facial action intensity estimation and emotion recognition
Yang Zhang 0002, Li Zhang 0013, M. Alamgir Hossain |
Expert Syst. Appl. | 1 |
| 2015 | Intelligent affect regression for bodily expressions using hybrid particle swarm optimization and adaptive ensembles
Yang Zhang 0002, Li Zhang 0013, Siew Chin Neoh, Kamlesh Mistry, M. Alamgir Hossain |
Expert Syst. Appl. | 1 |