VLDB 2026 Research / reviewers in the wild / expert
Zhiqiang He 0002
dblp:99/5332-2
· DBLP profile ↗
37ranked-venue papers
0as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 14 since 2021Artificial intelligence and machine learning · 16 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with AttributesabstractRecent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capacity for fine-grained linguistic comprehension and leads to a significant decline in performance when faced with detailed descriptions or contextual information. To tackle these problems, we develop DoGA: Detect objects with Grouped Attributes, which employs commonly apparent attributes to bridge different granular semantics and uses specific attributes to identify the object discrepancy. Our DoGA incorporates three principle components: 1) Generation of attribute-based prompts, consisting of linguistic definitions enriched with common-sense visible attributes and hard negative notations deriving from the image-specific attribute features; 2) Paralleled entity fusion and optimization, designed to manage long attribute-based descriptions and negative concepts efficiently; and 3) Prompt-wise grouped training to accommodate model to perform many-to-many assignments, facilitating simultaneous training and inferring with multiple attribute-based synonyms. Extensive experiments demonstrate that training with synonymous attribute-based prompts allows DoGA to generalize multi-granular prompts and surpass previous state-of-the-art approaches, yielding 50.2 on the COCO and 38.0 on the LVIS benchmarks under the zero-short setting. We will make our code publicly available upon acceptance. Yang Liu 0250, Feng Hou, Yunjie Peng, Gangjian Zhang, Yao Zhang 0010, Peng Wang 0095, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
AAAI | 12 |
| 2025 | Gradient-aware domain-invariant learning for domain generalization
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
Multim. Syst. | 9 |
| 2025 | DomainVerse: A Benchmark Towards Real-World Distribution Shifts for Training-Free Adaptive Domain GeneralizationabstractTraditional cross-domain tasks, including unsupervised domain adaptation (UDA), domain generalization (DG) and test-time adaptation (TTA), rely heavily on the training model by source domain data whether for specific or arbitrary target domains. With the recent advance of vision-language models (VLMs), recognized as natural source models that can be transferred to various downstream tasks without any parameter training, we propose a novel cross-domain task directly combining the strengths of both UDA and DG, named Training-Free Adaptive Domain Generalization (TF-ADG). However, current cross-domain datasets have many limitations, such as unrealistic domains, unclear domain definitions, and the inability to fine-grained domain decomposition, which hinder the real-world application of current cross-domain models due to the lack of accurate and fair evaluation of fine-grained realistic domains. These insights motivate us to establish a novel realistic benchmark for TF-ADG. Benefiting from the introduced hierarchical definition of domain shifts, our proposed dataset DomainVerse addresses these issues by providing about 0.5 million images from 390 realistic, hierarchical, and balanced domains, allowing for decomposition across multiple domains within each image. With the help of the constructed DomainVerse and VLMs, we further propose two algorithms called Domain CLIP and Domain++ CLIP for training-free adaptive domain generalization. Extensive and comprehensive experiments demonstrate the significance of the dataset and the effectiveness of the proposed methods. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 10 |
| 2024 | Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective
Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu 0008, Zhiqiang He 0002, Yongzhen Huang |
ECCV (6) | 6 |
| 2024 | Trust it or not: Confidence-guided automatic radiology report generation
Yixin Wang 0003, Zihao Lin 0003, Zhe Xu 0012, Jie Luo 0003, Jiang Tian, Zhongchao Shi, Lifu Huang, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002 |
Neurocomputing | 11 |
| 2024 | Learning rich features for gait recognition by integrating skeletons and silhouettes
Yunjie Peng, Yang Zhang 0002, Zhiqiang He 0002 |
Multim. Tools Appl. | 4 |
| 2024 | Domain-Aware Graph Network for Bridging Multi-Source Domain AdaptationabstractDomain adaptation (DA) addresses the challenge of distribution discrepancy between the training and test data, while multi-source domain adaptation (MSDA) is particularly appealing for realistic scenarios. With the emergence of extensive unlabeled datasets, self-supervised learning has gained significant popularity in deep learning. It is noteworthy that multi-source domain adaptation and self-supervised learning share a common objective: leveraging unlabeled data to acquire more informative representations. However, conventional self-supervised learning encounters two main limitations. Firstly, the traditional pretext task falls to transfer fine-grained knowledge to downstream task with general representation learning. Secondly, the scheme of the same feature extractor with distinct prediction heads makes the cross-task knowledge exchange and information sharing ineffective. In order to tackle these challenges, we introduce a novel approach called Domain-Aware Graph Network (DAGNet). DAGNet utilizes a graph neural network as a bridge to facilitate efficient cross-task knowledge exchange. By employing a mask token strategy, we enhance the robustness of representations by selectively masking certain domain or self-supervised information. In terms of datasets, the uneven and style-based domain shifts in current datasets make it challenging to measure the model's domain adaptation performance in real-world applications. To address this issue, we introduce a benchmark dataset DomainVerse with continuous spatio-temporal domain shifts encountered in the real world. Our extensive experiments demonstrate that DAGNet achieves state-of-the-art performance not only on mainstream multi-source domain adaptation datasets but also on different settings within DomainVerse. Code is available athttps://github.com/a791702141/SSG. Feng Hou, Yang Zhang 0002, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 8 |
| 2024 | A Survey of Visual TransformersabstractTransformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP). Inspired by such significant achievements, some pioneering works have recently been done on employing Transformer-liked architectures in the computer vision (CV) field, which have demonstrated their effectiveness on three fundamental CV tasks (classification, detection, and segmentation) as well as multiple sensory data stream (images, point clouds, and vision-language data). Because of their competitive modeling capabilities, the visual Transformers have achieved impressive performance improvements over multiple benchmarks as compared with modern convolution neural networks (CNNs). In this survey, we have reviewed over 100 of different visual Transformers comprehensively according to three fundamental CV tasks and different data stream types, where taxonomy is proposed to organize the representative methods according to their motivations, structures, and application scenarios. Because of their differences on training settings and dedicated vision tasks, we have also evaluated and compared all these existing visual Transformers under different configurations. Furthermore, we have revealed a series of essential but unexploited aspects that may empower such visual Transformers to stand out from numerous architectures, e.g., slack high-level semantic embeddings to bridge the gap between the visual Transformers and the sequential ones. Finally, two promising research directions are suggested for future investment. We will continue to update the latest articles and their released source codes at https://github.com/liuyang-ict/awesome-visual-transformers. Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 10 |
| 2024 | Neural Architecture Search for GNN-Based Graph ClassificationabstractGraph classification is an important problem with applications across many domains, for which graph neural networks (GNNs) have been state-of-the-art (SOTA) methods. In the literature, to adopt GNNs for the graph classification task, there are two groups of methods: global pooling and hierarchical pooling. The global pooling methods obtain the graph representation vectors by globally pooling all of the node embeddings together at the end of several GNN layers, whereas the hierarchical pooling methods provide one extra pooling operation between the GNN layers to extract hierarchical information and improve the graph representations. Both global and hierarchical pooling methods are effective in different scenarios. Due to highly diverse applications, it is challenging to design data-specific pooling methods with human expertise. To address this problem, we propose PAS (Pooling Architecture Search) to design adaptive pooling architectures by using the neural architecture search (NAS). To enable the search space design, we propose a unified pooling framework consisting of four modules: Aggregation, Pooling, Readout, and Merge. Two variants, PAS-G and PAS-NE, are provided to design the pooling operations in different scales. A set of candidate operations is designed in the search space using this framework. Then, existing human-designed pooling methods, including global and hierarchical ones, can be incorporated. To enable efficient search, a coarsening strategy is developed to continuously relax the search space, and then a differentiable search method can be adopted. We conduct extensive experiments on six real-world datasets, including the large-scale datasets MR and ogbg-molhiv. Experimental results in this article demonstrate the effectiveness and efficiency of the proposed PAS in designing the pooling architectures for graph classification. The Top-1 performance on two Open Graph Benchmark (OGB) datasets 1 further indicates the utility of PAS when facing diverse realistic data. The implementation of PAS is available at: https://github.com/AutoML-Research/PAS. Lanning Wei, Huan Zhao 0002, Zhiqiang He 0002, Quanming Yao |
ACM Trans. Inf. Syst. | 3 |
| 2024 | Deep Learning Based Occluded Person Re-Identification: A SurveyabstractOccluded person re-identification (Re-ID) focuses on addressing the occlusion problem when retrieving the person of interest across non-overlapping cameras. With the increasing demand for intelligent video surveillance and the application of person Re-ID technology, the real-world occlusion problem draws considerable interest from researchers. Although a large number of occluded person Re-ID methods have been proposed, there are few surveys that focus on occlusion. To fill this gap and help boost future research, this article provides a systematic survey of occluded person Re-ID. In this work, we review recent deep learning based occluded person Re-ID research. First, we summarize the main issues caused by occlusion as four groups: position misalignment, scale misalignment, noisy information, and missing information. Second, we categorize existing methods into six solution groups: matching, image transformation, multi-scale features, attention mechanism, auxiliary information, and contextual recovery. We also discuss the characteristics of each approach, as well as the issues they address. Furthermore, we present the performance comparison of recent occluded person Re-ID methods on four public datasets: Partial-ReID, Partial-iLIDS, Occluded-ReID, and Occluded-DukeMTMC. We conclude the study with thoughts on promising future research directions. Yunjie Peng, Jinlin Wu, Boqiang Xu, Chunshui Cao, Xu Liu 0008, Zhenan Sun, Zhiqiang He 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | SAP-DETR: Bridging the Gap Between Salient Points and Queries-Based Transformer Detector for Fast Model ConvergencyabstractRecently, the dominant DETR-based approaches apply central-concept spatial prior to accelerating Transformer detector convergency. These methods gradually refine the reference points to the center of target objects and imbue object queries with the updated central reference information for spatially conditional attention. However, centralizing reference points may severely deteriorate queries' saliency and confuse detectors due to the indiscriminative spatial prior. To bridge the gap between the reference points of salient queries and Transformer detectors, we propose SAlient Point-based DETR (SAP-DETR) by treating object detection as a transformation from salient points to instance objects. Concretely, we explicitly initialize a query-specific reference point for each object query, gradually aggregate them into an instance object, and then predict the distance from each side of the bounding box to these points. By rapidly attending to query-specific reference regions and the conditional box edges, SAP-DETR can effectively bridge the gap between the salient point and the query-based Transformer detector with a significant convergency speed. Experimentally, SAP-DETR achieves 1.4× convergency speed with competitive performance and stably promotes the SoTA approaches by ∼1.0 AP. Based on ResNet-DC-101, SAP-DETR achieves 46.9 AP. The code will be released at https://github.com/liuyang-ict/SAP-DETR. Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
CVPR | 8 |
| 2023 | Learning How to Learn Domain-Invariant Parameters for Domain GeneralizationabstractDue to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs are optimized to extract domain-invariant representations, we expect a general model that is capable of well perceiving and emphatically updating such domain-invariant parameters. In this paper, we propose two modules of Domain Decoupling and Combination (DDC) and Domain-invariance-guided Backpropagation (DIGB), which can encourage such general model to focus on the parameters that have a unified optimization direction between pairs of contrastive samples. Our extensive experiments on two benchmarks have demonstrated that our proposed method has achieved state-of-the-art performance with strong generalization capability. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
ICASSP | 9 |
| 2023 | Occluded Gait RecognitionabstractGait recognition suffers from common occlusions in real-world applications. However, academic research on gait recognition usually assumes access to full-body input data. For bridging the gap to practical applications, we propose to identify people when given an occluded gait sequence, namely occluded gait recognition. Since publicly available datasets do not meet the requirements of the intended research, we design a new framework named OccSilGait to generate realistic occluded gait silhouette sequences based on the principle of perspective transformation. Specifically, OccSilGait considers various occlusion scenarios including non-occlusion, crowd occlusion, static occlusion, and detection occlusion. And we employ OccSilGait to build the occluded gait dataset OccCASIA-B for further research. To address challenges brought by occlusion for gait recognition, we propose a novel SpaAlignTemOccRecover network consisting of 1) a Spatial auto-Align module that transforms silhouettes into spatially aligned ones with well-designed self-supervision; 2) a Spatial-Temporal Backbone that alternatively extracts spatial and temporal features to avoid the diffusion of occlusion; 3) a Temporal Occlusion Recovery module that reconstructs the current frame based on time index and temporal context, exploiting gait periodicity for occlusion recovery. Experiments on the newly built occluded dataset show the superiority of the proposed method. Both the OccSilGait framework and the code are available at https://github.com/YunjiePeng/OccludedGaitRecognition. Yunjie Peng, Chunshui Cao, Zhiqiang He 0002 |
IJCNN | 3 |
| 2023 | Search to Capture Long-range Dependency with Stacking GNNs for Graph ClassificationabstractIn recent years, Graph Neural Networks (GNNs) have been popular in the graph classification task. Currently, shallow GNNs are more common due to the well-known over-smoothing problem facing deeper GNNs. However, they are sub-optimal without utilizing the information from distant nodes, i.e., the long-range dependencies. The mainstream methods in the graph classification task can extract the long-range dependencies either by designing the pooling operations or incorporating the higher-order neighbors, while they have evident drawbacks by modifying the original graph structure, which may result in information loss in graph structure learning. In this paper, by justifying the smaller influence of the over-smoothing problem in the graph classification task, we evoke the importance of stacking-based GNNs and then employ them to capture the long-range dependencies without modifying the original graph structure. To achieve this, two design needs are given for stacking-based GNNs, i.e., sufficient model depth and adaptive skip-connection schemes. By transforming the two design needs into designing data-specific inter-layer connections, we propose a novel approach with the help of neural architecture search (NAS), which is dubbed LRGNN (Long-Range Graph Neural Networks). Extensive experiments on five datasets show that the proposed LRGNN can achieve the best performance, and obtained data-specific GNNs with different depth and skip-connection schemes, which can better capture the long-range dependencies. 1 Lanning Wei, Zhiqiang He 0002, Huan Zhao 0002, Quanming Yao |
WWW | 2 |
| 2022 | Cross-Domain Few-Shot Learning for Rare-Disease Skin Lesion SegmentationabstractRecently, deep learning (DL)-based skin lesion segmentation in dermoscopic images has advanced the efficient diagnosis of skin diseases. Commonly, most of the DL-based methods require a large amount of training data and can only perform accurate predictions on pre-defined classes. However, there exist some rare skin diseases with very limited labeled samples, which poses great challenges to typical DL-based methods. Few-shot learning (FSL) technique, which aims to train models with abundant seen classes and then generalizes to related unseen classes, is promising in addressing a similar problem. Unfortunately, simply borrowing the typical FSL is infeasible since collecting such abundant seen-class data (common skin diseases), is also difficult. In this paper, we propose a cross-domain few-shot segmentation (CD-FSS) framework, which enables the model to leverage the learning ability obtained from the natural domain, to facilitate rare-disease skin lesion segmentation with limited data of common diseases. Specifically, the framework consists of two processes, i.e., specific learning and generic learning, which are alternately optimized in a meta-training manner. A specific learner and a generic learner are tailored to build relationships between both processes. Experimental results demonstrate that our framework significantly improves the generalization ability from natural domain to unseen medical domain. Yixin Wang 0003, Zhe Xu 0012, Jiang Tian, Jie Luo 0003, Zhongchao Shi, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002 |
ICASSP | 8 |
| 2022 | mmFormer: Multimodal Medical Transformer for Incomplete Multimodal Learning of Brain Tumor Segmentation
Yao Zhang 0010, Nanjun He, Jiawei Yang 0002, Yuexiang Li, Dong Wei 0004, Yawen Huang, Yang Zhang 0002, Zhiqiang He 0002, Yefeng Zheng 0001 |
MICCAI (5) | 8 |
| 2022 | Designing the Topology of Graph Neural Networks: A Novel Feature Fusion PerspectiveabstractIn recent years, Graph Neural Networks (GNNs) have shown superior performance on diverse real-world applications. To improve the model capacity, besides designing aggregation operations, GNN topology design is also very important. In general, there are two mainstream GNN topology design manners. The first one is to stack aggregation operations to obtain the higher-level features but easily got performance drop as the network goes deeper. Secondly, the multiple aggregation operations are utilized in each layer which provides adequate and independent feature extraction stage on local neighbors while are costly to obtain the higher-level information. To enjoy the benefits while alleviating the corresponding deficiencies of these two manners, we learn to design the topology of GNNs in a novel feature fusion perspective which is dubbed F2GNN. To be specific, we provide a feature fusion perspective in designing GNN topology and propose a novel framework to unify the existing topology designs with feature selection and fusion strategies. Then we develop a neural architecture search method on top of the unified framework which contains a set of selection and fusion operations in the search space and an improved differentiable search algorithm. The performance gains on diverse datasets, five homophily and three heterophily ones, demonstrate the effectiveness of F2GNN. We further conduct experiments to show that F2GNN can improve the model capacity while alleviating the deficiencies of existing GNN topology design manners, especially alleviating the over-smoothing problem, by utilizing different levels of features adaptively. 1 Lanning Wei, Huan Zhao 0002, Zhiqiang He 0002 |
WWW | 3 |
| 2021 | Pooling Architecture Search for Graph ClassificationabstractGraph classification is an important problem with applications across many domains, like chemistry and bioinformatics, for which graph neural networks (GNNs) have been state-of-the-art (SOTA) methods. GNNs are designed to learn node-level representation based on neighborhood aggregation schemes, and to obtain graph-level representation, pooling methods are applied after the aggregation operation in existing GNN models to generate coarse-grained graphs. However, due to highly diverse applications of graph classification, and the performance of existing pooling methods vary on different graphs. In other words, it is a challenging problem to design a universal pooling architecture to perform well in most cases, leading to a demand for data-specific pooling methods in real-world applications. To address this problem, we propose to use neural architecture search (NAS) to search for adaptive pooling architectures for graph classification. Firstly we designed a unified framework consisting of four modules: Aggregation, Pooling, Readout, and Merge, which can cover existing human-designed pooling methods for graph classification. Based on this framework, a novel search space is designed by incorporating popular operations in human-designed architectures. Then to enable efficient search, a coarsening strategy is proposed to continuously relax the search space, thus a differentiable search method can be adopted. Extensive experiments on six real-world datasets from three domains are conducted, and the results demonstrate the effectiveness and efficiency of the proposed framework1 Lanning Wei, Huan Zhao 0002, Quanming Yao, Zhiqiang He 0002 |
CIKM | 4 |
| 2021 | ACN: Adversarial Co-training Network for Brain Tumor Segmentation with Missing Modalities
Yixin Wang 0003, Yang Zhang 0002, Yang Liu 0250, Zihao Lin 0003, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
MICCAI (7) | 9 |
| 2021 | TumorCP: A Simple but Effective Object-Level Data Augmentation for Tumor Segmentation
Jiawei Yang 0002, Yao Zhang 0010, Yuan Liang 0001, Yang Zhang 0002, Lei He 0001, Zhiqiang He 0002 |
MICCAI (1) | 6 |
| 2021 | Modality-Aware Mutual Learning for Multi-modal Medical Image Segmentation
Yao Zhang 0010, Jiawei Yang 0002, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 7 |
| 2021 | VerSe: A Vertebrae labelling and segmentation benchmark for multi-detector CT images
Anjany Sekuboyina, Malek El Husseini, Amirhossein Bayat, Maximilian Löffler, Hans Liebl, Hongwei Li 0004, Giles Tetteh, Jan Kukacka, Christian Payer, Darko Stern, Martin Urschler, Maodong Chen, Dalong Cheng, Nikolas Leßmann, Yujin Hu, Tianfu Wang 0001, Dong Yang 0005, Daguang Xu, Felix Ambellan, Tamaz Amiranashvili, Moritz Ehlke, Hans Lamecker, Sebastian Lehnert, Marilia Lirio, Nicolás Pérez de Olaguer, Heiko Ramm, Manish Sahu, Alexander Tack, Stefan Zachow, Xinjun Ma, Christoph Angerman, Xin Wang 0113, Alexandre Kirszenberg, Élodie Puybareau, Yiwei Bai, Brandon H. Rapazzo, Timyoas Yeah, Amber Zhang, Shangliang Xu, Feng Hou, Zhiqiang He 0002, Chan Zeng, Zheng Xiangshang, Xu Liming, Tucker J. Netherton, Raymond P. Mumme, Laurence E. Court, Zixun Huang, Chenhang He, Li-Wen Wang, Sai-Ho Ling, Lê Duy Huynh, Nicolas Boutry, Roman Jakubícek, Jirí Chmelík, Supriti Mulay, Mohanasankar Sivaprakasam, Johannes C. Paetzold, Suprosanna Shit, Ivan Ezhov, Benedikt Wiestler, Ben Glocker, Alexander Valentinitsch, Markus Rempfler, Bjoern Menze, Jan Kirschke |
Medical Image Anal. | 44 |
| 2021 | Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation: The M&Ms ChallengeabstractThe emergence of deep learning has considerably advanced the state-of-the-art in cardiac magnetic resonance (CMR) segmentation. Many techniques have been proposed over the last few years, bringing the accuracy of automated segmentation close to human performance. However, these models have been all too often trained and validated using cardiac imaging samples from single clinical centres or homogeneous imaging protocols. This has prevented the development and validation of models that are generalizable across different clinical centres, imaging conditions or scanner vendors. To promote further research and scientific benchmarking in the field of generalizable deep learning for cardiac segmentation, this paper presents the results of the Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation (M&Ms) Challenge, which was recently organized as part of the MICCAI 2020 Conference. A total of 14 teams submitted different solutions to the problem, combining various baseline models, data augmentation strategies, and domain adaptation techniques. The obtained results indicate the importance of intensity-driven data augmentation, as well as the need for further research to improve generalizability towards unseen scanner vendors or new imaging protocols. Furthermore, we present a new resource of 375 heterogeneous CMR datasets acquired by using four different scanner vendors in six hospitals and three different countries (Spain, Canada and Germany), which we provide as open-access for the community to enable future research in the field. Víctor M. Campello, Polyxeni Gkontra, Cristian Izquierdo, Carlos Martín-Isla, Alireza Sojoudi, Peter M. Full, Klaus H. Maier-Hein, Yao Zhang 0010, Zhiqiang He 0002, Jun Ma 0016, Mario Parreño, Alberto Albiol, Fanwei Kong, Shawn C. Shadden, Jorge Corral Acero, Vaanathi Sundaresan, Mina Saber, Mustafa A. Alattar, Hongwei Li 0004, Bjoern Menze, Firas Khader, Christoph Haarburger, Cian M. Scannell, Mitko Veta, Adam Carscadden, Kumaradevan Punithakumar, Xiao Liu 0037, Sotirios A. Tsaftaris, Xiaoqiong Huang, Xin Yang 0009, Lei Li 0020, Xiahai Zhuang, David Viladés, Martín Luís Descalzo, Andrea Guala 0002, Lucia La Mura, Matthias G. W. Friedrich, Ria Garg, Julie Lebel, Filipe Henriques, Mahir Karakas, Ersin Çavus, Steffen E. Petersen, Sergio Escalera, Santi Seguí, Jose Rodriguez-Palomares, Karim Lekadir |
IEEE Trans. Medical Imaging | 9 |
| 2020 | GaitPart: Temporal Part-Based Model for Gait RecognitionabstractGait recognition, applied to identify individual walking patterns in a long-distance, is one of the most promising video-based biometric technologies. At present, most gait recognition methods take the whole human body as a unit to establish the spatio-temporal representations. However, we have observed that different parts of human body possess evidently various visual appearances and movement patterns during walking. In the latest literature, employing partial features for human body description has been verified being beneficial to individual recognition. Taken above insights together, we assume that each part of human body needs its own spatio-temporal expression. Then, we propose a novel part-based model GaitPart and get two aspects effect of boosting the performance: On the one hand, Focal Convolution Layer, a new applying of convolution, is presented to enhance the fine-grained learning of the part-level spatial features. On the other hand, the Micro-motion Capture Module (MCM) is proposed and there are several parallel MCMs in the GaitPart corresponding to the pre-defined parts of the human body, respectively. It is worth mentioning that the MCM is a novel way of temporal modeling for gait task, which focuses on the short-range temporal features rather than the redundant long-range features for cycle gait. Experiments on two of the most popular public datasets, CASIA-B and OU-MVLP, richly exemplified that our method meets a new state-of-the-art on multiple standard benchmarks. The source code will be available on https://github.com/ChaoFan96/GaitPart. Chao Fan 0001, Yunjie Peng, Chunshui Cao, Xu Liu 0008, Saihui Hou, Jiannan Chi, Yongzhen Huang, Qing Li 0015, Zhiqiang He 0002 |
CVPR | 9 |
| 2020 | DARN: Deep Attentive Refinement Network for Liver Tumor Segmentation from 3D CT volumeabstractAutomatic liver tumor segmentation from 3D Computed Tomography (CT) is a necessary prerequisite in the interventions of hepatic abnormalities and surgery planning. However, accurate liver tumor segmentation remains challenging due to the large variability of tumor sizes and inhomogeneous texture. Recent advances based on Fully Convolutional Network (FCN) in liver tumor segmentation draw on success of learning discriminative multi-level features. In this paper, we propose a Deep Attentive Refinement Network (DARN) for improved liver tumor segmentation from CT volumes by fully exploiting both low and high level features embedded in different layers of FCN. Different from existing works, we exploit attention mechanism to leverage the relation of different levels of features encoded in different layers of FCN. Specifically, we introduce a Semantic Attention Refinement (SemRef) module to selectively emphasize global semantic information in low level features with the guidance of high level ones, and a Spatial Attention Refinement (SpaRef) module to adaptively enhance spatial details in high level features with the guidance of low level ones. We evaluate our network on the public MICCAI 2017 Liver Tumor Segmentation Challenge dataset (LiTS dataset) and it achieves state-of-the-art performance. The proposed refinement modules are an effective strategy to exploit multi-level features and has great potential to generalize to other medical image segmentation tasks. Yao Zhang 0010, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Zhiqiang He 0002 |
ICPR | 6 |
| 2020 | Double-Uncertainty Weighted Method for Semi-supervised Learning
Yixin Wang 0003, Yao Zhang 0010, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 7 |
| 2020 | Introspection unit in memory network: Learning to generalize inference in OOV scenarios
Qichuan Yang, Zhiqiang He 0002, Zhiqiang Zhan, Yang Zhang 0002, Rang Li, Changjian Hu |
Neurocomputing | 2 |
| 2019 | MIDS: End-to-End Personalized Response Generation in Untrimmed Multi-Role Dialogue*abstractMulti-role dialogue is a challenging issue in nature language process (NLP), which needs not only to understand the sentences, but also to simulate the interaction among roles. However, existing methods treat all roles’ speeches as one sequence and assume that only two speakers take turn to talk, which blurs the characteristics of roles and rarely happen in daily life. To address these issues, we propose a Multi-role Interposition Dialogue System (MIDS) which generates reasonable responses based on dialogue context and next speaker prediction. MIDS employs multiple role-defined encoders to understand each speaker, and an independent sequence model to predict the next speaker. The independent sequence model also works as a scheduler to integrate encoders with weights. Then, an attention-enhanced decoder generates responses based on dialogue context, speaker prediction and integrated encoders. Moreover, with the help of the unique speaker prediction, MIDS is able to generate diverse responses and join conversation actively when appropriate. Experimental results demonstrate that MIDS significantly improves the accuracy of speaker prediction and reduces the perplexity of generation over baselines. Furthermore, MIDS is able to interact with users without cue during real-life online conversations. This work marks a first step towards simulating multi-role dialogue generation. Qichuan Yang, Zhiqiang He 0002, Zhiqiang Zhan, Yang Zhang 0002, Changjian Hu |
IJCNN | 2 |
| 2019 | Discrimination Assessment for Saliency Maps
Ruiyi Li, Yangzhou Du, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
NLPCC (2) | 5 |
| 2019 | Patch-based self-adaptive matting for high-resolution image and video
Guangying Cao, Xiaowu Chen 0001, Zhiqiang He 0002 |
Vis. Comput. | 4 |
| 2018 | Adaptive Learning of Local Semantic and Global Structure Representations for Text ClassificationabstractRepresentation learning is a key issue for most Natural Language Processing (NLP) tasks. Most existing representation models either learn little structure information or just rely on pre-defined structures, leading to degradation of performance and generalization capability. This paper focuses on learning both local semantic and global structure representations for text classification. In detail, we propose a novel Sandwich Neural Network (SNN) to learn semantic and structure representations automatically without relying on parsers. More importantly, semantic and structure information contribute unequally to the text representation at corpus and instance level. To solve the fusion problem, we propose two strategies: Adaptive Learning Sandwich Neural Network (AL-SNN) and Self-Attention Sandwich Neural Network (SA-SNN). The former learns the weights at corpus level, and the latter further combines attention mechanism to assign the weights at instance level. Experimental results demonstrate that our approach achieves competitive performance on several text classification tasks, including sentiment analysis, question type classification and subjectivity classification. Specifically, the accuracies are MR (82.1%), SST-5 (50.4%), TREC (96%) and SUBJ (93.9%). Zhiqiang Zhan, Qichuan Yang, Yang Zhang 0002, Changjian Hu, Zhensheng Li, Liuxin Zhang, Zhiqiang He 0002 |
COLING | 8 |
| 2018 | SequentialSegNet: Combination with Sequential Feature for Multi-Organ SegmentationabstractMulti-organ segmentation from computed tomography (CT) images is essential for computer aided diagnosis (CAD), and recent advances in fully convolutional networks (FCNs) for volumetric image segmentation have demonstrated the importance of leveraging spatial information. In this paper, we propose a novel framework called SequentialSegNet, which efficiently combines features within a single CT image (intra-slice) and among multiple adjacent images (inter-slice) for a multi-organ segmentation. Experimental results show that our approach can effectively improve the segmentation performance on both large-size and small-size abdominal organs including liver, spleen and gallbladder. Yao Zhang 0010, Yang Zhang 0002, Zhongchao Shi, Zhensheng Li, Zhiqiang He 0002 |
ICPR | 7 |
| 2017 | Sequence-to-sequence prediction of personal computer software by recurrent neural networkabstractSequence to sequence (seq2seq) prediction is a key to many tasks of machine learning. Personal computer software sequence, as one of these tasks, was regarded as stochastic and unpredictable in the past. However, the deep neural networks (DNNs) have achieved excellent performance recently in sequence to sequence tasks, especially in the field of natural language process (NLP) such as language model, machine translation and dialogue systems. This paper examines the most popular DNNs approaches: LSTM, Encoder-Decoder network and Memory network in sequence prediction field to handle the software sequence learning and prediction task. Then three modified approaches based on these state of the art models are proposed to deal with additional information in sequence. These approaches focus on three aspects: adding information to enrich embedding input of Long-Short Term Memory, adding classifier to encoder-decoder neural network as an assistive model and processing data to be structured for memory unit in memory network. Experimental results based in real user data sets show that these proposed approaches outperform their corresponding standard DNNs and additional information can benefit the sequence neural network in different phases while constructing models. Experiments in different users shown that these modified strategies are robust and can be applied widely. Qichuan Yang, Zhiqiang He 0002, Fujiang Ge, Yang Zhang 0002 |
IJCNN | 2 |
| 2017 | A survey on context-aware mobile visual recognition
Weiqing Min, Shuqiang Jiang, Shuhui Wang, Ruihan Xu 0001, Yushan Cao, Luis Herranz, Zhiqiang He 0002 |
Multim. Syst. | 7 |
| 2017 | Modality-specific and hierarchical feature learning for RGB-D hand-held object recognition
Xiong Lv, Xinda Liu, Xiangyang Li 0002, Shuqiang Jiang, Zhiqiang He 0002 |
Multim. Tools Appl. | 6 |
| 2016 | Local Shape Transfer for Image Co-segmentation
Wei Teng, Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Zhiqiang He 0002 |
BMVC | 5 |
| 2014 | A virtual machine based task scheduling approach to improving data locality for virtualized HadoopabstractMapReduce emerges as an important distributed programming paradigm for large-scale data analysis applications. As an open-source implementation of MapReduce, Hadoop presents an attractive usage system for many enterprises. There are some drawbacks in a traditional Hadoop cluster deployed with a large scale of physical machines, such as burdensome cluster management and fluctuating resource utilization. Virtualized Hadoop cluster not only simplifies cluster management, but also facilitates cost-effective workload consolidation for resource utilization. In Hadoop system, the data locality is a critical factor impacting on performance of MapReduce applications. However, existing task scheduling approaches to improving data locality of virtualized Hadoop are not effective because of two levels distribution of data: virtual machines and physical servers. In this paper, we deploy virtualized Hadoop cluster in which computing node and storage node are placed in respective virtual machines to improve flexibility. We propose a novel task scheduling approach which aims to improve data locality for virtualized Hadoop cluster through migrating the virtual machine acted as computing node to the physical server running virtual machine acted as storage node that holds a data replica needed by that computing node. We evaluated our approach's efficiency on a virtualized Hadoop cluster with the aforementioned deployment for 11 computing nodes and 12 storage nodes. Our experiment results show that our approach improves performance of 86% typical MapReduce applications in our benchmark suite at varying degrees. Zhiqiang He 0002 |
ICIS | 4 |