EDBT 2026 Demo / reviewers in the wild / expert
Jian Zhang 0026
dblp:07/314-26
· DBLP profile ↗
35ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0001-6478-9192ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Divide-and-conquer towards optimal adaptation of pre-trained model to medical tasks
Zhanghui Huang, Zunlei Feng, Xiaoyan Sun 0006, Shuifa Sun, Zhenming Yuan, Jun Yu 0002, Jian Zhang 0026 |
Pattern Recognit. | 7 |
| 2025 | Adaptive Multimodal Fusion via Attention-Guided Feature Selection for Histopathology Image ClassificationabstractThis paper presents a framework for adaptive multimodal feature fusion that employs attention-based feature selection mechanisms to enhance the classification of histopathology images. The framework integrates medical images and clinical texts through three core modules: the Unified Feature Processing Module (UFPM) for standardized feature preprocessing, the Cross-Modal Attention Module (CMAM) for facilitating interactions between image and text features, and the Selective Feature Alignment Module (SFAM) for aligning features across different modalities. Experimental results on the Quilt-BCGG and Quilt-Derm4 datasets demonstrate the framework's good classification performance, which improves the accuracy and efficiency of histopathology diagnosis through optimized feature selection and alignment. The link to the code: https://github.com/leibabaya/Attention-guided-Adaptive-Fusion Jiaxin Lei, Kaihao He, Xiaoyan Sun 0006, Zhenming Yuan, Jian Zhang 0026 |
ICIP | 5 |
| 2025 | Self-adaptive image-text fusion for medical image classification
Jian Zhang 0026, Kaihao He, Zunlei Feng, Shuifa Sun, Xiaoyan Sun 0006, Zhenming Yuan, Jun Yu 0002 |
Pattern Recognit. | 1 |
| 2025 | Semi-Supervised RGB-D Hand Gesture Recognition via Mutual Learning of Self-Supervised ModelsabstractHuman hand gesture recognition is important to human–computer interaction. Gesture recognition based on RGB and Depth (RGB-D) data exploits both RGB and depth images to provide comprehensive results. However, the research under scenario with insufficient annotated data is not adequate. In view of the problem, our insight is to perform self-supervised learning with respect to each modality, transfer the learned information to modality-specific classifiers, and then fuse their results for final decision. To this end, we propose a semi-supervised hand gesture recognition method known as Mutual Learning of Rotation-Aware Gesture Predictors (MLRAGP), which exploits unlabeled training RGB and depth images via self-supervised learning and achieves multi-modal decision fusion through deep mutual learning. For each modality, we rotate both labeled and unlabeled images to fixed angles and train an angle predictor to predict the angles, then we use the feature extraction part of the angle predictor to construct the category predictor and train it through labeled data. We subsequently fuse the category predictors about both modalities by impelling each of them to simulate the probability estimation produced by the other, and making the prediction of labeled images to approach the ground truth annotation. During the training of category predictor and mutual learning, the parameters of feature extractors can be slighted fine-tuned to avoid under-fitting. Experimental results on NTU-Microsoft Kinect Hand Gesture dataset and Washington RGB-D dataset demonstrate the superiority of this framework to existing methods. Jian Zhang 0026, Kaihao He, Ting Yu 0016, Jun Yu 0002, Zhenming Yuan |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | 3D human pose estimation with multi-hypotheses gated transformer
Xiena Dong, Jian Zhang 0026, Jun Yu 0002, Ting Yu 0016 |
Multim. Syst. | 2 |
| 2024 | Multi-Granularity Contrastive Cross-Modal Collaborative Generation for End-to-End Long-Term Video Question AnsweringabstractLong-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-form questions, simultaneously emphasizing comprehensive cross-modal reasoning to yield precise answers. The canonical approaches often rely on off-the-shelf feature extractors to detour the expensive computation overhead, but often result in domain-independent modality-unrelated representations. Furthermore, the inherent gradient blocking between unimodal comprehension and cross-modal interaction hinders reliable answer generation. In contrast, recent emerging successful video-language pre-training models enable cost-effective end-to-end modeling but fall short in domain-specific ratiocination and exhibit disparities in task formulation. Toward this end, we present an entirely end-to-end solution for long-term VideoQA: Multi-granularity Contrastive cross-modal collaborative Generation (MCG) model. To derive discriminative representations possessing high visual concepts, we introduce Joint Unimodal Modeling (JUM) on a clip-bone architecture and leverage Multi-granularity Contrastive Learning (MCL) to harness the intrinsically or explicitly exhibited semantic correspondences. To alleviate the task formulation discrepancy problem, we propose a Cross-modal Collaborative Generation (CCG) module to reformulate VideoQA as a generative task instead of the conventional classification scheme, empowering the model with the capability for cross-modal high-semantic fusion and generation so as to rationalize and answer. Extensive experiments conducted on six publicly available VideoQA datasets underscore the superiority of our proposed method. Ting Yu 0016, Kunhao Fu, Jian Zhang 0026, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Image Process. | 3 |
| 2024 | Semi-Supervised Medical Report Generation via Graph-Guided Hybrid Feature ConsistencyabstractMedical report generation generates the corresponding report according to the given radiology image, which has been attracting increasing research interest. However, existing methods mainly adopt supervised training which rely on large amount of medical reports that are actually unavailable owing to the labor-intensive labeling process and privacy protection protocol. In the meanwhile, the intrinsic relationships between local pathological changes in the image are often ignored, which actually are important hints to high quality report generation. To this end, we propose a Relation-Aware Mean Teacher (RAMT) framework, which follows a standard mean teacher paradigm for semi-supervised report generation. The key to the encoder of the backbone network is the Graph-guided Hybrid Feature Encoding (GHFE) module, which exploits a prior disease knowledge graph to encode the intrinsic relations between pathological changes into the graph embedding and learns a word dictionary to retrieve the semantic embedding for each potential pathological change. GHFE combines the graph embedding, semantic embedding and visual features to form hybrid features, which are sent to a Transformer-based decoder for report generation. Extensive experiments on the MIMIC-CXR and IU X-Ray datasets demonstrate the effectiveness of our proposed approach. Ke Zhang 0029, Hanliang Jiang, Jian Zhang 0026, Qingming Huang, Jianping Fan 0007, Jun Yu 0002, Weidong Han 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Knowledge-Constrained Answer Generation for Open-Ended Video Question AnsweringabstractOpen-ended Video question answering (open-ended VideoQA) aims to understand video content and question semantics to generate the correct answers. Most of the best performing models define the problem as a discriminative task of multi-label classification. In real-world scenarios, however, it is difficult to define a candidate set that includes all possible answers. In this paper, we propose a Knowledge-constrained Generative VideoQA Algorithm (KcGA) with an encoder-decoder pipeline, which enables out-of-domain answer generation through an adaptive external knowledge module and a multi-stream information control mechanism. We use ClipBERT to extract the video-question features, extract framewise object-level external knowledge from a commonsense knowledge base and compute the contextual-aware episode memory units via an attention based GRU to form the external knowledge features, and exploit multi-stream information control mechanism to fuse video-question and external knowledge features such that the semantic complementation and alignment are well achieved. We evaluate our model on two open-ended benchmark datasets to demonstrate that we can effectively and robustly generate high-quality answers without restrictions of training data. Guocheng Niu, Xinyan Xiao, Jian Zhang 0026, Xi Peng 0001, Jun Yu 0002 |
AAAI | 4 |
| 2023 | Position constrained network for 3D human pose estimation
Xiena Dong, Jun Yu 0002, Jian Zhang 0026 |
Multim. Syst. | 3 |
| 2023 | Joint Embedding of Deep Visual and Semantic Features for Medical Image Report GenerationabstractMedical image report generation (MeIRG) aims at generating associated diagnosis descriptions with natural language sentences from medical images, which is essential in the computer-aided diagnosis system. Nevertheless, this task remains challenging in that medical images and linguistic expressions should be understood jointly which however show great discrepancies in the modality. To fill this visual-to-semantic gap, we propose a novel framework that follows the encoder-decoder pipeline. Our framework is characterized by encoding both deep visual and semantic embeddings through a triple-branch network (TriNet) during the encoding phase. The visual attention branch captures attended visual embeddings from medical images with the soft-attention mechanism. The medical report (MeRP) embedding branch predicts semantic report embeddings. The embedding branch of medical subject headings (MeSH) obtains semantic embeddings of related medical tags as complementary information. Then, outputs of these branches are fused and fed into a decoder for the report generation. Experimental results on two benchmark datasets have demonstrated the excellent performance of our method. Related codes are available athttps://github.com/yangyan22/Medical-Report-Generation-TriNet. Jun Yu 0002, Jian Zhang 0026, Weidong Han 0001, Hanliang Jiang, Qingming Huang |
IEEE Trans. Multim. | 3 |
| 2023 | Adaptive Neural Network Control of an Uncertain 2-DOF Helicopter With Unknown Backlash-Like Hysteresis and Output ConstraintsabstractAn adaptive neural network (NN) control is proposed for an unknown two-degree of freedom (2-DOF) helicopter system with unknown backlash-like hysteresis and output constraint in this study. A radial basis function NN is adopted to estimate the unknown dynamics model of the helicopter, adaptive variables are employed to eliminate the effect of unknown backlash-like hysteresis present in the system, and a barrier Lyapunov function is designed to deal with the output constraint. Through the Lyapunov stability analysis, the closed-loop system is proven to be semiglobally and uniformly bounded, and the asymptotic attitude adjustment and tracking of the desired set point and trajectory are achieved. Finally, numerical simulation and experiments on a Quanser's experimental platform verify that the control method is appropriate and effective. Zhijia Zhao 0002, Jian Zhang 0026, Zhijie Liu 0001, Chaoxu Mu, Keum Shik Hong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Semisupervised image classification by mutual learning of multiple self-supervised modelsabstractImage classification has been widely adopted by current social media applications. Compared with fully supervised classification, semisupervised classification attracts more attention because it is commonly observed that category labels are only available for a small portion of images while most images on social media platforms do not have labels. To this end, we propose a two-stage semisupervised learning framework. In the first stage, we train two Self-supervised Models (SSMs). One model is initialized by predicting the rotation angles of pretransformed training images and then further trained by the labeled images. The other model is initialized by making consistent predictions for the transformed images in color, shape, and quality from the same sample image, and then further trained by the labeled images. In the second stage, we fuse the two SSMs through deep mutual learning, which enhances each of the two SSMs with the complementary information provided by the other such that the correct prediction could be shared. Experimental results on CIFAR and Caltech-256 data sets demonstrate the effect of the proposed framework. Jian Zhang 0026, Jun Yu 0002, Jianping Fan 0001 |
Int. J. Intell. Syst. | 1 |
| 2022 | Joint usage of global and local attentions in hourglass network for human pose estimation
Xiena Dong, Jun Yu 0002, Jian Zhang 0026 |
Neurocomputing | 3 |
| 2022 | A contrastive triplet network for automatic chest X-ray reporting
Jun Yu 0002, Hanliang Jiang, Weidong Han 0001, Jian Zhang 0026 |
Neurocomputing | 5 |
| 2022 | Graph and dynamics interpretation in robotic reinforcement learning task
Zonggui Yao, Jun Yu 0002, Jian Zhang 0026, Wei He 0001 |
Inf. Sci. | 3 |
| 2022 | Towards Knowledge-Aware Video Captioning via Transitive Visual Relationship DetectionabstractVideo captioning can be enhanced by incorporating the knowledge, which is usually represented as relationships of objects. However, the previous methods construct only superficial or static object relationships, and often introduce noise into the task through irrelevant common sense or fixed syntax templates. These problems mitigate the model interpretability and lead to the undesirable consequence. To overcome these limitations, we propose to enhance video captioning with deep-level object relationships that are adaptively explored during training. Specifically, we present a Transitive Visual Relationship Detection (TVRD) module in which we estimate the actions of the visual objects, and construct an Object-Action Graph (OAG) to describe the shallow relationship between the objects and actions. Then we bridge the gap between the objects via the actions to transitively infer an Object-Object Graph (OOG) which reflects the deep-level relationship. We further feed the OOG to a graph convolutional network to refine the object representation by deep-level relationships. With the refined representation, we capitalize on an LSTM-based decoder for caption generation. Experimental results on two benchmark datasets: MSVD, MSR-VTT demonstrate that the proposed method achieves state-of-the-art performance. Lastly, we present comprehensive ablation studies as well as visualization of visual relationships to demonstrate the effectiveness and interpretability of our model. Bofeng Wu, Guocheng Niu, Jun Yu 0002, Xinyan Xiao, Jian Zhang 0026, Hua Wu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal MatchingabstractThis paper proposes an approach to Dense Video Captioning (DVC) without pairwise event-sentence annotation. First, we adopt the knowledge distilled from relevant and well solved tasks to generate high-quality event proposals. Then we incorporate contrastive loss and cycle-consistency loss typically applied to cross-modal retrieval tasks to build semantic matching between the proposals and sentences, which are eventually used to train the caption generation module. In addition, the parameters of matching module are initialized via pre-training based on annotated images to improve the matching performance. Extensive experiments on ActivityNet-Caption dataset reveal the significance of distillation-based event proposal generation and cross-modal retrieval-based semantic matching to weakly supervised DVC, and demonstrate the superiority of our method to existing state-of-the-art methods. Bofeng Wu, Guocheng Niu, Jun Yu 0002, Xinyan Xiao, Jian Zhang 0026, Hua Wu 0003 |
IJCAI | 5 |
| 2021 | Multi-features guided robust visual tracking
Yun Liang 0003, Jian Zhang 0026, Mei-hua Wang, Chen Lin 0001 |
Multim. Tools Appl. | 2 |
| 2021 | Vector of Locally and Adaptively Aggregated Descriptors for Image Feature Representation
Jian Zhang 0026, Yunyin Cao |
Pattern Recognit. | 1 |
| 2021 | SPRNet: Single-Pixel Reconstruction for One-Stage Instance SegmentationabstractObject instance segmentation is one of the most fundamental but challenging tasks in computer vision, and it requires the pixel-level image understanding. Most existing approaches address this problem by adding a mask prediction branch to a two-stage object detector with the region proposal network (RPN). Although producing good segmentation results, the efficiency of these two-stage approaches is far from satisfactory, restricting their applicability in practice. In this article, we propose a one-stage framework, single-pixel reconstruction net (SPRNet), which performs efficient instance segmentation by introducing a single-pixel reconstruction (SPR) branch to off-the-shelf one-stage detectors. The added SPR branch reconstructs the pixel-level mask from every single pixel in the convolution feature map directly. Using the same ResNet-50 backbone, SPRNet achieves comparable mask AP with Mask R-CNN at a higher inference speed and gains all-round improvements on box AP at every scale compared with RetinaNet. Jun Yu 0002, Jinghan Yao, Jian Zhang 0026, Zhou Yu 0001, Dacheng Tao |
IEEE Trans. Cybern. | 3 |
| 2020 | Combining active learning and local patch alignment for data-driven facial animation with fine-grained local detail
Jian Zhang 0026, Guihua Liao |
Neurocomputing | 1 |
| 2020 | Spatial Pyramid-Enhanced NetVLAD With Weighted Triplet Loss for Place RecognitionabstractWe propose an end-to-end place recognition model based on a novel deep neural network. First, we propose to exploit the spatial pyramid structure of the images to enhance the vector of locally aggregated descriptors (VLAD) such that the enhanced VLAD features can reflect the structural information of the images. To encode this feature extraction into the deep learning method, we build a spatial pyramid-enhanced VLAD (SPE-VLAD) layer. Next, we impose weight constraints on the terms of the traditional triplet loss (T-loss) function such that the weighted T-loss (WT-loss) function avoids the suboptimal convergence of the learning process. The loss function can work well under weakly supervised scenarios in that it determines the semantically positive and negative samples of each query through not only the GPS tags but also the Euclidean distance between the image representations. The SPE-VLAD layer and the WT-loss layer are integrated with the VGG-16 network or ResNet-18 network to form a novel end-to-end deep neural network that can be easily trained via the standard backpropagation method. We conduct experiments on three benchmark data sets, and the results demonstrate that the proposed model defeats the state-of-the-art deep learning approaches applied to place recognition. Jun Yu 0002, Jian Zhang 0026, Qingming Huang, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Automatic image enhancement by learning adaptive patch selectionabstractToday, digital cameras are widely used in taking photos. However, some photos lack detail and need enhancement. Many existing image enhancement algorithms are patch based and the patch size is always fixed throughout the image. Users must tune the patch size to obtain the appropriate enhancement. In this study, we propose an automatic image enhancement method based on adaptive patch selection using both dark and bright channels. The double channels enhance images with various exposure problems. The patch size used for channel extraction is selected automatically by thresholding a contrast feature, which is learned systematically from a set of natural images crawled from the web. Our proposed method can automatically enhance foggy or under-exposed/backlit images without any user interaction. Experimental results demonstrate that our method can provide a significant improvement in existing patch-based image enhancement algorithms. Jian Zhang 0026 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2019 | Robust visual tracking via identifying multi-scale patches
Yun Liang 0003, Ke Li 0005, Jian Zhang 0026, Meihua Wang, Chen Lin 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Compressed sensing for image reconstruction via back-off and rectification of greedy algorithm
Qingyong Deng, Hongqing Zeng, Jian Zhang 0026, Shujuan Tian, Jiasheng Cao, Zhetao Li, Anfeng Liu |
Signal Process. | 3 |
| 2019 | Multimodal Face-Pose Estimation With Multitask Manifold Deep LearningabstractFace-pose estimation aims at estimating the gazing direction with two-dimensional face images. It gives important communicative information and visual saliency. However, it is challenging because of lights, background, face orientations, and appearance visibility. Therefore, a descriptive representation of face images and mapping it to poses are critical. In this paper, we use multimodal data and propose a novel face-pose estimation framework named multitask manifold deep learning ($\text{M}^2\text{DL}$). It is based on feature extraction with improved convolutional neural networks (CNNs) and multimodal mapping relationship with multitask learning. In the proposed CNNs, manifold regularized convolutional layers learn the relationship between outputs of neurons in a low-rank space. Besides, in the proposed mapping relationship learning method, different modals of face representations are naturally combined by applying multitask learning with incoherent sparse and low-rank learning with a least-squares loss. Experimental results on three challenging benchmark datasets demonstrate the performance of$\text{M}^2\text{DL}$. Jun Yu 0002, Jian Zhang 0026, Xiongnan Jin, Kyong-Ho Lee |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Multi-modal feature fusion for geographic image annotation
Ke Li 0005, Changqing Zou, Shuhui Bu, Yun Liang 0003, Jian Zhang 0026, Minglun Gong |
Pattern Recognit. | 5 |
| 2018 | A semi-supervised framework for topology preserving performance-driven facial animation
Jian Zhang 0026, Yun Liang 0003 |
Signal Process. | 1 |
| 2018 | Local Deep-Feature Alignment for Unsupervised Dimension ReductionabstractThis paper presents an unsupervised deep-learning framework named Local Deep-Feature Alignment (LDFA) for dimension reduction. We construct neighbourhood for each data sample and learn a local Stacked Contractive Auto-encoder (SCAE) from the neighbourhood to extract the local deep features. Next, we exploit an affine transformation to align the local deep features of each neighbourhood with the global features. Moreover, we derive an approach from LDFA to map explicitly a new data sample into the learned low-dimensional subspace. The advantage of the LDFA method is that it learns both local and global characteristics of the data sample set: the local SCAEs capture local characteristics contained in the data set, while the global alignment procedures encode the interdependencies between neighbourhoods into the final lowdimensional feature representations. Experimental results on data visualization, clustering and classification show that the LDFA method is competitive with several well-known dimension reduction techniques, and exploiting locality in deep learning is a research topic worth further exploring. Jian Zhang 0026, Jun Yu 0002, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2017 | Learning 3D faces from 2D images via Stacked Contractive Autoencoder
Jian Zhang 0026, Ke Li 0005, Yun Liang 0003 |
Neurocomputing | 1 |
| 2017 | Graph-based clustering and ranking for diversified image search
Yan Yan 0006, Gaowen Liu, Sen Wang 0001, Jian Zhang 0026, Kai Zheng 0001 |
Multim. Syst. | 4 |
| 2016 | Data-driven facial animation via hypergraph learningabstractData-driven facial animation has attracted much attention in recent years. Existing facial animation methods may not preserve the topology structure, and cannot achieve a natural face. This paper proposes a new data-driven facial animation method based on hypergraph learning. It drives a neutral face to a certain expression face. This paper assumes that neutral face has similar topology with the expression face, we compute the alignment laplacian matrix using hypergraph learning. To get a natural face, we add a constraint item which is consisted of a set of motion data. Experiment results demonstrate that our method can achieve a natural expression face. And the results show the superiority over the state-of-art. Jun Yu 0002, Fei Gao 0006, Jian Zhang 0026 |
SMC | 4 |
| 2016 | Data-driven facial animation via semi-supervised local patch alignment
Jian Zhang 0026, Jun Yu 0002, Jane You, Dapeng Tao, Jun Cheng 0002 |
Pattern Recognit. | 1 |
| 2015 | Monocular face reconstruction with global and local shape constraints
Jian Zhang 0026, Dapeng Tao, Xiangjuan Bian, Xiaosi Zhan |
Neurocomputing | 1 |
| 2015 | l2, 1 Norm regularized fisher criterion for optimal feature selection
Jian Zhang 0026, Jun Yu 0002, Jian Wan 0001 |
Neurocomputing | 1 |