VLDB 2026 Research / reviewers in the wild / expert
Jinxia Zhang
dblp:89/1778
· DBLP profile ↗
28ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality CheckabstractThe pursuit of real-time agentic interaction has driven interest in Diffusion-based Large Language Models (dLLMs) as alternatives to autoregressive backbones, promising to break the sequential latency bottleneck.However, does such efficiency gains translate into effective agentic behavior?In this work, we present a comprehensive evaluation of dLLMs (e.g., LLaDA, Dream) across two distinct agentic paradigms: Embodied Agents (requiring longhorizon planning) and Tool-Calling Agents (requiring precise formatting).Contrary to the efficiency hype, our results on Agentboard and BFCL reveal a "bitter lesson": current dLLMs fail to serve as reliable agentic backbones, frequently leading to systematic failure.(1) In Embodied settings, dLLMs suffer repeated attempts, failing to branch under temporal feedback.(2) In Tool-Calling settings, dLLMs fail to maintain symbolic precision (e.g.strict JSON schemas) under diffusion noise.To assess the potential of dLLMs in agentic workflows, we introduce DiffuAgent, a multi-agent evaluation framework that integrates dLLMs as plug-and-play cognitive cores.Our analysis shows that dLLMs are effective in non-causal roles (e.g., memory summarization and tool selection) but require the incorporation of causal, precise, and logically grounded reasoning mechanisms into the denoising process to be viable for agentic tasks. Qingyu Lu 0001, Liang Ding 0006, Kan-Jian Zhang, Jinxia Zhang, Dacheng Tao |
ACL (1) | 4 |
| 2026 | Partitioned observation network for camouflaged object detection
Jinxia Zhang, Yin Yuan, Xuwen Zhu, Kaihua Zhang 0001 |
Pattern Recognit. | 1 |
| 2026 | Attention-Enhanced Diffusion With LLM-Driven Prompts for Controllable Defect Generation in Photovoltaic CellsabstractSolar power is a vital clean energy source, and defects in photovoltaic cells critically impact power generation efficiency. While deep learning has advanced automated defect recognition, models face challenges due to scarcity and imbalance in defect samples. To overcome these challenges, a novel approach is proposed that exploits controllable image generation to produce diverse and realistic defect samples, while maintaining precise control over key attributes. First, to overcome the difficulty of generating accurate defect descriptions, a large language model is leveraged to automatically generate detailed text prompts. Next, to enable the model to learn domain-specific knowledge in photovoltaic cell defect recognition, the diffusion model is fine-tuned by DreamBooth. To further enhance the editability of these defects, a controllable mask, accompanied by class name, is proposed to allow precise control over the type, size, and position of the defects. Finally, an attention enhancement module is proposed to improve both the precision and flexibility of defect generation, ensuring a more accurate and realistic simulation of defects in photovoltaic images. The quality of the generated images is assessed using diverse metrics, further demonstrating that the generated images significantly improve defect recognition performance. Zechao Zhan, Jinxia Zhang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented SamplesabstractExisting Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their effectiveness in capturing coarse-grained semantic correlations. However, their capability for fine-grained understanding, which is critical for many nuanced vision-language applications, remains limited. Prevailing VLP models often overlook the intricate distinctions in expressing different modal features and typically depend on the similarity of holistic features for cross-modal interactions. Moreover, these models directly align and integrate features from different modalities, focusing more on coarse-grained general representations, thus failing to capture the nuanced differences necessary for tasks demanding a more detailed perception. In response to these limitations, we introduce Negative Augmented Samples(NAS), a refined vision-language pretraining model that innovatively incorporates NAS to specifically address the challenge of fine-grained understanding. NAS utilizes a Visual Dictionary(VD) as a semantic bridge between visual and linguistic domains. Additionally, it employs a Negative Visual Augmentation(NVA) method based on the VD to generate challenging negative image samples. These samples deviate from positive samples exclusively at the token level, thereby necessitating that the model discerns the subtle disparities between positive and negative samples with greater precision. Comprehensive experiments validate the efficacy of NAS components and underscore its potential to enhance fine-grained vision-language comprehension. Yeyuan Wang, Dehong Gao, Lei Yi, Linbo Jin, Jinxia Zhang, Libin Yang, Xiaoyan Cai |
AAAI | 5 |
| 2025 | MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation EvaluatorsabstractLarge Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment, providing both scores and fine-grained feedback. Although approaches such as GEMBA-MQM have shown state-of-the-art performance on reference-free evaluation, the predicted errors do not align well with those annotated by human, limiting their interpretability as feedback signals. To enhance the quality of error annotations predicted by LLM evaluators, we introduce a universal and training-free framework, MQM-APE, based on the idea of filtering out non-impactful errors by Automatically Post-Editing (APE) the original translation based on each error, leaving only those errors that contribute to quality improvement. Specifically, we prompt the LLM to act as 1) evaluator to provide error annotations, 2) post-editor to determine whether errors impact quality improvement and 3) pairwise quality verifier as the error filter. Experiments show that our approach consistently improves both the reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages. Orthogonal to trained approaches, MQM-APE complements translation-specific evaluators such as Tower, highlighting its broad applicability. Further analysis confirms the effectiveness of each module and offers valuable insights into evaluator design and LLMs selection. Qingyu Lu 0001, Liang Ding 0006, Kan-Jian Zhang, Jinxia Zhang, Dacheng Tao |
COLING | 4 |
| 2025 | Shifting Spotlight for Co-supervision: A Simple yet Efficient Single-branch Network to See Through CamouflageabstractCamouflaged object detection (COD) remains a challenging task in computer vision. Existing methods often resort to additional branches for edge supervision, incurring substantial computational costs. To address this, we propose the Co-Supervised Spotlight Shifting Network (CS3Net), a compact single-branch framework inspired by how shifting light source exposes camouflage. Our spotlight shifting strategy replaces multi-branch designs by generating supervisory signals that highlight boundary cues. Within CS3Net, a Projection Aware Attention (PAA) module is devised to strengthen feature extraction, while the Extended Neighbor Connection Decoder (ENCD) enhances final predictions. Extensive experiments on public datasets demonstrate that CS3Net not only achieves superior performance, but also reduces Multiply-Accumulate operations (MACs) by 32.13% compared to state-of-the-art COD methods, striking an optimal balance between efficiency and effectiveness. Jinxia Zhang, Yin Yuan, Zechao Zhan |
ICASSP | 2 |
| 2025 | FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-trainingabstractLarge-scale Vision-Language Pre-training (VLP) has demonstrated remarkable success in the general domain. However, in the fashion domain, items are distinguished by fine-grained attributes such as texture and material, which are crucial for tasks such as retrieval. Existing models often fail to take advantage of these fine-grained attributes from both text and image modalities. To address the above issue, we propose a novel approach for the fashion domain, Fine-grained Attributes Enhanced VLP (FashionFAE), which focuses on the detailed characteristics of the fashion data. An attribute-emphasized text prediction task is proposed to predict fine-grained attributes of the items. This forces the model to focus on the salient attributes from the text modality. In addition, a novel attribute-promoted image reconstruction task is proposed, which further enhances the fine-grained ability of the model by leveraging the representative attributes from the image modality. Extensive experiments show that FashionFAE outperforms State-Of-The-Art (SOTA) methods, achieving 2.9% and 5.2% improvements in retrieval on sub-test set and full test set, respectively, and an average improvement of 1.6% in recognition tasks. Dehong Gao, Jinxia Zhang, Zechao Zhan |
ICASSP | 3 |
| 2025 | CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language ModelsabstractThe impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks. However, current MLLM often struggles to effectively address fine-grained multi-modal challenges. We argue that this limitation is closely linked to the models’ visual grounding capabilities. The restricted spatial awareness and perceptual acuity of visual encoders frequently lead to interference from irrelevant background information in images, causing the models to overlook subtle but crucial details. As a result, achieving fine-grained regional visual comprehension becomes difficult. In this paper, we break down multi-modal understanding into two stages, from Coarse to Fine (CoF). In the first stage, we prompt the MLLM to locate the approximate area of the answer. In the second stage, we further enhance the model’s focus on relevant areas within the image through visual prompt engineering, adjusting attention weights of pertinent regions. This, in turn, improves both visual grounding and overall performance in downstream tasks. Our experiments show that this approach significantly boosts the performance of baseline models, demonstrating notable generalization and effectiveness. Our CoF approach is available online at https://github.com/Gavin001201/CoF. Yeyuan Wang, Dehong Gao, Rujiao Long, Lei Yi, Xiaoyan Cai, Libin Yang, Jinxia Zhang, Shanqing Yu, Qi Xuan 0001 |
ICASSP | 8 |
| 2025 | MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced DiffusionabstractText-guided image editing model has achieved great success in general domain. However, directly applying these models to the fashion domain may encounter two issues: (1) Inaccurate localization of editing region; (2) Weak editing magnitude. To address these issues, the MADiff model is proposed. Specifically, to more accurately identify editing region, the MaskNet is proposed, in which the foreground region, densepose and mask prompts from large language model are fed into a lightweight UNet to predict the mask for editing region. To strengthen the editing magnitude, the Attention-Enhanced Diffusion Model is proposed, where the noise map, attention map, and the mask from MaskNet are fed into the proposed Attention Processor to produce a refined noise map. By integrating the refined noise map into the diffusion model, the edited image can better align with the target prompt. Given the absence of benchmarks in fashion image editing, we constructed a dataset named FashionE, comprising 28390 image-text pairs in the training set, and 2639 image-text pairs for four types of fashion tasks in the evaluation set. Extensive experiments on Fashion-E demonstrate that our proposed method can accurately predict the mask of editing region and significantly enhance editing magnitude in fashion image editing compared to the state-of-the-art methods. Zechao Zhan, Dehong Gao, Jinxia Zhang |
ICASSP | 3 |
| 2025 | Fault diagnosis method of mining vibrating screen mesh based on an improved algorithmabstractArtificial intelligence fault diagnosis technology based on machine vision, due to its low cost and high efficiency, has become an indispensable part of production processes across various industries. Compared to traditional fault diagnosis methods, artificial intelligence diagnosis of common mechanical failures, such as ‘clogging’, ‘wear’, and ‘breakage’ in vibrating screen meshes within the mining screening sector, improves detection efficiency, accuracy, and sustainability. Since small target faults in large screening areas are challenging to detect through manual diagnosis, it reduces screening efficiency and shorter equipment lifespan, negatively impacting mining enterprises' safe and efficient production. A fault diagnosis model with a better speed-precision trade-off is proposed to improve detection precision based on the You Only Look Once version 5 single-stage object detection algorithm. This model is optimized in feature extraction and fusion by integrating autocode masking, re-parameterization, and omni-dimensional attention. The model's performance is primarily evaluated using precision, recall, balanced score, and mean average precision. The improved algorithm achieves a precision of 97.2%, a recall of 93.3%, a balanced score of 95.21%, and a mean average precision of 97.0%. Experimental results demonstrate that the improved algorithm increases the mean average precision by 3.1% compared to the original model. The results show that the improved algorithm is more effective than the original in fault diagnosis, with enhanced screen mesh detection precision. Thus, it ensures production safety and stable screening efficiency. Moreover, the proposed algorithm provides a reference for advancing intelligent and efficient fault diagnosis technology in the mining screening field. Fusheng Niu, Jinxia Zhang, Zhiheng Nie, Guang Song, Xiongsheng Zhu |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | DICO: Distance-weighted Contrast and Instance Correlation for salient object ranking
Jinxia Zhang, Xinchao Zhu, Haikun Wei, Shixiong Fang, Kan-Jian Zhang |
Neurocomputing | 1 |
| 2025 | Referring Solar Cell Defect Segmentation in Electroluminescence ImagesabstractIn the photovoltaic (PV) power generation field, accurately identifying solar cell defects based electroluminescence (EL) images is essential for maintaining high efficiency for PV power plants. Current solar cell defect segmentation methods typically segment all defects in the EL image uniformly, making it difficult to precisely identify specific defects according to maintenance needs. This limitation hinders personalized defect detection for smart operation and maintenance of PV power plants. To solve this problem, a novel task referred to as referring solar cell defect segmentation (RSCDS) is proposed in this article. The goal of the RSCDS task is to precisely segment the specified solar cell defects based on the referring text, tailored to the personalized maintenance requirements of actual PV power plants. Given the lack of relevant datasets, an RSCDS dataset is developed, abbreviated as Ref-EL-defect, comprising 60 000 pairs of defects and corresponding referring texts. The referring text can indicate single defect, multiple defects, or even no defects at all in the EL image, and such multigranularity correspondence enables accurate and personalized segmentation of defects. In addition, a multimodal multigranularity segment network is designed for the RSCDS task. By exploiting the characteristics of the solar cell defects, the multimodal fusion module and multigranularity perception grouping module are proposed to better adapt to the RSCDS task. State-of-the-art (SOTA) referring expression segmentation models designed for natural scene images are transferred to the RSCDS task, and experimental results demonstrate that the proposed method outperforms the SOTA models. Shenghao Dong, Jinxia Zhang, Dehong Gao |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Subjective Quality Assessment of Thermal Infrared ImagesabstractThermal infrared images (TIIs) can be distorted by multiple factors, resulting in noise, low contrast, limited dynamic range, and fuzziness, which greatly impede their usefulness. It is crucial to evaluate the quality of TIIs. Unfortunately, there have been very few attempts to study this problem. In this study, we collected 1,000 authentically distorted TIIs using thermal infrared acquisition equipment and conducted strict subjective experiments to obtain a thermal infrared image quality assessment (IQA) database. Each image’s quality score was obtained under strict scoring rules. Finally, we investigated the feasibility of several no-reference (NR) IQA methods in quality assessment of TIIs. We found that existing NR-IQA methods achieve ordinary performance in such a task, and there is an urgent need to develop a specific IQA methods for TIIs. The findings together with the constructed database are expected to pave the way for the development of more advanced IQA methods for further development of this field. Guanghui Yue 0001, Jinxia Zhang, Zhaofei Xu, Shuigen Wang, Tianwei Zhou, Yuanhao Gong, Wei Zhou 0021 |
ICIP | 3 |
| 2024 | Medical Named Entity Recognition Model Based on Knowledge Graph EnhancementabstractTo improve the recognition ability of clinical named entity recognition (CNER) in a limited number of Chinese electronic medical records, it provides meaningful support for clinical advanced knowledge extraction. In this paper, using CCKS2019 Chinese electronic medical record as an experimental data source, a fusion model enhanced by knowledge graph (KG) is proposed, and the model is applied to specific Chinese CNER tasks. This study consists of three main parts: single-mode model construction and comparison experiment, KG enhancement experiment, and model fusion experiment. The model has achieved good performance in CNER from the results. The accuracy rate, recall rate, and F1 value are 83.825%, 84.705%, and 84.263%, respectively, which is the global optimal, which proves the effectiveness of the model. This provides a good help for further research of medical information. Yonghe Lu, Ruijie Zhao 0004, Xiuxian Wen, Xinyu Tong 0002, Dingcheng Xiang, Jinxia Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2024 | Erratum: Medical Named Entity Recognition Model Based on Knowledge Graph Enhancement
Yonghe Lu, Ruijie Zhao 0004, Xiuxian Wen, Xinyu Tong 0002, Dingcheng Xiang, Jinxia Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2023 | Instance-dimension dual contrastive learning of visual representations
Qingrui Liu, Liantao Wang, Qinxu Wang, Jinxia Zhang |
Mach. Vis. Appl. | 4 |
| 2023 | RGB-D salient object ranking based on depth stack and truth stack for complex indoor scenes
Jingzheng Deng, Jinxia Zhang, Zewen Hu, Liantao Wang, Xinchao Zhu, Yin Yuan |
Pattern Recognit. | 2 |
| 2023 | Automatic Detection of Defective Solar Cells in Electroluminescence Images via Global Similarity and Concatenated Saliency Guided NetworkabstractElectroluminescence imaging becomes a very useful technique to automatically detect defects for solar cells since it can provide high resolution electroluminescence images. However, few methods explicitly consider the visual characteristics of the defects and the noises in solar cells. In this article, a global pairwise similarity and concatenated saliency guided neural network is proposed by fully considering the observed visual characteristics in electroluminescence solar cell images. The proposed network exploits a global pairwise similarity module and a concatenated saliency module to refine the features extracted by the convolutional neural network. The global pairwise similarity module aims to refine the features of an image pixel by modeling long-range dependencies. The concatenated saliency module is exploited to suppress the background and decouple different salient regions to better represent the features of an image. Extensive experiments based on five different baselines, i.e., VGG16, ResNet56, ResNet50, DenseNet40, and GoogleNet, prove that the proposed method significantly outperforms the baseline models and show that both the global similarity module and the concatenated saliency module can help detect defective solar cells in electroluminescence images. Jinxia Zhang, Shixiong Fang, Kan-Jian Zhang, Haikun Wei, Weili Guo |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | ST-CSNN: a novel method for vehicle counting
Liantao Wang, Jinxia Zhang |
Mach. Vis. Appl. | 3 |
| 2021 | Salient Object Detection by Fusing Local and Global ContextsabstractBenefiting from the powerful discriminative feature learning capability of convolutional neural networks (CNNs), deep learning techniques have achieved remarkable performance improvement for the task of salient object detection (SOD) in recent years. However, most existing deep SOD models do not fully exploit informative contextual features, which often leads to suboptimal detection performance in the presence of a cluttered background. This paper presents a context-aware attention module that detects salient objects by simultaneously constructing connections between each image pixel and its local and global contextual pixels. Specifically, each pixel and its neighbors bidirectionally exchange semantic information by computing their correlation coefficients, and this process aggregates contextual attention features both locally and globally. In addition, an attention-guided hierarchical network architecture is designed to capture fine-grained spatial details by transmitting contextual information from deeper to shallower network layers in a top-down manner. Extensive experiments on six public SOD datasets show that our proposed model demonstrates superior SOD performance against most of the current state-of-the-art models under different evaluation metrics. Qinghua Ren, Shijian Lu, Jinxia Zhang |
IEEE Trans. Multim. | 3 |
| 2020 | A Robust Automatic Method for Removing Projective Distortion of Photovoltaic Modules from Close Shot Images
Jinxia Zhang, Kan-Jian Zhang, Haikun Wei |
PRCV (1) | 3 |
| 2019 | Salient object detection via reliable boundary seeds and saliency refinementabstractSalient object detection can identify the most distinctive objects in a scene. In this study, a novel graph‐based approach is proposed to detect a salient object via reliable boundary seeds and saliency refinement. A natural image is firstly mapped to a graph with superpixels as nodes. Saliency information is then diffused over the graph using seeds. For the reason that the boundary nodes may contain salient nodes, it is not appropriate to use all boundary nodes as the background seeds. Therefore, a boundary saliency measurement is proposed to obtain more accurate background seeds. After that, the information of background seeds is diffused by a two‐stage scheme. A background‐based map and a foreground‐based map are generated based on the two‐stage scheme. Furthermore, in order to enhance the detection accuracy, a refinement model is presented to fuse the information of background‐based and foreground‐based maps. Experiments on seven public datasets show the proposed algorithm out‐performs the state‐of‐the‐art salient object detection algorithms. Xiyin Wu, Xiaodi Ma, Jinxia Zhang, Zhong Jin |
IET Comput. Vis. | 3 |
| 2018 | Salient Object Detection Via Deformed Smoothness ConstraintabstractIn recent years, various graph-based salient object detection methods have been successfully proposed. Since existing methods may miss some object regions with low contrast to background, a novel propagation model via deformed smoothness constraint is proposed to address this problem. By regularizing nodes and their neighbors locally, the deformed smoothness constraint is able to prevent erroneous label propagation. Thus, the object regions with low contrast to background can be emerged. Besides, the deformed smoothness constraint is further utilized in a map refinement model, which can suppress the background noises in label propagation result. Experiments on three public datasets show that the proposed method outperforms eleven state-of-the-art salient object detection methods. Xiyin Wu, Xiaodi Ma, Jinxia Zhang, Andong Wang, Zhong Jin |
ICIP | 3 |
| 2018 | Stability analysis of opposite singularity in multilayer perceptrons
Weili Guo, Junsheng Zhao, Jinxia Zhang, Haikun Wei, Aiguo Song, Kan-Jian Zhang |
Neurocomputing | 3 |
| 2017 | A novel graph-based optimization framework for salient object detection
Jinxia Zhang, Krista A. Ehinger, Haikun Wei, Kan-Jian Zhang, Jing-Yu Yang 0001 |
Pattern Recognit. | 1 |
| 2017 | Erratum to: A novel graph-based optimization framework for salient object detection [Pattern Recognition 64C (2017) 39-50]
Jinxia Zhang, Krista A. Ehinger, Haikun Wei, Kan-Jian Zhang, Jing-Yu Yang 0001 |
Pattern Recognit. | 1 |
| 2014 | A prior-based graph for salient object detectionabstractRecently, various graph-based methods have be proposed for salient object detection. These algorithms represent image points and their similarity as nodes and edges in a graph. Although the edge structure and weighting are the heart of these methods, the graph construction has not been studied in detail. In this paper, we exploit image priors, including spatial priors, color priors, and a central bias prior, to construct the graph. We connect nodes which are spatially close in the image, nodes which have similar color features, and the boundary nodes along the borders of the image, while weighting edges according to both their color similarity and spatial proximity. Moreover, we propose a new sine spatial distance instead of the commonly-used Euclidean spatial distance, which better captures the central bias in scenes. Extensive experiments show that our method outperforms thirteen state-of-the-art methods on four different image databases. Jinxia Zhang, Krista A. Ehinger, Jundi Ding, Jing-Yu Yang 0001 |
ICIP | 1 |
| 2014 | Exploiting global rarity, local contrast and central bias for salient region learning
Jinxia Zhang, Jundi Ding |
Neurocomputing | 1 |