EDBT 2026 Demo / reviewers in the wild / expert
Lingxiang Wu
dblp:116/3149
· DBLP profile ↗
17ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SQL-Checker: Error Detection and Labeling for Text-to-SQL with Interpretability AnalysisabstractText-to-SQL technology converts natural language queries into SQL statements for database retrieval. Recent advances in large language models (LLMs) have improved Text-to-SQL performance, but generated SQL often contains semantic or syntax errors that degrade user experience and system stability. Existing SQL error detection methods are costly, lack interpretability, and do not support error labeling. To overcome these issues, we propose SQL-Checker a specialized model for Text-to-SQL error detection. We first analyze common error factors in Text-to-SQL, and we design a novel data synthesis framework based on these error factors. This framework simulates error factors to construct a basic error SQL data, and then using an error analysis template to distill high-quality SQL error analysis data from large-scale models. For complex errors, a self-guided iterative distillation strategy further enhances data quality. SQL-Checker is then trained on this distilled dataset. Additionally, we refine SQL error labeling and, integrate error label recognition into the detection task, enabling macro-level cause analysis. Experiments show SQL-Checker achieves state-of-the-art results on multiple error detection datasets. Incorporating SQL-Checker into the Text-to-SQL pipeline also improves execution accuracy. Lingxiang Wu, Xuepeng Wang, Xueming Tang, Jinqiao Wang |
WWW | 3 |
| 2026 | AnyDesign: Versatile area fashion editing via mask-free diffusion
Yunfang Niu, Dong Yi, Lingxiang Wu, Jinqiao Wang |
Neural Networks | 3 |
| 2026 | Adversarial discriminant attack on text-to-image diffusion models
Hanxiao Wu, Shengwu Xiong 0001, Dong Yi, Lingxiang Wu, Jianqing Zhu, Guibo Zhu, Jinqiao Wang |
Neural Networks | 4 |
| 2025 | Enhancing Chain of Thought Prompting in Large Language Models via Reasoning PatternsabstractChain of Thought (CoT) prompting can encourage language models to engage in multi-step logical reasoning. The quality of the provided demonstrations significantly influences the success of downstream inference tasks. Current unsupervised CoT methods primarily select examples based on the semantics of the questions, which can introduce noise and lack interpretability. In this paper, we propose leveraging reasoning patterns to enhance CoT prompting effectiveness. Reasoning patterns represent the process by which language models arrive at their final results. By utilizing prior knowledge and prompt-based methods from large models, we first construct task-specific pattern sets. We then select diverse demonstrations based on different reasoning patterns. This approach not only mitigates the impact of noise but also provides explicit interpretability to help us understand the mechanisms of CoT. Extensive experiments demonstrate that our method is more robust and consistently leads to improvements across various reasoning tasks. Xuepeng Wang, Lingxiang Wu, Jinqiao Wang |
AAAI | 3 |
| 2025 | SCANER: robust and sensitive identification of malignant cells from the scRNA-seq profiled tumor ecosystemabstractSingle-cell RNA sequencing (scRNA-seq) has enabled the dissection of complex tumor ecosystems. Recognition of malignant cells as an essential step has a profound impact on downstream interpretation. However, most existing computational strategies are based on prior knowledge of canonical cell-type markers. We have developed a marker-free approach, the Seed-Cluster based Approach for NEoplastic cells Recognition (SCANER), to identify malignant cells based on significant gene expression variations caused by genomic instability. Upon analyzing different cancer types, SCANER achieved superior accuracy and robustness in identifying malignant cells, effectively addressing dropout events and tumor purity variations. Besides, SCANER can significantly detect copy number variations (CNVs) in malignant cells compared to nonmalignant cells, which is further confirmed through the paired whole exome sequencing data. In conclusion, SCANER has the potential to facilitate the biological exploration of the tumor ecosystem by accurately identifying malignant cells and it is applicable across various solid cancer types regardless of prior knowledge. SCANER is available at https://github.com/woolingxiang/SCANER. Quanzhong Liu, Mengyan Zhu, Miao Yu 0010, Kening Li, Lingxiang Wu, Qianghu Wang |
Briefings Bioinform. | 11 |
| 2025 | Enhancing Visual Aligning and Grounding for Aerial Vision-and-Dialog NavigationabstractVision-and-Language Navigation tasks require an agent to navigate to a destination following natural language instructions. We focus on a challenging VLN dataset, Aerial Vision-and-Dialog Navigation, which encompasses a diverse array of environments and includes an additional altitude variable. Significant spatial and scale variations in the aerial agent's view make destination visual grounding a crucial capability for the navigation task. However, the existing frameworks not only have insufficient attention to the vision model, but also lack the correlations between visual and textual modalities. To address this, we propose a model that aligns destination visual images with navigation instructions, featuring three innovative components. Firstly, we propose a multi-stage pre-training pipeline that enhances the model's ability to associate language instructions with top-view images of destinations. Secondly, trajectories are augmented elastically to simulate the noise of the controlling process. Finally, a polygon regression loss function is introduced for rotated object detection, which significantly enhances the accuracy of altitude and orientation estimation. Experiments demonstrate the effectiveness of our approach, which achieves state-of-the-art advancements with improvements of 2.9% in the val unseen dataset and 3.0% in test unseen dataset in success weighted by path length. Guanhui Qiao, Dong Yi, Lingxiang Wu, Hanxiao Wu, Jinqiao Wang |
IEEE Signal Process. Lett. | 3 |
| 2024 | Enhancing Text-to-SQL Capabilities of Large Language Models via Domain Database Knowledge InjectionabstractText-to-SQL is a subtask in semantic parsing that has seen rapid progress with the evolution of Large Language Models (LLMs). However, LLMs face challenges due to hallucination issues and a lack of domain-specific database knowledge(such as table schema and cell values). As a result, they can make errors in generating table names, columns, and matching values to the correct columns in SQL statements. This paper introduces a method of knowledge injection to enhance LLMs’ ability to understand schema contents by incorporating prior knowledge. This approach improves their performance in Text-to-SQL tasks. Experimental results show that pre-training LLMs on domain-specific database knowledge and fine-tuning them on downstream Text-to-SQL tasks significantly improves the Execution Match (EX) and Exact Match (EM) metrics across various models. This effectively reduces errors in generating column names and matching values to the columns. Furthermore, the knowledge-injected models can be applied to many downstream Text-to-SQL tasks, demonstrating the generalizability of the approach presented in this paper. Lingxiang Wu, Xuepeng Wang, Xueming Tang, Jinqiao Wang |
ECAI | 3 |
| 2024 | PFDM: Parser-Free Virtual Try-On via Diffusion ModelabstractVirtual try-on can significantly improve the garment shopping experiences in both online and in-store scenarios, attracting broad interest in computer vision. However, to achieve high-fidelity try-on performance, most state-of-the-art methods still rely on accurate segmentation masks, which are often produced by near-perfect parsers or manual labeling. To overcome the bottleneck, we propose a parser-free virtual try-on method based on the diffusion model (PFDM). Given two images, PFDM can "wear" garments on the target person seamlessly by implicitly warping without any other information. To learn the model effectively, we synthesize many pseudo-images and construct sample pairs by wearing various garments on persons. Supervised by the large-scale expanded dataset, we fuse the person and garment features using a proposed Garment Fusion Attention (GFA) mechanism. Experiments demonstrate that our proposed PFDM can successfully handle complex cases, synthesize high-fidelity images, and outperform both state-of-the-art parser-free and parser-based models. Yunfang Niu, Dong Yi, Lingxiang Wu, Zhiwei Liu 0004, Pengxiang Cai, Jinqiao Wang |
ICASSP | 3 |
| 2024 | Completing Saliency from Details
Jin Zhang 0021, Lingxiang Wu, Renwei Dian, Yiheng Yao, Shihao Huang, Yang Yang 0074, Ruiheng Zhang 0001 |
PRCV (13) | 3 |
| 2024 | Multi-Model Style-Aware Diffusion Learning for Semantic Image SynthesisabstractSemantic image synthesis aims to generate images from given semantic layouts, which is a challenging task that requires training models to capture the relationship between layouts and images. Previous works are usually based on Generative Adversarial Networks (GAN) or autoregressive (AR) models. However, the GAN model's training process is unstable, and the AR model’s performance is seriously affected by the independent image encoder and the unidirectional generation bias. Due to the above limitations, these methods tend to synthesize unrealistic, poorly aligned images and only consider single-style image generation. In this paper, we propose a Multi-model Style-aware Diffusion Learning (MSDL) framework for semantic image synthesis, including a training module and a sampling module. In the training module, a layout-to-image model is introduced to transfer the learned knowledge from a model pretrained with massive weak correlated text-image pairs data, making the training process more efficient. In the sampling module, we designed a map-guidance technique and creatively designed a multi-model style-guidance strategy for creating images in multiple styles, e.g., oil painting, Disney Cartoon, and pixel style. We evaluate our method on Cityscapes, ADE20K, and COCO-Stuff, making visual comparisons and computing with multiple metrics such as FID, LPIPS, etc. Experimental results demonstrate that our model is highly competitive, especially in terms of fidelity and diversity. Yunfang Niu, Lingxiang Wu, Yousong Zhu, Guibo Zhu, Jinqiao Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | TaiSu: A 166M Large-scale High-Quality Dataset for Chinese Vision-Language Pre-trainingabstractVision-Language Pre-training (VLP) has been shown to be an efficient method to improve the performance of models on different vision-and-language downstream tasks. Substantial studies have shown that neural networks may be able to learn some general rules about language and visual concepts from a large-scale weakly labeled image-text dataset. However, most of the public cross-modal datasets that contain more than 100M image-text pairs are in English; there is a lack of available large-scale and high-quality Chinese VLP datasets. In this work, we propose a new framework for automatic dataset acquisition and cleaning with which we construct a new large-scale and high-quality cross-modal dataset named as TaiSu, containing 166 million images and 219 million Chinese captions. Compared with the recently released Wukong dataset, our dataset is achieved with much stricter restrictions on the semantic correlation of image-text pairs. We also propose to combine texts collected from the web with texts generated by a pre-trained image-captioning model. To the best of our knowledge, TaiSu is currently the largest publicly accessible Chinese cross-modal dataset. Furthermore, we test our dataset on several vision-language downstream tasks. TaiSu outperforms BriVL by a large margin on the zero-shot image-text retrieval task and zero-shot image classification task. TaiSu also shows better performance than Wukong on the image-retrieval task without using image augmentation for training. Results demonstrate that TaiSu can serve as a promising VLP dataset, both for understanding and generative tasks. More information can be referred to https://github.com/ksOAn6g5/TaiSu. Guibo Zhu, Qi Song 0003, Guojing Ge, Guanhui Qiao, Ru Peng, Lingxiang Wu, Jinqiao Wang |
NeurIPS | 9 |
| 2021 | Noise Augmented Double-Stream Graph Convolutional Networks for Image CaptioningabstractImage captioning, aiming at generating natural sentences to describe image contents, has received significant attention with remarkable improvements in recent advances. The problem nevertheless is not trivial for cross-modal training due to the two challenges: 1) image detectors often consider only salient areas in an image and seldom explore the rich background context; 2) the language model is highly vulnerable to small but intentional perturbation attacks. To alleviate these issues, we propose the Noise Augmented Double-stream Graph Convolutional Networks (NADGCN) that novelly exploits the additional background context and enhances the generalization of the language model. Technically, NADGCN capitalizes on grid-stream GCN as a supplementary to the region stream, following the recipe that a rescaled grid graph can encode the relationship across grid areas over the full image rather than salient areas only. Moreover, we devise a noise module and integrate into the double-stream GCN to augment the capability of the basic generator. Such noise module introduces adaptive noise into the Recurrent Neural Networks (RNN) and is learnt through regarding the module as an agent with a stochastic Gaussian policy in Reinforcement Learning (RL). Extensive experiments on MSCOCO validate the design of the grid-stream GCN and the noise agent, and our generator outperforms the comparative baselines clearly. Lingxiang Wu, Min Xu 0001, Lei Sang 0001, Ting Yao 0003, Tao Mei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Multi-camera multi-player tracking with deep player identification in sports video
Ruiheng Zhang 0001, Lingxiang Wu, Yukun Yang 0003, Wanneng Wu, Yueqiang Chen, Min Xu 0001 |
Pattern Recognit. | 2 |
| 2020 | Recall What You See Continually Using GridLSTM in Image CaptioningabstractThe goal of image captioning is to automatically describe an image with a sentence, and the task has attracted research attention from both the computer vision and natural-language processing research communities. The existing encoder-decoder model and its variants, which are the most popular models for image captioning, use the image features in three ways: first, they inject the encoded image features into the decoder only once at the initial step, which does not enable the rich image content to be explored sufficiently while gradually generating a text caption; second, they concatenate the encoded image features with text as extra inputs at every step, which introduces unnecessary noise; and, third, they using an attention mechanism, which increases the computational complexity due to the introduction of extra neural nets to identify the attention regions. Different from the existing methods, in this paper, we propose a novel network, Recall Network, for generating captions that are consistent with the images. The recall network selectively involves the visual features by using a GridLSTM and, thus, is able to recall image contents while generating each word. By importing the visual information as the latent memory along the depth dimension LSTM, the decoder is able to admit the visual features dynamically through the inherent LSTM structure without adding any extra neural nets or parameters. The Recall Network efficiently prevents the decoder from deviating from the original image content. To verify the efficiency of our model, we conducted exhaustive experiments on full and dense image captioning. The experimental results clearly demonstrate that our recall network outperforms the conventional encoder-decoder model by a large margin and that it performs comparably to the state-of-the-art methods. Lingxiang Wu, Min Xu 0001, Jinqiao Wang, Stuart W. Perry |
IEEE Trans. Multim. | 1 |
| 2020 | Image to Modern Chinese Poetry Creation via a Constrained Topic-aware ModelabstractArtificial creativity has attracted increasing research attention in the field of multimedia and artificial intelligence. Despite the promising work on poetry/painting/music generation, creating modern Chinese poetry from images, which can significantly enrich the functionality of photo-sharing platforms, has rarely been explored. Moreover, existing generation models cannot tackle three challenges in this task: (1) Maintaining semantic consistency between images and poems; (2) preventing topic drift in the generation; (3) avoidance of certain words appearing frequently. These three points are even common challenges in other sequence generation tasks. In this article, we propose a Constrained Topic-aware Model (CTAM) to create modern Chinese poetries from images regarding the challenges above. Without image-poetry paired dataset, we construct a visual semantic vector to embed visual contents via image captions. For the topic-drift problem, we propose a topic-aware poetry generation model. Additionally, we design an Anti-frequency Decoding (AFD) scheme to constrain high-frequency characters in the generation. Experimental results show that our model achieves promising performance and is effective in poetry’s readability and semantic consistency. Lingxiang Wu, Min Xu 0001, Shengsheng Qian, Jianwei Cui 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Appearance features in Encoding Color Space for visual surveillance
Lingxiang Wu, Min Xu 0001, Guibo Zhu, Jinqiao Wang, Tianrong Rao |
Neurocomputing | 1 |
| 2016 | Person re-identification via rich color-gradient featureabstractPerson re-identification refers to match the same pedestrian across disjoint views in non-overlapping camera networks. Lots of local and global features in the literature are put forward to solve the matching problem, where color feature is robust to viewpoint variance and gradient feature provides a rich representation robust to illumination change. However, how to effectively combine the color and gradient features is an open problem. In this paper, to effectively leverage the color-gradient property in multiple color spaces, we propose a novel Second Order Histogram feature (SOH) for person reidentification in large surveillance dataset. Firstly, we utilize discrete encoding to transform commonly used color space into Encoding Color Space (ECS), and calculate the statistical gradient features on each color channel. Then, a second order statistical distribution is calculated on each cell map with a spatial partition. In this way, the proposed SOH feature effectively leverages the statistical property of gradient and color as well as reduces the redundant information. Finally, a metric learned by KISSME [1] with Mahalanobis distance is used for person matching. Experimental results on three public datasets, VIPeR, CAVIAR and CUHK01, show the promise of the proposed approach. Lingxiang Wu, Jinqiao Wang, Guibo Zhu, Min Xu 0001, Hanqing Lu |
ICME | 1 |