VLDB 2026 Research / reviewers in the wild / expert
Juhua Liu
dblp:122/1682
· DBLP profile ↗
38ranked-venue papers
6as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 3 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Achieving >97% on GSM8K: deeply understanding the problems makes LLMs better solvers for math word problems
Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001 |
Frontiers Comput. Sci. | 5 |
| 2026 | MKDS: Multi-source knowledge-driven data synthesis framework for effective domain adaptation of large language models
Qihuang Zhong, Jinzhao Gong, Juhua Liu, Bo Du 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Tool retrieval bridge: Aligning vague instructions with retriever preferences via bridge model
Kunfeng Chen, Luyao Zhuang, Juhua Liu, Bo Du 0001 |
Neural Networks | 4 |
| 2026 | SFA: Scan, Focus, and Amplify toward guidance-aware answering for Video TextVQA
Haibin He 0001, Qihuang Zhong, Juhua Liu, Bo Du 0001, Peng Wang 0076, Jing Zhang 0037 |
Pattern Recognit. | 3 |
| 2025 | Rethink Sparse Signals for Pose-Guided Text-to-Image GenerationabstractRecent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-image generation. However, dense representations raised new challenges, including editing difficulties and potential inconsistencies with textual prompts. This fact motivates us to revisit sparse signals for pose guidance, owing to their simplicity and shape-agnostic nature, which remains underexplored. This paper proposes a novel Spatial-Pose ControlNet(SP-Ctrl), equipping sparse signals with robust controllability for pose-guided image generation. Specifically, we extend OpenPose to a learnable spatial representation, making keypoint embeddings discriminative and expressive. Additionally, we introduce keypoint concept learning, which encourages keypoint tokens to attend to the spatial positions of each keypoint, thus improving pose alignment. Experiments on animal- and human-centric image generation tasks demonstrate that our method outperforms recent spatially controllable T2I generation approaches under sparse-pose guidance and even matches the performance of dense signal-based methods. Moreover, SP-Ctrl shows promising capabilities in diverse and cross-species generation through sparse signals. Codes will be available at https://github.com/DREAMXFAR/SP-Ctrl. Wenjie Xuan, Jing Zhang 0037, Juhua Liu, Bo Du 0001, Dacheng Tao |
ICCV | 3 |
| 2025 | Hi-SAM: Marrying Segment Anything Model for Hierarchical Text SegmentationabstractThe Segment Anything Model (SAM), a profound vision foundation model pretrained on a large-scale dataset, breaks the boundaries of general segmentation and sparks various downstream applications. This paper introduces Hi-SAM, a unified model leveraging SAM for hierarchical text segmentation. Hi-SAM excels in segmentation across four hierarchies, including pixel-level text, word, text-line, and paragraph, while realizing layout analysis as well. Specifically, we first turn SAM into a high-quality pixel-level text segmentation (TS) model through a parameter-efficient fine-tuning approach. We use this TS model to iteratively generate the pixel-level text labels in a semi-automatical manner, unifying labels across the four text hierarchies in the HierText dataset. Subsequently, with these complete labels, we launch the end-to-end trainable Hi-SAM based on the TS architecture with a customized hierarchical mask decoder. During inference, Hi-SAM offers both automatic mask generation (AMG) mode and promptable segmentation (PS) mode. In the AMG mode, Hi-SAM segments pixel-level text foreground masks initially, then samples foreground points for hierarchical text mask generation and achieves layout analysis in passing. As for the PS mode, Hi-SAM provides word, text-line, and paragraph masks with a single point click. Experimental results show the state-of-the-art performance of our TS model: 84.86% fgIOU on Total-Text and 88.96% fgIOU on TextSeg for pixel-level text segmentation. Moreover, compared to the previous specialist for joint hierarchical detection and layout analysis on HierText, Hi-SAM achieves significant improvements: 4.73% PQ and 5.39% F1 on the text-line level, 5.49% PQ and 7.39% F1 on the paragraph level layout analysis, requiring fewer training epochs. Maoyuan Ye, Jing Zhang 0037, Juhua Liu, Cong Liu 0006, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Detect Changes Like Humans: Incorporating Semantic Priors for Improved Change DetectionabstractWhen given two similar images, humans identify their differences by comparing the appearance (e.g., color, texture) with the help of semantics (e.g., objects, relations). However, mainstream binary change detection models adopt a supervised training paradigm, where the annotated binary change map is the main constraint. Thus, such methods primarily emphasize difference-aware features between bi-temporal images, and the semantic understanding of changed landscapes is undermined, resulting in limited accuracy in the face of noise and illumination variations. To this end, this paper explores incorporating semantic priors from visual foundation models to improve the ability to detect changes. Firstly, we propose a Semantic-Aware Change Detection network (SA-CDNet), which transfers the knowledge of visual foundation models (i.e., FastSAM) to change detection. Inspired by the human visual paradigm, a novel dual-stream feature decoder is derived to distinguish changes by combining semantic-aware features and difference-aware features. Secondly, we explore a single-temporal pre-training strategy for better adaptation of visual foundation models. With pseudo-change data constructed from single-temporal segmentation datasets, we employ an extra branch of proxy semantic segmentation task for pre-training. We explore various settings like dataset combinations and landscape types, thus providing valuable insights. Experimental results on five challenging benchmarks demonstrate the superiority of our method over the existing state-of- the-art methods. The code is available at SA-CD. Yuhang Gan, Wenjie Xuan, Zhiming Luo, Zengmao Wang, Juhua Liu, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Revisiting Knowledge Distillation for Autoregressive Language ModelsabstractKnowledge distillation (KD) is a common approach to compress a teacher model to reduce its inference cost and memory footprint, by training a smaller student model.However, in the context of autoregressive language models (LMs), we empirically find that larger teachers might dramatically result in a poorer student.In response to this problem, we conduct a series of analyses and reveal that different tokens have different teaching modes, neglecting which will lead to performance degradation.Motivated by this, we propose a simple yet effective adaptive teaching approach (ATKD) to improve the KD.The core of ATKD is to reduce rote learning and make teaching more diverse and flexible.Extensive experiments on 8 LM tasks show that, with the help of ATKD, various baseline KD methods can achieve consistent and significant performance gains (up to +3.04% average score) across all model types and sizes.More encouragingly, ATKD can improve the student model generalization effectively. Qihuang Zhong, Liang Ding 0006, Li Shen 0008, Juhua Liu, Bo Du 0001, Dacheng Tao |
ACL (1) | 4 |
| 2024 | When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following AbilityabstractControlNet excels at creating content that closely matches precise contours in user-provided masks. However, when these masks contain noise, as a frequent occurrence with non-expert users, the output would include unwanted artifacts. This paper first highlights the crucial role of controlling the impact of these inexplicit masks with diverse deterioration levels through in-depth analysis. Subsequently, to enhance controllability with inexplicit masks, an advanced Shape-aware ControlNet consisting of a deterioration estimator and a shape-prior modulation block is devised. The deterioration estimator assesses the deterioration factor of the provided masks. Then this factor is used in a modulation block to adaptively adjust the model's contour-following ability, which helps it dismiss the noise part in the inexplicit masks. Extensive experiments prove its effectiveness in encouraging ControlNet to interpret inaccurate spatial conditions robustly rather than blindly following the given contours, suitable for diverse kinds of conditions. We showcase application scenarios like modifying shape priors and composable shape-controllable generation. Codes are available at github. Wenjie Xuan, Yufei Xu, Shanshan Zhao 0001, Juhua Liu, Bo Du 0001, Dacheng Tao |
ACM Multimedia | 5 |
| 2024 | GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term MatchingabstractBeyond the text detection and recognition tasks in image text spotting, video text spotting presents an augmented challenge with the inclusion of tracking. While advanced end-to-end trainable methods have shown commendable performance, the pursuit of multi-task optimization may pose the risk of producing sub-optimal outcomes for individual tasks. In this paper, we identify a main bottleneck in the state-of-the-art video text spotter: the limited recognition capability. In response to this issue, we propose to efficiently turn an off-the-shelf query-based image text spotter into a specialist on video and present a simple baseline termed GoMatching, which focuses the training efforts on tracking while maintaining strong recognition performance. To adapt the image text spotter to video datasets, we add a rescoring head to rescore each detected instance's confidence via efficient tuning, leading to a better tracking candidate pool.
Additionally, we design a long-short term matching module, termed LST-Matcher, to enhance the spotter's tracking capability by integrating both long- and short-term matching results via Transformer. Based on the above simple designs, GoMatching delivers new records on ICDAR15-video, DSText, BOVText, and our proposed novel test set with arbitrary-shaped text termed ArTVideo, which demonstates GoMatching's capability to accommodate general, dense, small, arbitrary-shaped, Chinese and English text scenarios while saving considerable training budgets. The code will be released. Haibin He 0001, Maoyuan Ye, Jing Zhang 0037, Juhua Liu, Bo Du 0001, Dacheng Tao |
NeurIPS | 4 |
| 2024 | Diff-Font: Diffusion Model for Robust One-Shot Font Generation
Haibin He 0001, Juhua Liu, Bo Du 0001, Dacheng Tao, Yu Qiao 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Successive model-agnostic meta-learning for few-shot fault time series prognosis
Hai Su, Jiajun Hu, Songsen Yu, Juhua Liu, Xiangyang Qin |
Neurocomputing | 4 |
| 2024 | RFL-CDNet: Towards accurate change detection via richer feature learning
Yuhang Gan, Wenjie Xuan, Juhua Liu, Bo Du 0001 |
Pattern Recognit. | 4 |
| 2024 | PanDa: Prompt Transfer Meets Knowledge Distillation for Efficient Model AdaptationabstractPrompt Transfer (PoT) is a recently-proposed approach to improve prompt-tuning, by initializing the target prompt with the existing prompt trained on similar source tasks. However, such a vanilla PoT approach usually achieves sub-optimal performance, as (i) the PoT is sensitive to the similarity of source-target pair and (ii) directly fine-tuning the prompt initialized with source prompt on target task might lead to forgetting of the useful general knowledge learned from source task. To tackle these issues, we propose a new metric to accurately predict the prompt transferability (regarding (i)), and a novel PoT approach (namelyPanDa) that leverages the knowledge distillation technique to alleviate the knowledge forgetting effectively (regarding (ii)). Extensive and systematic experiments on 189 combinations of 21 source and 9 target datasets across 5 scales of PLMs demonstrate that: 1)our proposed metric works well to predict the prompt transferability; 2)ourPanDaconsistently outperforms the vanilla PoT approach by 2.3% average score (up to 24.1%) among all tasks and model sizes; 3)with ourPanDaapproach, prompt-tuning can achieve competitive and even better performance than model-tuning in various PLM scales scenarios. Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and GenerationabstractSequence-to-sequence (seq2seq) learning is a popular fashion for large-scale pretraining language models. However, the previous seq2seq pretraining models generally focus on reconstructive objectives on the decoder side and neglect the effect of encoder-side supervision, which we argue may lead to sub-optimal performance. To verify our hypothesis, we first empirically study the functionalities of the encoder and decoder in seq2seq pretrained language models, and find that the encoder takes an important but under-exploitation role than the decoder regarding the downstream performance and neuron activation. Therefore, we propose an encoding-enhanced seq2seq pretraining strategy, namelyE2S2, which improves the seq2seq models via integrating more efficient self-supervised information into the encoders. Specifically, E2S2 adopts two self-supervised objectives on the encoder side from two aspects: 1) locally denoising the corrupted sentence (denoising objective); and 2) globally learning better sentence representations (contrastive objective). With the help of both objectives, the encoder can effectively distinguish the noise tokens and capture high-level (i.e., syntactic and semantic) knowledge, thus strengthening the ability of seq2seq model to accurately achieve the conditional generation. On a large diversity of downstream natural language understanding and generation tasks, E2S2 dominantly improves the performance of its powerful backbone models, e.g., BART and T5. For example, upon BART backbone, we achieve +1.1% averaged gain on the general language understanding evaluation (GLUE) benchmark and +1.75%$F_{0.5}$score improvement on CoNLL2014 dataset. We also provide in-depth analyses to show the improvement stems from better linguistic representation. We hope that our work will foster future self-supervision research on seq2seq language model pretraining. Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in TransformerabstractRecently, Transformer-based methods, which predict polygon points or Bezier curve control points for localizing texts, are popular in scene text detection. However, these methods built upon detection transformer framework might achieve sub-optimal training efficiency and performance due to coarse positional query modeling. In addition, the point label form exploited in previous works implies the reading order of humans, which impedes the detection robustness from our observation. To address these challenges, this paper proposes a concise Dynamic Point Text DEtection TRansformer network, termed DPText-DETR. In detail, DPText-DETR directly leverages explicit point coordinates to generate position queries and dynamically updates them in a progressive way. Moreover, to improve the spatial inductive bias of non-local self-attention in Transformer, we present an Enhanced Factorized Self-Attention module which provides point queries within each instance with circular shape guidance. Furthermore, we design a simple yet effective positional label form to tackle the side effect of the previous form. To further evaluate the impact of different label forms on the detection robustness in real-world scenario, we establish an Inverse-Text test set containing 500 manually labeled images. Extensive experiments prove the high training efficiency, robustness, and state-of-the-art performance of our method on popular benchmarks. The code and the Inverse-Text test set are available at https://github.com/ymy-k/DPText-DETR. Maoyuan Ye, Jing Zhang 0037, Shanshan Zhao 0001, Juhua Liu, Bo Du 0001, Dacheng Tao |
AAAI | 4 |
| 2023 | Revisiting Token Dropping Strategy in Efficient BERT PretrainingabstractToken dropping is a recently-proposed strategy to speed up the pretraining of masked language models, such as BERT, by skipping the computation of a subset of the input tokens at several middle layers.It can effectively reduce the training time without degrading much performance on downstream tasks.However, we empirically find that token dropping is prone to a semantic loss problem and falls short in handling semantic-intense tasks ( §2).Motivated by this, we propose a simple yet effective semantic-consistent learning method (SCTD) to improve the token dropping.SCTD aims to encourage the model to learn how to preserve the semantic information in the representation space.Extensive experiments on 12 tasks show that, with the help of our SCTD, token dropping can achieve consistent and significant performance gains across all task types and model sizes.More encouragingly, SCTD saves up to 57% of pretraining time and brings up to +1.56% average improvement over the vanilla token dropping. Qihuang Zhong, Liang Ding 0006, Juhua Liu, Xuebo Liu 0002, Min Zhang 0005, Bo Du 0001, Dacheng Tao |
ACL (1) | 3 |
| 2023 | DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text SpottingabstractEnd-to-end text spotting aims to integrate scene text detection and recognition into a unified framework. Dealing with the relationship between the two sub-tasks plays a pivotal role in designing effective spotters. Although Transformer-based methods eliminate the heuristic postprocessing, they still suffer from the synergy issue between the sub-tasks and low training efficiency. In this paper, we present DeepSolo, a simple DETR-like baseline that lets a single Decoder with Explicit Points Solo for text detection and recognition simultaneously. Technically, for each text instance, we represent the character sequence as ordered points and model them with learnable explicit point queries. After passing a single decoder, the point queries have encoded requisite text semantics and locations, thus can be further decoded to the center line, boundary, script, and confidence of text via very simple prediction heads in parallel. Besides, we also introduce a text-matching criterion to deliver more accurate supervisory signals, thus enabling more efficient training. Quantitative experiments on public benchmarks demonstrate that DeepSolo outperforms previous state-of-the-art methods and achieves better training efficiency. In addition, DeepSolo is also compatible with line annotations, which require much less annotation cost than polygons. The code is available at https://github.com/ViTAE-Transformer/DeepSolo. Maoyuan Ye, Jing Zhang 0037, Shanshan Zhao 0001, Juhua Liu, Tongliang Liu, Bo Du 0001, Dacheng Tao |
CVPR | 4 |
| 2023 | Zero-shot Sharpness-Aware Quantization for Pre-trained Language ModelsabstractQuantization is a promising approach for reducing memory overhead and accelerating inference, especially in large pre-trained language model (PLM) scenarios.While having no access to original training data due to security and privacy concerns has emerged the demand for zero-shot quantization.Most of the cuttingedge zero-shot quantization methods primarily ❶ apply to computer vision tasks, and ❷ neglect of overfitting problem in the generative adversarial learning process, leading to sub-optimal performance.Motivated by this, we propose a novel zero-shot sharpness-aware quantization (ZSAQ) framework for the zeroshot quantization of various PLMs.The key algorithm in solving ZSAQ is the SAM-SGA optimization, which aims to improve the quantization accuracy and model generalization via optimizing a minimax problem.We theoretically prove the convergence rate for the minimax optimization problem and this result can be applied to other nonconvex-PL minimax optimization frameworks.Extensive experiments on 11 tasks demonstrate that our method brings consistent and significant performance gains on both discriminative and generative PLMs, i.e., up to +6.98 average score.Furthermore, we empirically validate that our method can effectively improve the model generalization. Miaoxi Zhu, Qihuang Zhong, Li Shen 0008, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao |
EMNLP | 5 |
| 2023 | PNT-Edge: Towards Robust Edge Detection with Noisy Labels by Learning Pixel-level Noise TransitionsabstractRelying on large-scale training data with pixel-level labels, previous edge detection methods have achieved high performance. However, it is hard to manually label edges accurately, especially for large datasets, and thus the datasets inevitably contain noisy labels. This label-noise issue has been studied extensively for classification, while still remaining under-explored for edge detection. To address the label-noise issue for edge detection, this paper proposes to learn Pixel-level Noise Transitions to model the label-corruption process. To achieve it, we develop a novel Pixel-wise Shift Learning (PSL) module to estimate the transition from clean to noisy labels as a displacement field. Exploiting the estimated noise transitions, our model, named PNT-Edge, is able to fit the prediction to clean labels. In addition, a local edge density regularization term is devised to exploit local structure information for better transition learning. This term encourages learning large shifts for the edges with complex local structures. Experiments on SBD and Cityscapes demonstrate the effectiveness of our method in relieving the impact of label noise. Codes will be available at github.com/DREAMXFAR/PNT-Edge. Wenjie Xuan, Shanshan Zhao 0001, Yu Yao 0005, Juhua Liu, Tongliang Liu, Yixin Chen 0001, Bo Du 0001, Dacheng Tao |
ACM Multimedia | 4 |
| 2023 | An array of two periodic leaky-wave antennas with sum and difference beam scanning for application in target detection and trackingabstractAn array of two substrate-integrated waveguide (SIW) periodic leaky-wave antennas (LWAs) with sum and difference beam scanning is proposed for application in target detection and tracking. The array is composed of two periodic LWAs with different periods, in which each LWA generates a narrow beam through the n =−1 space harmonic. Due to the two different periods for the two LWAs, two beams with two different directions can be realized, which can be combined into a sum beam when the array is fed in phase or into a difference beam when the array is fed 180° out of phase. The array integrated with 180° hybrid is designed, fabricated, and measured. Measurement results show that the sum beam can reach a gain up to 15.9 dBi and scan from −33.4° to 20.8°. In the scanning range, the direction of the null in the difference beam is consistent with the direction of the sum beam, with the lowest null depth of −40. 8 dB. With the excellent performance, the antenna provides an alternative solution with low complexity and low cost for target detection and tracking. Mianfeng Huang, Juhua Liu |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2023 | Joint image and feature adaptative attention-aware networks for cross-modality semantic segmentation
Qihuang Zhong, Fanzhou Zeng, Juhua Liu, Bo Du 0001, Jedi S. Shang |
Neural Comput. Appl. | 4 |
| 2023 | Unified Instance and Knowledge Alignment Pretraining for Aspect-Based Sentiment AnalysisabstractThe goal of aspect-based sentiment analysis (ABSA) is to determine the sentiment polarity towards an aspect. Because of the expensive and limited amounts of labelled data, the pretraining strategy has become the de facto standard for ABSA. However, there always exists a severe domain shift between the pretraining and downstream ABSA datasets, which hinders effective knowledge transfer when directly fine-tuning, making the downstream task suboptimal. To mitigate this domain shift, we introduce a unified alignment pretraining framework into the vanilla pretrain-finetune pipeline, that has both instance- and knowledge-level alignments. Specifically, we first devise a novel coarse-to-fine retrieval sampling approach to select target domain-related instances from the large-scale pretraining dataset, thus aligning the instances between pretraining and the target domains (First Stage). Then, we introduce a knowledge guidance-based strategy to further bridge the domain gap at the knowledge level. In practice, we formulate the model pretrained on the sampled instances into a knowledge guidance model and a learner model. On the target dataset, we design an on-the-fly teacher-student joint fine-tuning approach to progressively transfer the knowledge from the knowledge guidance model to the learner model (Second Stage). Therefore, the learner model can maintain more domain-invariant knowledge when learning new knowledge from the target dataset. In theThird Stage,the learner model is finetuned to better adapt its learned knowledge to the target dataset. Extensive experiments and analyses on several ABSA benchmarks demonstrate the effectiveness and universality of our proposed pretraining framework. Juhua Liu, Qihuang Zhong, Liang Ding 0006, Bo Du 0001, Dacheng Tao |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Knowledge Graph Augmented Network Towards Multiview Representation Learning for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is a fine-grained task of sentiment analysis. To better comprehend long complicated sentences and obtain accurate aspect-specific information, linguistic and commonsense knowledge are generally required in this task. However, most current methods employ complicated and inefficient approaches to incorporate external knowledge, e.g., directly searching the graph nodes. Additionally, the complementarity between external knowledge and linguistic information has not been thoroughly studied. To this end, we propose a knowledge graph augmented network (KGAN), which aims to effectively incorporate external knowledge with explicitly syntactic and contextual information. In particular, KGAN captures the sentiment feature representations from multiple different perspectives,i.e., context-, syntax- and knowledge-based. First, KGAN learns the contextual and syntactic representations in parallel to fully extract the semantic features. Then, KGAN integrates the knowledge graphs into the embedding space, based on which the aspect-specific knowledge representations are further obtained via an attention mechanism. Last, we propose a hierarchical fusion module to complement these multi-view representations in alocal-to-globalmanner. Extensive experiments on five popular ABSA benchmarks demonstrate the effectiveness and robustness of our KGAN. Notably, with the help of the pretrained model of RoBERTa, KGAN achieves a new record of state-of-the-art performance among all datasets. Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Visual Semantics Allow for Textual Reasoning Better in Scene Text RecognitionabstractExisting Scene Text Recognition (STR) methods typically use a language model to optimize the joint probability of the 1D character sequence predicted by a visual recognition (VR) model, which ignore the 2D spatial context of visual semantics within and between character instances, making them not generalize well to arbitrary shape scene text. To address this issue, we make the first attempt to perform textual reasoning based on visual semantics in this paper. Technically, given the character segmentation maps predicted by a VR model, we construct a subgraph for each instance, where nodes represent the pixels in it and edges are added between nodes based on their spatial similarity. Then, these subgraphs are sequentially connected by their root nodes and merged into a complete graph. Based on this graph, we devise a graph convolutional network for textual reasoning (GTR) by supervising it with a cross-entropy loss. GTR can be easily plugged in representative STR models to improve their performance owing to better textual reasoning. Specifically, we construct our model, namely S-GTR, by paralleling GTR to the language model in a segmentation-based STR baseline, which can effectively exploit the visual-linguistic complementarity via mutual learning. S-GTR sets new state-of-the-art on six challenging STR benchmarks and generalizes well to multi-linguistic datasets. Code is available at https://github.com/adeline-cs/GTR. Yue He 0005, Jing Zhang 0037, Juhua Liu, Fengxiang He, Bo Du 0001 |
AAAI | 4 |
| 2022 | I3CL: Intra- and Inter-Instance Collaborative Learning for Arbitrary-Shaped Scene Text Detection
Bo Du 0001, Jing Zhang 0037, Juhua Liu, Dacheng Tao |
Int. J. Comput. Vis. | 4 |
| 2022 | FCL-Net: Towards accurate edge detection via Fine-scale Corrective Learning
Wenjie Xuan, Shaoli Huang, Juhua Liu, Bo Du 0001 |
Neural Networks | 3 |
| 2022 | An End-to-end Supervised Domain Adaptation Framework for Cross-Domain Change Detection
Wenjie Xuan, Yuhang Gan, Yibing Zhan, Juhua Liu, Bo Du 0001 |
Pattern Recognit. | 5 |
| 2020 | TextFuseNet: Scene Text Detection with Richer Fused FeaturesabstractArbitrary shape text detection in natural scenes is an extremely challenging task. Unlike existing text detection approaches that only perceive texts based on limited feature representations, we propose a novel framework, namely TextFuseNet, to exploit the use of richer features fused for text detection. More specifically, we propose to perceive texts from three levels of feature representations, i.e., character-, word- and global-level, and then introduce a novel text representation fusion technique to help achieve robust arbitrary text detection. The multi-level feature representation can adequately describe texts by dissecting them into individual characters while still maintaining their general semantics. TextFuseNet then collects and merges the texts’ features from different levels using a multi-path fusion architecture which can effectively align and fuse different representations. In practice, our proposed TextFuseNet can learn a more adequate description of arbitrary shapes texts, suppressing false positives and producing more accurate detection results. Our proposed framework can also be trained with weak supervision for those datasets that lack character-level annotations. Experiments on several datasets show that the proposed TextFuseNet achieves state-of-the-art performance. Specifically, we achieve an F-measure of 94.3% on ICDAR2013, 92.1% on ICDAR2015, 87.1% on Total-Text and 86.6% on CTW-1500, respectively. Zhe Chen 0013, Juhua Liu, Bo Du 0001 |
IJCAI | 3 |
| 2020 | SemiText: Scene text detection with semi-supervised learning
Juhua Liu, Qihuang Zhong, Hai Su, Bo Du 0001 |
Neurocomputing | 1 |
| 2020 | Locally and multiply distorted image quality assessment via multi-stage CNNs
Hai Su, Juhua Liu |
Inf. Process. Manag. | 3 |
| 2020 | Target Detection in Hyperspectral Imagery via Sparse and Dense Hybrid RepresentationabstractRepresentation-based target detectors for hyperspectral imagery (HSI) have recently aroused a lot of interests. However, existing methods ignore the dictionary structure and cannot guarantee an informative and discriminative representation of test pixels for target detection. To alleviate the problem, this letter proposes a novel sparse and dense hybrid representation-based target detector (SDRD). The proposed detector adopts the idea that the relationship between the background and the target sub-dictionaries is a collaborative competition. The structure of the dictionary is discovered and preserved by learning a sparse and dense hybrid representation for test pixel. Benefitting from this, a compact and discriminative representation can be obtained to better represent the test pixel for an improved detection performance. Experimental results on several HSI data sets verify the effectiveness of SDRD in comparison with several state-of-the-art methods. Tan Guo, Fulin Luo, Lei Zhang 0038, Xiaoheng Tan, Juhua Liu, Xiaocheng Zhou |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | ASTS: A Unified Framework for Arbitrary Shape Text SpottingabstractArbitrary shape text spotting remains a challenging computer vision task. In this paper, we propose an end-to-end trainable unified framework for arbitrary shape text spotting to overcome the limitations inherent in the existing methods. Specifically, we propose to perceive and understand text based on different levels of semantics, i.e ., holistic-, pixel- and sequence-level semantics, and then unify the recognized semantics for robust text spotting. To implement the framework, we customize the detection and mask branches of Mask R-CNN to explore both holistic- and pixel-level semantics for text recognition. According to the recognition results, the text spotting task can then be formulated in the two-dimensional feature space. Then, by feeding the two-dimensional feature maps into an additional text recognition branch, our framework further delivers one-dimensional sequence-level semantics for text recognition based on an attention-based sequence-to-sequence network. Finally, the results from all the three levels of semantics are merged as the final result. Therefore, our framework is capable of simultaneously recognizing texts from both the one- and two-dimensional perspectives, achieving highly comprehensive text recognition. In addition, because some existing datasets lack character-level annotations, the extensive descriptions of texts from our framework further allow us to use only word-level annotations as weak supervision for training a robust text spotting model. Experiments on ICDAR 2013, ICDAR 2015, and Total-Text show that our framework achieves state-of-the-art performance for both detection and recognition. Juhua Liu, Zhe Chen 0013, Bo Du 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2020 | Multistage GAN for Fabric Defect DetectionabstractFabric defect detection is an intriguing but challenging topic. Many methods have been proposed for fabric defect detection, but these methods are still suboptimal due to the complex diversity of both fabric textures and defects. In this paper, we propose a generative adversarial network (GAN)-based framework for fabric defect detection. Considering existing challenges in real-world applications, the proposed fabric defect detection system is capable of learning existing fabric defect samples and automatically adapting to different fabric textures during different application periods. Specifically, we customize a deep semantic segmentation network for fabric defect detection that can detect different defect types. Furthermore, we attempted to train a multistage GAN to synthesize reasonable defects in new defect-free samples. First, a texture-conditioned GAN is trained to explore the conditional distribution of defects given different texture backgrounds. Given a novel fabric, we aim to generate reasonable defective patches. Then, a GAN-based fusion network fuses the generated defects to specific locations. Finally, the well-trained multistage GAN continuously updates the existing fabric defect datasets and contributes to the fine-tuning of the semantic segmentation network to better detect defects under different conditions. Comprehensive experiments on various representative fabric samples are conducted to verify the detection performance of our proposed method. Juhua Liu, Hai Su, Bo Du 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2016 | A minimal Munsell value error based laser printer model
Juhua Liu, Hai Su, Wenbin Hu 0001, Lefei Zhang, Dacheng Tao |
Neurocomputing | 1 |
| 2016 | Optimal parameters based stochastic dot model for tone compensation of dither matrix
Hai Su, Juhua Liu, Yaohua Yi, Bo Du 0001 |
Neurocomputing | 2 |
| 2016 | Robust text detection via multi-degree of sharpening and blurring
Juhua Liu, Hai Su, Yaohua Yi, Wenbin Hu 0001 |
Signal Process. | 1 |
| 2011 | Wind Integration in Power Systems: Operational Challenges and Possible SolutionsabstractThis paper surveys major technical challenges for power system operations in support of large-scale wind energy integration. The fundamental difficulties of integrating wind power arise from its high inter-temporal variation and limited predictability. The impact of wind power integration is manifested in, but not limited to, scheduling, frequency regulations, and system stabilization requirements. Possible alternatives are suggested for a more reliable and cost-effective power system operation. New computationally efficient methods for improving system performances by using prediction and operational interdependencies over different time horizons remain critical open research problems. Le Xie 0001, Pedro M. S. Carvalho, Luis A. F. M. Ferreira, Juhua Liu, Bruce H. Krogh, Nipun Popli, Marija D. Ilic |
Proc. IEEE | 4 |