EDBT 2026 Demo / reviewers in the wild / expert
Liuan Wang
dblp:142/0012
· DBLP profile ↗
28ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Published content vs. live-streamed: Empirical from digital content activities in online healthcare communities
Liuan Wang, Yuke Luo, Linan Zhang |
Decis. Support Syst. | 1 |
| 2026 | Causal-prior-informed prediction of coronary heart disease: Leveraging LLM-guided structure learning and imbalance-aware augmentation
Liuan Wang, Yujian Ma |
Knowl. Based Syst. | 2 |
| 2025 | Efficient Object Placement Via LLM and Diffusion ModelabstractWe address the problem of object placement with user instructions using LLM and diffusion model. Traditional methods struggle to find a suitable location for filling the object with a semantically reasonable size. In this work, we leverage the LLM to predict the coordinates of the added object with the help of user instruction. First, We extend the object placement benchmark OPA-INST for object placement image editing. Second, we propose a framework Inst-GEdit to predict the coordinates of added object and blend the object with background via diffusion model in a semantically consistent and natural manner. Inst-GEdit can accomplish the object placement editing task with the user instruction in a mask free and training free manner. We conduct the experiments and evaluate our method on the OPA and image composition benchmark. We evaluate the results in both subjective and quantitative way, which demonstrate the effectiveness of our proposed method. Liuan Wang |
ICASSP | 2 |
| 2025 | Proxy-Mamba: Training-Free Architecture Search for Mamba via Gradient-Weight Correlation
Huigang Zhang, Liuan Wang, Jun Sun 0004 |
PRICAI | 2 |
| 2025 | The roles of trust and privacy calculus in citizen-centric services usage: evidence from the close contact query platform in ChinaabstractThe mechanism of the ‘privacy calculus’ between privacy risks and benefits has gained attention in e-government service, but there were few empirical studies. Trust reduces perceived risk as well as enhances perceived benefits. However, the effects of different types of trust on privacy calculus are yet undefined. Regarding above issues, we build a research model based on the theoretical framework of ‘Antecedents → Privacy Concerns → Outcomes (APCO)’, theories of privacy calculus and e-government trustworthiness typology. We analyse survey data from the Close Contact Query Platform providing citizen-centric services for COVID-19 in China. Trust in transactions and information is increased by trust in government organisations. Trust in government organisations and transactions decrease perceived privacy risks, while trust in government organisations and information enhances perceived privacy benefits. Privacy risks and benefits significantly affect usage intentions, and the effect of privacy benefits is greater. We also identified mediating effects. Our findings have important implications regarding the unexplored inner mechanism of trust, the roles of privacy calculus in e-government adoption, and the expansion of the APCO framework. We offer recommendations for citizen-centric platform developers, administrators to design citizen-centric services, and offline public servants. Jingrui Ju, Liuan Wang |
Behav. Inf. Technol. | 2 |
| 2024 | NR-CION: Non-rigid Consistent Image Composition Via Diffusion Model
Wei Liu 0169, Liuan Wang, Jun Sun 0004 |
ICPR (25) | 2 |
| 2024 | Temporal Insight Enhancement: Mitigating Temporal Hallucination in Video Understanding by Multimodal Large Language Models
Li Sun 0007, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
ICPR (7) | 2 |
| 2024 | The faster or richer the response, the better performance? An empirical analysis of online healthcare platforms from a competitive perspective
Liuan Wang |
Decis. Support Syst. | 2 |
| 2024 | Variously and freely to use: Exploring routine and innovative use of fitness apps from a self-management perspective
Aoshuang Li, Yongqiang Sun, Liuan Wang, Jinyu Guo |
Inf. Manag. | 3 |
| 2024 | Towards Generated Image Provenance Analysis via Conceptual-Similar-Guided-SLIP RetrievalabstractWith the prevalence of state-of-the-art generative models, photorealistic synthetic images can now be easily generated. However, the generated images may replicate contents from the original training images, which can lead to potential legal issues. In this paper, we propose a novel method calledConceptual-Similar-guided Self-supervised Language-Image Pre-training(CS-SLIP) that leverages both image and text modalities for the generated image provenance. Besides the self-supervised learning branch and contrastive learning branch, a conceptual-similar branch is designed to guide the model to learn a better feature representation of image-text-pairs. We also adopt the re-ranking method to refine the initial matching candidates via the cross-modal bi-directional retrieval. Extensive qualitative and quantitative experiments are conducted, which demonstrate that the replication indeed exists in the generated images, and our proposed method can effectively retrieve the most similar images from the training corpus to achieve the goal of generated image provenance analysis. Xiaojie Xia, Liuan Wang, Jun Sun 0004, Akira Nakagawa |
IEEE Signal Process. Lett. | 2 |
| 2023 | Semantic-Embedded Knowledge Acquisition and Reasoning for Image SegmentationabstractImage segmentation is a difficult and challenging task because of the complex object appearance and diverse object categories. Traditional methods directly use visual features for segmentation but ignore the correlation between objects. We introduce a knowledge reasoning module (KRM) for external knowledge aggregation and leverage a graphic neural network to aggregate the knowledge feature, which is concatenated with a visual feature for semantic segmentation. To this end, we use word embedding of category names as semantic feature and establish the relationship between categories. Through iteration, the aggregated features can be enriched. In experiments, three well known semantic segmentation methods are used as baseline. Our experiment results outperform the baseline methods on the food dataset Food-Seg103 and Cityscapes, and demonstrate the effectiveness of our proposed method. Wei Liu 0169, Huigang Zhang, Xiaojie Xia, Liuan Wang, Jun Sun 0004 |
ICIP | 4 |
| 2023 | Prompt Prototype Learning Based on Ranking Instruction For Few-Shot Visual TasksabstractQuerying large language models (LLMs), such as GPT-3, for high-quality prompts and utilizing pre-trained vision-language models, such as CLIP, to construct a zero-shot visual classification model, offer promising performance across various downstream visual tasks. However, when applied to specific domains, their efficacy is restricted due to the gap between the general prompts they generate and the required domain-specific knowledge. In this paper, we propose a novel, lightweight method for prompt prototype learning through ranking instruction, specifically designed to bridge this gap in the context of few-shot visual classification. We generate domain-specific prompts leveraging the knowledge contained in LLMs and then fine-tune the prompt prototype with effective ranking instructions from several domain images. Our few-shot experiments on facial expression benchmarks demonstrate the efficacy of the prompt prototype. Notably, our method delivers results that are on par with state-of-the-art few-shot image classification techniques and can be integrated with them to further improve performance in the facial expression domain. Our approach provides a promising solution to few-shot visual classification, making use of the knowledge contained in LLMs to generate domain-specific prompts. Li Sun 0007, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
ICIP | 2 |
| 2023 | Zero-shot temporal event localisation: Label-free, training-free, domain-freeabstractAbstract Temporal event localisation (TEL) has recently attracted increasing attention due to the rapid development of video platforms. Existing methods are based on either fully/weakly supervised or unsupervised learning, and thus they rely on expensive data annotation and time‐consuming training. Moreover, these models, which are trained on specific domain data, limit the model generalisation to data distribution shifts. To cope with these difficulties, the authors propose a zero‐shot TEL method that can operate without training data or annotations. Leveraging large‐scale vision and language pre‐trained models, for example, CLIP, we solve the two key problems: (1) how to find the relevant region where the event is likely to occur; (2) how to determine event duration after we find the relevant region. Query guided optimisation for local frame relevance relying on the query‐to‐frame relationship is proposed to find the most relevant frame region where the event is most likely to occur. Proposal generation method relying on the frame‐to‐frame relationship is proposed to determine the event duration. The authors also propose a greedy event sampling strategy to predict multiple durations with high reliability for the given event. The authors’ methodology is unique, offering a label‐free, training‐free, and domain‐free approach. It enables the application of TEL purely at the testing stage. The practical results show it achieves competitive performance on the standard Charades‐STA and ActivityCaptions datasets. Li Sun 0007, Ping Wang 0034, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
IET Comput. Vis. | 3 |
| 2023 | Exploiting spatio-temporal knowledge for video action recognitionabstractAbstract Action recognition has been a popular area of computer vision research in recent years. The goal of this task is to recognise human actions in video frames. Most existing methods often depend on the visual features and their relationships inside the videos. The extracted features only represent the visual information of the current video itself and cannot represent the general knowledge of particular actions beyond the video. Thus, there are some deviations in these features, and the recognition performance still requires improvement. In this sudy, we present a novel spatio‐temporal knowledge module (STKM) to endow the current methods with commonsense knowledge. To this end, we first collect hybrid external knowledge from universal fields, which contains both visual and semantic information. Then graph convolution networks (GCN) are used to represent and aggregate this knowledge. The GCNs involve (i) a spatial graph to capture spatial relations and (ii) a temporal graph to capture serial occurrence relations among actions. By integrating knowledge and visual features, we can get better recognition results. Experiments on AVA, UCF101‐24 and JHMDB datasets show the robustness and generalisation ability of STKM. The results report a new state‐of‐the‐art 32.0 mAP on AVA v2.1. On UCF101‐24 and JHMDB datasets, our method also improves by 1.5 AP and 2.6 AP, respectively, over the baseline method. Huigang Zhang, Liuan Wang, Jun Sun 0004 |
IET Comput. Vis. | 2 |
| 2022 | FBMOT: Flow Bridges the Gap between Detection and Tracking in Multiple Object TrackingabstractThe detection performance is a key bottleneck to the performance of multiple object tracking (MOT) algorithms and the advanced detection algorithms contribute a large portion to the success of MOT algorithms. Even though, the detectors can still make false detections based on only the image frames, which may directly result in the loss or mismatch of tracks. In MOT tasks, we point out that the video frames are temporally correlated and the temporal relationship can be leveraged to further improve the performance of detection and subsequent tracking. In our work, we propose FBMOT which uses optical flow to compute a prior heatmap about the locations of previously tracked objects in the current frame and takes the heatmap as an additional input in detection. We further proposed a novel regularization loss to help our model distinguish useful information in the prior heatmap. As a result, our method improves 0.8, 1.0 and 2.7 MOTA in MOT16, MOT17 and MOT20 test datasets respectively. Lisheng Wu, Liuan Wang, Jun Sun 0004 |
AVSS | 2 |
| 2021 | Key-Guided Identity Document Classification Method by Graph Attention Network
Xiaojie Xia, Wei Liu 0169, Ying Zhang 0124, Liuan Wang, Jun Sun 0004 |
ICDAR (4) | 4 |
| 2021 | Knowledge-Based Reasoning Network For Object DetectionabstractThe mainstream object detection algorithms rely on recognizing object instances individually, but do not consider the high-level relationship among objects in context. This will inevitably lead to biased detection results, due to the lack of commonsense knowledge that humans often use to assist the task for object identification. In this paper, we present a novel reasoning module to endow the current detection systems with the power of commonsense knowledge. Specifically, we use graph attention network (GAT) to represent the knowledge among objects. The knowledge covers visual and semantic relations. Through the iterative update of GAT, the object features can be enriched. Experiments on the COCO detection benchmark indicate that our knowledge-based reasoning network has achieved consistent improvements upon various CNN detectors. We achieved 1.9 and 1.8 points higher Average Precision (AP) than Faster-RCNN and Mask-RCNN respectively, when using ResNet50-FPN as backbone. Huigang Zhang, Liuan Wang, Jun Sun 0004 |
ICIP | 2 |
| 2021 | Exploring Pathologist Knowledge for Automatic Assessment of Breast Cancer Metastases in Whole-slide ImageabstractAutomatic assessment of breast cancer metastases plays an important role to help pathologist reduce the time-consuming work in histopathological whole-slide image diagnosis. From the utilization of knowledge point of view, the low-magnification level and high-magnification level are carefully checked by the pathologists for tumor pattern and cell tumor characteristic. In this paper, we propose a novel automatic patient-level tumor segmentation and classification method, which makes full use of the diagnosis knowledge clues from pathologists. For tumor segmentation, a multi-level view DeepLabV3+ (MLV-DeepLabV3+) is designed to explore the distinguishing features of cell characteristics between tumor and normal tissue. Furthermore, the expert segmentation models are selected and integrated by Pareto-front optimization to imitate the expert consultation to get perfect diagnosis. For wholeslide classification, multi-level magnifications are adaptive checked to focus on the effective features in different magnification. The experimental results demonstrate that our pathologist knowledge-based automatic assessment of whileslide image is effective and robust on the public benchmark dataset. Liuan Wang, Li Sun 0007, Huigang Zhang, Ping Wang 0034, Rong Zhou 0005, Jun Sun 0004 |
ACM Multimedia | 1 |
| 2020 | NAS-EOD: an end-to-end Neural Architecture Search method for Efficient Object DetectionabstractModel efficiency for object detection has become more and more important recently, especially when intelligent mobile devices get more and more popular. Current lightweight object detection model is either migrated from lightweight classification models, or pruned directly from complex object detection models. These pipelines can not match the performance requirements of edge devices. In this work, we propose a neural architecture search (NAS) method to build a detection model automatically that can perform well on edge devices. Specifically, the proposed method supports the search of not only multiscale feature network, but also backbone network. This enables us to search out a global optimal model. To this end, we've made a special design that the backbone and feature network can share the same search space. This method greatly reduces the search time while ensuring the search accuracy. It can find a good architecture in 14GPU days. Additionally, we add latency information into the main objective during performance estimation. Therefore, we can control between the speed and accuracy to better adapt to the edge environment. Experiments on the PASCAL VOC benchmark indicate that the searched architecture (named NAS-EOD) can get good accuracy even if training from scratch. When using a pre-training scheme, our model is superior to state-of-the-art small object detection models. Huigang Zhang, Liuan Wang, Jun Sun 0004, Li Sun 0007, Hiromichi Kobashi, Nobutaka Imamura |
ICPR | 2 |
| 2019 | Automatic Neural Network Search Method for Open Set RecognitionabstractReal-world recognition or classification tasks in computer vision are not apparent in controlled environments and often get involved in open set. Previous research work on real-world recognition problem is knowledge- and labor-intensive to pursue good performance for there are numbers of task domains. Auto Machine Learning (AutoML) approaches supply an easier way to apply advanced machine learning technologies, reduce the demand for experienced human experts and improve classification performance on close set. This paper proposes an automatic neural network search method for designing effective convolution neural network (CNN) models for open set recognition (OSR). Feature distribution information is explicitly incorporated into the main objective. So during the search process, the sampled models will enlarge interclass differences and reduce intra-class variations. We design a flexible search space based on classic CNN models to diversify neural architectures and also add some search principles to limit the size of the search space. Experimental results on CIFAR-10 and Dunhuang historical Chinese datasets show that our approach improves performances on both close and open set. Comparing with the other two OSR algorithms, our method also achieves the best performance. Li Sun 0007, Xiaoyi Yu, Liuan Wang, Jun Sun 0004, Hiroya Inakoshi, Ken Kobayashi, Hiromichi Kobashi |
ICIP | 3 |
| 2018 | Text Line Extraction Based on Integrated K-Shortest Paths OptimizationabstractText in images can be utilized in many image understanding applications due to the exact semantic information. In this paper, we propose a novel integrated k-shortest paths optimization based text line extraction method. Firstly, the candidate text components are extracted by the Maximal Stable Extremal Region (MSER) algorithm on gray, red, green and blue channels. Secondly, one integrated directed graph on red, green, and blue channels are constructed upon the candidate text components, which can effectively incorporate different channels into one framework. Then, the integrated directed graph is transformed guided by the extracted text lines in gray channel to reduced the computational complexity. Finally, we use the k-shortest paths optimization algorithm to extract the text lines by taking advantage of the particular structure of the integrated directed graph. Experimental results demonstrate the effectiveness of the proposed method in comparison with state-of-the-art methods. Liuan Wang, Jun Sun 0004, Seiichi Uchida |
DAS | 1 |
| 2018 | Compact Binary Feature for Open Set RecognitionabstractMost pattern recognition systems are closed set recognition systems in which any input sample is to be classified as belonging to one of the given classes. This paper, however, addresses the open set recognition problem in which a test sample may either come from one of the labeled "known classes" or come from an unknown class. The number of unknown classes could potentially be unlimited. A compact binary feature (CBF) generated by an ensemble binary classifier (EBC) is proposed to solve the open set recognition problem. This method can be regarded as a type of ECOC (Error-Correcting Output Codes) combined with modern CNN (Convolutional Neural Network) techniques and adapted for open set recognition. By randomly partitioning the known classes into two groups and training a binary classifier with CNN to separate them apart, and by repeating such a procedure for many times, rich information is extracted from the training set in the form of an EBC which can associate any test sample with a CBF that can be matched according to Hamming distance which is very efficient to compute. According to the experiments on the Dunhuang ancient Chinese character dataset, EBC can boost the recognition performance significantly compared with a single feedforward CNN. Apart from that, CBF is very efficient for storage and saves lots of time in feature matching at the cost of more computation in the training phase. Jun Sun 0004, Xiaoyi Yu, Liuan Wang |
DAS | 4 |
| 2016 | Globally Optimal Text Line Extraction Based on K-Shortest Paths AlgorithmabstractThe task of text line extraction in images is a crucial prerequisite for content-based image understanding applications. In this paper, we propose a novel text line extraction method based on k-shortest paths global optimization in images. Firstly, the candidate connected components are extracted by reformulating it as Maximal Stable Extremal Region (MSER) results in images. Then, the directed graph is built upon the connected component nodes with edges comprising of unary and pairwise cost function. Finally, the text line extraction problem is solved using the k-shortest paths optimization algorithm by taking advantage of the particular structure of the directed graph. Experimental results on public dataset demonstrate the effectiveness of proposed method in comparison with state-of-the-art methods. Liuan Wang, Seiichi Uchida, Wei Fan 0005, Jun Sun 0004 |
DAS | 1 |
| 2016 | Character region segmentation based on Stroke Stable RegionsabstractRegion segmentation is the key procedure in various text related image processing tasks. A good region extractor, which separates text area from complex background clutter, will reduce the burden of subsequent text grouping and post-processing functions. This paper propose a character region segmentation method based on a new concept named Stroke Stable Region (SSR) to achieve a better precision than many off-the-shelf region extractors such as MSER in the text segmentation task. Our work presented in this paper is inspired by the structure of MSER. However, instead of evaluating the area variation of each connected component, we proposed a novel parameter, stroke time, to measure the possibility that a pixel belongs to a character or a stroke by analyzing its character affinity. The experiments show that SSR tends to extract the visual objects with prominent text characteristics and is capable of suppressing various background noise. In a text extraction task on the ICDAR 2003 dataset, the SSR based method reduces the extracted noise components to about 1/3 of those obtained by the MSER based method, maintaining the same level of recall rate. The proposed algorithm was successfully applied to a wide range of text segmentation tasks. Hong Shang, Liuan Wang, Tanaka Hiroshi, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi |
ICPR | 2 |
| 2015 | Text line extraction in document imagesabstractText line extraction in document images is an important prerequisite for many content based image understanding applications. In this paper, we propose an accurate and robust method for generic text line extraction, which can be applied on large categories of document images, diverse languages, and text lines with different orientations. Firstly, the candidate connected components are extracted from document image using Maximal Stable Extremal Region (MSER) with the noises filtered by Adaboost and Convolution Neural Network (CNN). Then, the coarse text lines are generated from hierarchical edges reconstruction and cut by local linearity of text lines in the document spanning tree. Finally, for accurate text line extraction, the cut multi-components are re-connected based on text line energy minimization in terms of text line consistency and the fitting error. Experimental results on multilingual test dataset demonstrate the effectiveness and robust of the proposed method, which yields higher performance compared with state-of-the-art methods. Liuan Wang, Wei Fan 0005, Jun Sun 0004, Satoshi Naoi, Tanaka Hiroshi |
ICDAR | 1 |
| 2015 | Scene text recognition with deeper convolutional neural networksabstractScene text recognition plays an important role in many applications such as video indexing and house number localization in maps. Recently, some feature learning methods have been proposed to handle this problem, which often exploit deep architectures with no more than 5 layers and relatively large receptive fields. Meanwhile, to avoid model overfitting, they generally take advantage of large amount of additional data. Inspired by the great success of GoogleLeNet with a deeper network and VGG networks with smaller receptive fields in the ImageNet competition, in this paper, we adopt a much deeper network with up to 15 layers and smaller receptive fields (3×3) to learn better features for scene text recognition. Particularly, even without additional training data, our model can achieve better performance. Experiments on scene text datasets (ICDAR 2003, SVT, Chars74K) demonstrate that our method achieves the state-of-the-art performance on character classification and competitive performance on cropped word recognition. Yuqi Zhang 0001, Wei Wang 0115, Liang Wang 0001, Liuan Wang |
ICIP | 4 |
| 2014 | Fast and Accurate Text Detection in Natural Scene Images with User-IntentionabstractText detection in natural scene images plays an important role in content-based image retrieval, especially user-guided text detection for human-computer interaction. In this paper, we propose a fast and accurate text detection method with user-intention in terms of tap gesture. Firstly, a user-intention slice descriptor is designed based on the estimated text property, which contains all the user interested texts, and fast heuristic features and accurate texture feature of decomposed connected components (CCs) are fed into cascade of Gentle Adaboost classifiers to eliminate non-text candidates, finally candidate texts, sharing the same property consistent with the seed CCs, are accumulated to a user-intention text line according to local and global permutation constraint. Experimental results demonstrate the effectiveness and robustness of the proposed method in comparison with the state-of-art methods. Liuan Wang, Wei Fan 0005, Yuan He 0001, Jun Sun 0004, Yutaka Katsuyama, Yoshinobu Hotta |
ICPR | 1 |
| 2013 | Text detection in natural scene images with user-intentionabstractWe propose an accurate and robust coarse-to-fine text detection scheme with user-intention which captures the intrinsic characteristics of natural scene texts. In the coarse detection stage, a double edge detector is designed to estimate the symmetry of stroke and the stroke width, which help segment the foreground. Then the initial user-intention region is extended to generate a coarse bounding box based on the estimated foreground. In the refinement stage, candidate connected components (CCs) from Niblack decomposition, are grouped together by location to form text lines after noise removal and layer selection. Experimental results demonstrate the effectiveness of the proposed method which yields higher performance compared with state-of-the-art methods. Liuan Wang, Yutaka Katsuyama, Wei Fan 0005, Yuan He 0001, Jun Sun 0004, Yoshinobu Hotta |
ICIP | 1 |