VLDB 2026 Research / reviewers in the wild / expert
Haoran Liang 0001
dblp:120/8763-1
· DBLP profile ↗
27ranked-venue papers
6as first author
20since 2021 · last 2025
0000-0002-5906-1380ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCLSTE: Semi-supervised Contrastive Learning-Guided Scene Text EditingabstractThe objective of scene text editing is to substitute the existing text with desirable text, while preserving the background and text styles intact. However, existing methods struggle to effectively replicate textual styles due to the complexity of real scenarios, and most are limited to training exclusively on labeled synthetic datasets. One alternative semi-supervised method incorporates unlabeled real-scene text images for training by using the original image as supervision and editing it with the same text. However, this approach risks degrading the model into an identity mapping network. To address these problems, we introduce a novel semi-supervised training strategy incorporating contrastive learning. It allows for editing real-scene text images with any text, circumventing the identity mapping issue while ensuring the accuracy of both text content and style. Moreover, we propose a robust Style-Aware Text Editing Module to address complexity of real scenarios and enhance the imitation of text styles. To the best of our knowledge, our work is the first to apply contrastive learning to the scene text editing. Extensive experiments demonstrate that our method outperforms existing models in terms of quality and quantity. Especially for the HEval metrics on both real scene (Tamper-Scene) and synthetic scene (Temper-Syn2k), we get both 12% improvement compared to state-of-the-art method. Min Yin, Liang Xie 0003, Haoran Liang 0001, Xing Zhao 0001, Ben Chen 0004, Ronghua Liang |
MMM (3) | 3 |
| 2025 | FactExplorer: Fact Embedding-Based Exploratory Data Analysis for Tabular DataabstractDespite exploratory data analysis (EDA) is a powerful approach for uncovering insights from unfamiliar datasets, existing EDA tools face challenges in assisting users to assess the progress of exploration and synthesize coherent insights from isolated findings. To address these challenges, we present FactExplorer, a novel fact-based EDA system that shifts the analysis focus from raw data to data facts. FactExplorer employs a hybrid logical-visual representation, providing users with a comprehensive overview of all potential facts at the outset of their exploration. Moreover, FactExplorer introduces fact-mining techniques, including topic-based drill-down and transition path search capabilities. These features facilitate in-depth analysis of facts and enhance the understanding of interconnections between specific facts. Finally, we present a usage scenario and conduct a user study to assess the effectiveness of FactExplorer. The results indicate that FactExplorer facilitates the understanding of isolated findings and enables users to steer a thorough and effective EDA. Guodao Sun, Lvhan Pan, Baofeng Chang, Haoran Liang 0001, Ronghua Liang |
Int. J. Hum. Comput. Interact. | 7 |
| 2025 | Semantic-aware representations for unsupervised Camouflaged Object Detection
Zelin Lu, Xing Zhao 0001, Liang Xie 0003, Haoran Liang 0001, Ronghua Liang |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | A Weakly-Supervised Cross-Domain Query Framework for Video Camouflage Object DetectionabstractVCOD (Video Camouflage Object Detection) is a crucial security technology that identifies camouflaged objects in videos, bolstering security measures across diverse applications. On one hand, appearance-based VCOD methods face challenges because camouflaged appearances cause objects to blend into their surroundings, and current VCOD methods typically utilize optical flow to represent motion information. However, over-reliance on accurate estimation renders the model overly fragile. On the other hand, there is a shortage of effectively annotated camouflaged video datasets, coupled with the time-consuming and labor-intensive annotation process, severely constraining the development of this field. To address this, we propose a novel weakly-supervised framework for VCOD based on cross-domain querying of preceding and succeeding frames. Specifically, we propose a time-efficient and labor-saving manual annotation approach based on large visual models to rapidly generate pseudo-labels. Furthermore, we design a network based on Spatio-Temporal Memory (STM) that performs cross-modal feature querying with the current frame against preceding and succeeding frames to acquire useful information, thereby enhancing the focus on temporal information. Extensive experiments conducted on two common VCOD datasets have proven the effectiveness of our method, achieving state-of-the-art performance on the challenging camouflaged video data. Zelin Lu, Liang Xie 0003, Xing Zhao 0001, Binwei Xu, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multidimensional Exploration of Segment Anything Model for Weakly Supervised Video Salient Object DetectionabstractFully supervised video salient object detection (VSOD) has made considerable breakthroughs using costly and time-consuming pixel-wise annotations. Recently, to achieve a trade-off between the annotation burden and the model performance, scribble-based VSOD tasks have attracted increasing attention. However, learning the complete object structure and precise boundary details from sparse scribble annotations remains challenging. In this paper, we propose a series of strategies to effectively explore valid information from the recently proposed segmentation foundation model “Segment Anything Model (SAM)” in various perspectives to address these challenges. Specifically, due to the limited performance of SAM on videos, we propose a SAM-guided label enhancement method instead of directly using the results of SAM, which can introduce edge information while reducing the interference of erroneous information. Moreover, we propose a SAM-driven spatiotemporal network guided by general semantic features from the SAM encoder to help the model be aware of global connections. Additionally, we propose a SAM-based global-aware loss, which further considers the affinity constraint between predicted results and foreground labels or background labels from a global perspective, guiding the model to perceive the complete salient objects. Experimental results demonstrate that our method outperforms state-of-the-art weakly supervised VSOD methods and is comparable to fully supervised VSOD methods. Binwei Xu, Qiuping Jiang, Xing Zhao 0001, Chenyang Lu 0002, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Learning Video Salient Object Detection Progressively From Unlabeled VideosabstractRecently, deep learning-based video salient object detection (VSOD) has achieved some breakthroughs, but these methods rely on expensive annotated videos with pixel-wise annotations or weak annotations. In this paper, based on the similarities and differences between VSOD and image salient object detection (SOD), we propose a novel VSOD method via a progressive framework that locates and segments salient objects in sequence without utilizing any video annotation. To efficiently use the knowledge learned in the SOD dataset for VSOD efficiently, we introduce dynamic saliency to compensate for the lack of motion information of SOD during the locating process while maintaining the same fine segmenting process. Specifically, we utilize the coarse locating model trained on the image dataset, to identify frames with both static and dynamic saliency. Locating results of these frames are selected as spatiotemporal location labels. Moreover, by tracking salient objects in adjacent frames, the number of spatiotemporal location labels is increased. On the basis of these location labels, a two-stream locating network with an optical flow branch is proposed to capture salient objects in videos. The results with respect to five public benchmarks demonstrate that our method outperforms the state-of-the-art weakly and unsupervised methods. Binwei Xu, Qiuping Jiang, Haoran Liang 0001, Dingwen Zhang, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 3 |
| 2025 | Position Fusing and Refining for Clear Salient Object DetectionabstractMultilevel feature fusion plays a pivotal role in salient object detection (SOD). High-level features present rich semantic information but lack object position information, whereas low-level features contain object position information but are mixed with noises such as backgrounds. Appropriately addressing the gap between low- and high-level features is important in SOD. We first propose a global position embedding attention (GPEA) module to minimize the discrepancy between multilevel features in this article. We extract the position information by utilizing the semantic information at high-level features to resist noises at low-level features. Object refine attention (ORA) module is introduced to refine features used to predict saliency maps further without any additional supervision and heighten discriminative regions near the salient object, such as boundaries. Moreover, we find that the saliency maps generated by the previous methods contain some blurry regions, and we design a pixel value (PV) loss to help the model generate saliency maps with improved clarity. Experimental results on five commonly used SOD datasets demonstrated that the proposed method is effective and outperforms the state-of-the-art approaches on multiple metrics. Xing Zhao 0001, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | RE-IDVIS: Person Re-Identification System based on Interactive Visualizationabstractpixel-based visual encoding attribute-based visual encoding image-based visual encoding Figure 1: The interface of the system.(A) the probe panel which allows users to select the person-of-interest as a probe and set up the visual parameter of the search space.(B) the ranking list composed of the pixel-based visual encoding which allows users to quickly retrieve strong negative samples.(C) the search space view which supports visual exploration and enables users to provide feedback on the samples.(D) the spatiotemporal view which summarizes the spatiotemporal information of the retrieval results.(E) a cluster sample at three different visualization scales. Guodao Sun, Pan Liang, Sujia Zhu, Yiming Wu 0005, Haoran Liang 0001, Ronghua Liang |
ICMR | 7 |
| 2024 | Towards Small Object Editing: A Benchmark Dataset and A Training-Free ApproachabstractA plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in producing high-quality images, their application to small object generation has been limited due to difficulties in aligning cross-modal attention maps between text and these objects. Our approach offers a training-free method that significantly mitigates this alignment issue with local and global attention guidance, enhancing the model's ability to accurately render small objects in accordance with textual descriptions. We detail the methodology in our approach, emphasizing its divergence from traditional generation techniques and highlighting its advantages. What's more important is that we also provide SOEBench (Small Object Editing), a standardized benchmark for quantitatively evaluating text-based small object generation collected from MSCOCO[22] and OpenImage[18]. Preliminary results demonstrate the effectiveness of our method, showing marked improvements in the fidelity and accuracy of small object generation compared to existing models. This advancement not only contributes to the field of AI and computer vision but also opens up new possibilities for applications in various industries where precise image generation is critical.We will release our dataset on our project page: https://soebench.github.io/ Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Yiming Wu 0005, Wei Ji 0008, Haoran Liang 0001, Ronghua Liang |
ACM Multimedia | 7 |
| 2024 | Hierarchical multi-modal video summarization with dynamic samplingabstractAbstract Previous video summarization methods often neglected inter‐frame variations during the preprocessing stage. Sampling repeated frames can lead to information redundancy, while missing key frames can result in deviations in semantic comprehension and inaccuracies in the generated summaries. This work proposes a dynamic sampling module that leverages frame‐level motion information to alleviate these issues. The module conducts high‐frequency sampling during intervals with significant changes, allowing for a finer capture of details. Combined with a hierarchical multi‐modal structure, it integrates shot‐level visual and textual information to enhance the semantic understanding of video clips and improve the accuracy of the summarized content. Extensive experiments on benchmark datasets SumMe and TVSum demonstrate the effectiveness of the proposed method. Lingjian Yu, Xing Zhao 0001, Liang Xie 0003, Haoran Liang 0001, Ronghua Liang |
IET Image Process. | 4 |
| 2024 | Salient object detection in egocentric videosabstractAbstract In the realm of video salient object detection (VSOD), the majority of research has traditionally been centered on third‐person perspective videos. However, this focus overlooks the unique requirements of certain first‐person tasks, such as autonomous driving or robot vision. To bridge this gap, a novel dataset and a camera‐based VSOD model, CaMSD , specifically designed for egocentric videos, is introduced. First, the SalEgo dataset, comprising 17,400 fully annotated frames for video salient object detection, is presented. Second, a computational model that incorporates a camera movement module is proposed, designed to emulate the patterns observed when humans view videos. Additionally, to achieve precise segmentation of a single salient object during switches between salient objects, as opposed to simultaneously segmenting two objects, a saliency enhancement module based on the Squeeze and Excitation Block is incorporated. Experimental results show that the approach outperforms other state‐of‐the‐art methods in egocentric video salient object detection tasks. Dataset and codes can be found at https://github.com/hzhang1999/SalEgo . Haoran Liang 0001, Xing Zhao 0001, Jian Liu 0053, Ronghua Liang |
IET Image Process. | 2 |
| 2024 | A Distributed and Parallel Accelerator Design for 3-D Acoustic Imaging on FPGA-Based Systemsabstract3-D imaging sonar is crucial in the exploration of marine resources, and the development of portable device with high imaging quality and high real-time performance is the general trend. However, traditional framework methods are limited by the huge amount of computation brought by high-quality imaging, making it difficult to implement in engineering. To address this issue, we develop 3-D real-time sonar system in an algorithm-hardware co-designed way. An ultrawideband distributed and parallel subarray beamforming algorithm (UWBDPS) is proposed for 3-D acoustic imaging. This is a multi-stage array time-frequency beamforming method under a distributed parallel computing architecture. Based on this, we propose field-programmable gate array (FPGA)-based accelerator. It divides a large sonar receiving planar array into multiple parallel subarrays, and complete the beamforming in two stages, which can reduce the calculation load and speeds up 3-D imaging. For engineering implementation, we optimized the sparseness of the planar transducer array, with a sparse rate as high as 97.7%. The experimental results show that the calculation amount of the proposed UWB-DPS algorithm is reduced to 1/5.7 of the traditional framework algorithm, the imaging performance is effectively improved, and the FPGA-based accelerator outperforms the CPU software implementation by 935×. Weibo Mao, Peng Chen 0008, Yingtian Hu, Haoran Liang 0001, Yuanjie Dang, Ronghua Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | A Visual Representation-Guided Framework With Global Affinity for Weakly Supervised Salient Object DetectionabstractFully supervised salient object detection (SOD) methods have made considerable progress in performance, yet these models rely heavily on expensive pixel-wise labels. Recently, to achieve a trade-off between labeling burden and performance, scribble-based SOD methods have attracted increasing attention. Previous scribble-based models directly implement the SOD task only based on SOD training data with limited information, it is extremely difficult for them to understand the image and further achieve a superior SOD task. In this paper, we propose a simple yet effective framework guided by general visual representations with rich contextual semantic knowledge for scribble-based SOD. These general visual representations are generated by self-supervised learning based on large-scale unlabeled datasets. Our framework consists of a task-related encoder, a general visual module, and an information integration module to efficiently combine the general visual representations with task-related features to perform the SOD task based on understanding the contextual connections of images. Meanwhile, we propose a novel global semantic affinity loss to guide the model to perceive the global structure of the salient objects. Experimental results on five public benchmark datasets demonstrate that our method, which only utilizes scribble annotations without introducing any extra label, outperforms the state-of-theart weakly supervised SOD methods. Specifically, it outperforms the previous best scribble-based method on all datasets with an average gain of 5.5% for max f-measure, 5.8% for mean f-measure, 24% for MAE, and 3.1% for E-measure. Moreover, our method achieves comparable or even superior performance to the state-of-the-art fully supervised models. Binwei Xu, Haoran Liang 0001, Weihua Gong, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Motion-Aware Memory Network for Fast Video Salient Object DetectionabstractPrevious methods based on 3DCNN, convLSTM, or optical flow have achieved great success in video salient object detection (VSOD). However, these methods still suffer from high computational costs or poor quality of the generated saliency maps. To address this, we design a space-time memory (STM)-based network that employs a standard encoder-decoder architecture. During the encoding stage, we extract high-level temporal features from the current frame and its adjacent frames, which is more efficient and practical than methods reliant on optical flow. During the decoding stage, we introduce an effective fusion strategy for both spatial and temporal branches. The semantic information of the high-level features is used to improve the object details in the low-level features. Subsequently, spatiotemporal features are methodically derived step by step to reconstruct the saliency maps. Moreover, inspired by the boundary supervision prevalent in image salient object detection (ISOD), we design a motion-aware loss that predicts object boundary motion, and simultaneously perform multitask learning for VSOD and object motion prediction. This can further enhance the model's capability to accurately extract spatiotemporal features while maintaining object integrity. Extensive experiments on several datasets demonstrate the effectiveness of our method and can achieve state-of-the-art metrics on some datasets. Our proposed model does not require optical flow or additional preprocessing, and can reach an impressive inference speed of nearly 100 FPS. Xing Zhao 0001, Haoran Liang 0001, Guodao Sun, Ronghua Liang, Xiaofei He 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Synthesize Boundaries: A Boundary-Aware Self-Consistent Framework for Weakly Supervised Salient Object DetectionabstractFully supervised salient object detection (SOD) has made considerable progress based on expensive and time-consuming data with pixel-wise annotations. Recently, to relieve the labeling burden while maintaining performance, some scribble-based SOD methods have been proposed. However, learning precise boundary details from scribble annotations that lack edge information is still difficult. In this article, we propose to learn precise boundaries from our designed synthetic images and labels without introducing any extra auxiliary data. The synthetic image creates boundary information by inserting synthetic concave regions that simulate the real concave regions of salient objects. Furthermore, we propose a novel self-consistent framework that consists of a global integral branch (GIB) and a boundary-aware branch (BAB) to train a saliency detector. GIB aims to identify integral salient objects, whose input is the original image. BAB aims to help predict accurate boundaries, whose input is the synthetic image. These two branches are connected through a self-consistent loss to guide the saliency detector to predict precise boundaries while identifying salient objects. Experimental results on five benchmarks demonstrate that our method outperforms the state-of-the-art weakly supervised SOD methods and further narrows the gap with the fully supervised methods. Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 2 |
| 2024 | Sentiment-Aware Representation Learning Framework Fusion With Multi-Aspect Information for POI RecommendationabstractRecently, how to provide better POI recommendation performance by exploiting representation learning frame work is still one challenging task in LBSNs. Most existing methods either focus on modeling few limited information without considering other important information such as social influence and sentiment factor, or only rely on traditional shallow models to combine different factors, lacking effective integrating way to learn latent representations by fully utilizing multi-aspect interaction relations. To address these issues, we propose a novel sentiment-aware POI recommendation framework, dubbed as Senti2LSTM, which is capable of learning more comprehensive representations of users and POIs with fusion of multi-relations. Specifically, we first employ dual LSTMs to capture different sentimental embeddings for users and POIs respectively from emotional comments in LBSNs, and then we integrate them into the aggregation of propagation embeddings for users and POIs when learning their latent representations from user-POI bipartite graph and social link graph in LBSNs. Additionally, we also consider discriminating the importance of different social neighbors by leveraging social based attention mechanism, which makes social friends with common sentiments have more similar preferences.Finally,extensive experimental results conducted on two real-world datasets, e.g., Foursquare-NYC and Yelp2018, have demonstrated the effectiveness of our proposed Senti2LSTM in sentiment learning, and significantly outperforming the state-of-the-art POI recommendation methods Weihua Gong, Genhang Shen, Lianghuai Yang, Haoran Liang 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | A progressive segmentation with weight contrast label enhancement for weakly supervised video salient object detectionabstractAbstract Scribble labels have gained increasing attention in the field of weakly supervised video salient object detection (VSOD). Based on scribble labels, latest methods can spread labeled pixels to unlabeled regions using local coherence loss, but predicted objects often lose detail and boundary information. In this work, a novel method based on back‐foreground weight contrast is proposed that adds label enhancement points to facilitate the model to learn the edge, detail and location of salient object. Additionally, a new VSOD framework based on global structural localization is introduced. Enhanced scribble labels are used to assist the model for global localization, and then the located regions are finely segmented by the trained model. Extensive experiments demonstrate that the method achieves the state‐of‐the‐art performance on common VSOD datasets, with an improvement of 3.75%, 4.68%, and 0.88% in S‐measure, F‐measure, and MAE, respectively. Zelin Lu, Haoran Liang 0001, Binwei Xu, Ronghua Liang |
IET Image Process. | 2 |
| 2022 | CFN: A coarse-to-fine network for eye fixation predictionabstractAbstract Many image‐to‐image computer vision approaches have made great progress by an end‐to‐end framework with the encoder–decoder architecture. However, the same image‐to‐image eye fixation prediction task is not the same as those computer vision tasks in that it focuses more on salient regions rather than precise predictions for every pixel. Thus, it is not appropriate to directly apply the end‐to‐end encoder–decoder to the eye fixation prediction task. In addition, although high‐level feature is important, the contribution of low‐level feature should also be kept and balanced in computational model. Nevertheless, some low‐level features that attract attention are easily neglected while transiting through the deep network. Therefore, the effective way to integrate low‐level and high‐level features for improving eye fixation prediction performance is still a challenging task. In this paper, a coarse‐to‐fine network (CFN) that encompasses two pathways with different training strategies are proposed: coarse perceiving network (CFN‐Coarse) can be a simple encoder network or any of the existing pretrained network to capture the distribution of salient regions and generate high‐quality feature maps; fine integrating network (CFN‐Fine) uses fixed parameters from the CFN‐Coarse and combines features from deep to shallow in the deconvolution process by adding skip connections between down‐sampling and up‐sampling paths to efficiently integrate deep and shallow features. The saliency map obtained by the method is evaluated over 6 standard benchmark datasets, namely SALICON, MIT1003, MIT300, Toronto, OSIE, and SUN500. The results demonstrate that the method can surpass the state‐of‐the‐art accuracy of eye fixation prediction and achieves the competitive performance to date under most evaluation metrics on SALICON Saliency Prediction Challenge (LSUN2017). Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
IET Image Process. | 2 |
| 2021 | Locate Globally, Segment Locally: A Progressive Architecture With Knowledge Review Network for Salient Object DetectionabstractSalient object location and segmentation are two different tasks in salient object detection (SOD). The former aims to globally find the most attractive objects in an image, whereas the latter can be achieved only using local regions that contain salient objects. However, previous methods mainly accomplish the two tasks simultaneously in a simple end-to-end manner, which leads to the ignorance of the differences between them. We assume that the human vision system orderly locates and segments objects, so we propose a novel progressive architecture with knowledge review network (PA-KRN) for SOD. It consists of three parts. (1) A coarse locating module (CLM) that uses body-attention label locates rough areas containing salient objects without boundary details. (2) An attention-based sampler highlights salient object regions with high resolution based on body-attention maps. (3) A fine segmenting module (FSM) finely segments salient objects. The networks applied in CLM and FSM are mainly based on our proposed knowledge review network (KRN) that utilizes the finest feature maps to reintegrate all previous layers, which can make up for the important information that is continuously diluted in the top-down path. Experiments on five benchmarks demonstrate that our single KRN can outperform state-of-the-art methods. Furthermore, our PA-KRN performs better and substantially surpasses the aforementioned methods. Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
AAAI | 2 |
| 2021 | VSumVis: Interactive Visual Understanding and Diagnosis of Video Summarization ModelabstractWith the rapid development of mobile Internet, the popularity of video capture devices has brought a surge in multimedia video resources. Utilizing machine learning methods combined with well-designed features, we could automatically obtain video summarization to relax video resource consumption and retrieval issues. However, there always exists a gap between the summarization obtained by the model and the ones annotated by users. How to help users understand the difference, provide insights in improving the model, and enhance the trust in the model remains challenging in the current study. To address these challenges, we propose VSumVis under a user-centered design methodology, a visual analysis system with multi-feature examination and multi-level exploration, which could help users explore and analyze video content, as well as the intrinsic relationship that existed in our video summarization model. The system contains multiple coordinated views, i.e., video view, projection view, detail view, and sequential frames view. A multi-level analysis process to integrate video events and frames are presented with clusters and nodes visualization in our system. Temporal patterns concerning the difference between the manual annotation score and the saliency score produced by our model are further investigated and distinguished with sequential frames view. Moreover, we propose a set of rich user interactions that enable an in-depth, multi-faceted analysis of the features in our video summarization model. We conduct case studies and interviews with domain experts to provide anecdotal evidence about the effectiveness of our approach. Quantitative feedback from a user study confirms the usefulness of our visual system for exploring the video summarization model. Guodao Sun, Chaoqing Xu, Haoran Liang 0001, Binwei Xu, Ronghua Liang |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2020 | Video summarisation with visual and semantic cuesabstractVideo summarisation greatly improves the efficiency of people browsing videos and saves storage space. A good video summary should satisfy human visual interestingness and preserve the theme of the original video at the semantic level. Unlike many existing methods that consider only visual features to generate video summaries, this study proposes a method that combines visual and semantic cues to extract important information for dynamic video summarisation. The authors propose visual‐verbal saliency consistency to add semantic information and propose a novel attention motion, along with other visual features to fully represent visual interestingness. Based on the importance score of each frame calculated by combining these features, they select an optimal subset of segments to generate an important and interesting summary. They evaluate their method using the SumMe and TVSum datasets and experimental results show that their method generates high‐quality video summaries. Binwei Xu, Haoran Liang 0001, Ronghua Liang |
IET Image Process. | 2 |
| 2020 | A structure-guided approach to the prediction of natural image saliency
Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
Neurocomputing | 1 |
| 2019 | CapVis: Toward Better Understanding of Visual-Verbal Saliency ConsistencyabstractWhen looking at an image, humans shift their attention toward interesting regions, making sequences of eye fixations. When describing an image, they also come up with simple sentences that highlight the key elements in the scene. What is the correlation between where people look and what they describe in an image? To investigate this problem intuitively, we develop a visual analytics system, CapVis, to look into visual attention and image captioning, two types of subjective annotations that are relatively task-free and natural. Using these annotations, we propose a word-weighting scheme to extract visual and verbal saliency ranks to compare against each other. In our approach, a number of low-level and semantic-level features relevant to visual-verbal saliency consistency are proposed and visualized for a better understanding of image content. Our method also shows the different ways that a human and a computational model look at and describe images, which provides reliable information for a captioning model. Experiment also shows that the visualized feature can be integrated into a computational model to effectively predict the consistency between the two modalities on an image dataset with both types of annotations. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2017 | Saliency prediction with scene structural guidanceabstractPrevious works have suggested the role of scene information in directing gaze. The structure of a scene provides global contextual information that complements local object information in saliency prediction. In this study, we explore how scene envelopes such as openness, depth, and perspective affect visual attention in natural outdoor images. To facilitate this study, an eye tracking dataset is first built with 500 natural scene images and eye tracking data with 15 subjects free-viewing the images. We make observations on scene layout properties and propose a set of scene structural features relating to visual attention. We further integrate features from deep neural networks and use the set of complementary features for saliency prediction. Our features are independent of and can work together with many computational modules, and this work demonstrates the use of Multiple kernel learning (MKL) as an example to integrate the features at low- and high-levels. Experimental results demonstrate that our model outperforms existing methods and our scene structural features can improve the performance of other saliency models in outdoor scenes. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
SMC | 1 |
| 2017 | Visual-verbal consistency of image saliencyabstractWhen looking at an image, humans shift their attention towards interesting regions, making sequences of eye fixations. When describing an image, they also come up with simple sentences that highlight the key elements in the scene. What is the correlation between where people look and what they describe in an image? To investigate this problem, we look into eye fixations and image captions, two types of subjective annotations that are relatively task-free and natural. From the annotations, we extract visual and verbal saliency ranks to compare against each other. We then propose a number of low-level and semantic-level features relevant to the visualverbal consistency. Integrated into a computational model, the proposed features effectively predict the consistency between the two modalities on a large dataset with both types of annotations, namely SALICON [1]. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
SMC | 1 |
| 2016 | Coupled Dictionary Learning for the Detail-Enhanced Synthesis of 3-D Facial ExpressionsabstractThe desire to reconstruct 3-D face models with expressions from 2-D face images fosters increasing interest in addressing the problem of face modeling. This task is important and challenging in the field of computer animation. Facial contours and wrinkles are essential to generate a face with a certain expression; however, these details are generally ignored or are not seriously considered in previous studies on face model reconstruction. Thus, we employ coupled radius basis function networks to derive an intermediate 3-D face model from a single 2-D face image. To optimize the 3-D face model further through landmarks, a coupled dictionary that is related to 3-D face models and their corresponding 3-D landmarks is learned from the given training set through local coordinate coding. Another coupled dictionary is then constructed to bridge the 2-D and 3-D landmarks for the transfer of vertices on the face model. As a result, the final 3-D face can be generated with the appropriate expression. In the testing phase, the 2-D input faces are converted into 3-D models that display different expressions. Experimental results indicate that the proposed approach to facial expression synthesis can obtain model details more effectively than previous methods can. Haoran Liang 0001, Ronghua Liang, Mingli Song, Xiaofei He 0001 |
IEEE Trans. Cybern. | 1 |
| 2016 | Looking Into Saliency Model via Space-Time VisualizationabstractWe introduce a visual analytics method to analyze eye-tracking data and saliency models for dynamic stimuli, such as video or animated graphics. The focus lies on the analysis of the different performance of saliency models in contrast to human observers to identify trends in the general viewing behavior, including time sequences of attentional synchrony and objects with a strong attentional focus. By using a space-time cube visualization in combination with clustering, the dynamic stimuli and associated eye gazes as well as the attention maps from saliency models can be analyzed in a static three-dimensional representation. We propose algorithms to keep the appearance of the computer's attention data in line with the human's eye-tracking data. The analytical process is supported by multiple coordinated views that allow the user to focus on different aspects of spatial and temporal information in eye gaze data and saliency map. By comparing attention data from both human and computer incorporated with the spatiotemporal characteristics, we are able to find the different patterns within human and computer algorithms. We list our key findings to help developing better saliency detection algorithms. Haoran Liang 0001, Ronghua Liang, Guodao Sun |
IEEE Trans. Multim. | 1 |