Shiguang Liu

dblp:58/5351 · DBLP profile ↗
← Back
93ranked-venue papers
29as first author
48since 2021 · last 2026
0000-0003-2353-5318ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 83 · 27 first-author · 41 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Computer networks · 5 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Visual simulation of fruit chilling injury
abstract
Chilling injury (CI) is a major postharvest physiological disorder in fruits, causing quality degradation and economic losses during low-temperature storage. While physically-based methods exist for simulating plant deformation, they are computationally intensive and not optimized for capturing the subtle, spatially distributed symptoms of CI, such as browning, pitting, and wrinkling. In this paper, we propose a biologically-informed, texture-based framework for dynamic CI simulation that links biological symptom progression to visual representation. Browning and pitting are modeled using a texture-based de-chilling technique driven by a Logistic model of the chilling injury index (CII), with a histogram-matching-based algorithm ensuring alignment between simulated symptoms and CII values. Wrinkling is simulated by combining kinetic models of water loss with bump maps generated using Worley noise, which approximate the quasi-random yet locally continuous surface deformations caused by epidermal shrinkage. The proposed framework efficiently integrates biologically-driven modeling, dynamic texture evolution, and water-loss-induced surface deformation, producing realistic CI simulations without high-resolution meshes. It applies to multiple fruit types — including tropical climacteric (banana), Solanaceous (tomato), Cucurbit (cucumber), and Citrus (orange, lemon) — offering an effective approach for visualizing CI progression.
Shiguang Liu
Graph. Model.2
2026 DESTalker : Disentangling Emotion and Style for Expressive 3D Facial Animation via Residual Generation
abstract
ABSTRACT Speech‐driven 3D facial animation has achieved remarkable lip‐synchronization. However, generating expressive and personalized animations remains a formidable challenge. Existing methods often struggle with the entanglement of emotion and speaking style, as well as semantic conflicts arising from multi‐modal inputs. To address these limitations, we propose DESTalker , a framework that disentangles emotion and style to generate expressive facial motion via residual modulation. We introduce an information‐theoretic disentanglement strategy utilizing mutual information minimization and orthogonal constraints to robustly separate audio content, motion emotion, and speaking style. Furthermore, we adopt a residual generation strategy where neutral lip‐sync motion is synthesized first, followed by the injection of emotion and style as residual components to preserve articulatory precision. Finally, to resolve cross‐modal conflicts, we design a semantic consistency‐guided soft‐gating mechanism that dynamically regulates visual emotion integration under audio dominance. Extensive experiments on the 3DMEAD dataset demonstrate that DESTalker outperforms state‐of‐the‐art methods in emotional expressiveness, style consistency, and lip‐synchronization accuracy.
Shiguang Liu
Comput. Animat. Virtual Worlds2
2026 Adaptive Geometric Attention-Driven No-Reference Multi-Modal Point Cloud Quality Assessment
abstract
Point clouds play an essential role in 3D visual media applications. Point Cloud Quality Assessment (PCQA) is vital to improving the subjective experience, but lacking an effective geometric characterization hinders its consistency with subjective perception. To tackle this problem, this article introduces a No-Reference Multi-Modal Point Cloud Quality Assessment (NR-PCQA) approach driven by adaptive geometric attention. Specifically, this article defines two dimensionless geometric descriptors, namely Radial Depth Ratio (RDR) and Relative Radial Distance (RRD). These two descriptors are then used to construct an Adaptive Geometric Attention Mechanism (AGAM), which dynamically guides the network to focus on features related to geometric quality during feature extraction. Based on AGAM, a multi-modal fusion framework is further built, where the 3D geometric cues are combined with the texture and semantic information of the 2D projection images, leading to accurate PCQA. Creatively, the Hierarchical Multi-Modal Attention Fusion (HMAF) mechanism is designed to achieve the complementary strengths of 3D and 2D features, which first maximizes the extraction of single-modal features and then deeply merges cross-modal information. Naturally, the experimental results on SJTU-PCQA and WPC databases demonstrate that the innovative design of AGAM and HMAF achieves effective multi-modal feature fusion on the basis of satisfactory description for geometric properties, resulting in higher subjective consistency than the state-of-the-art methods. Meanwhile, the proposed method exhibits strong robustness across various types of point cloud distortions and diverse point cloud content, providing comprehensive validation of its effectiveness and practicality.
Ziqing Huang, Shiguang Liu
ACM Trans. Multim. Comput. Commun. Appl.7
2025 Dynamic Style-Adaptive Image Transformation Network
Hanadi Al-Mekhlafi, Shiguang Liu
CGI (1)2
2025 Bi-IRNet: A Transformer-Based Binaural Impulse Response Generation Guidance Model
Yisheng Zhang, Shiguang Liu
ICIG (1)2
2025 Fitted-Singer: Singing Voice Synthesis with Style Control and Rhythm Control
abstract
Using text prompts to control rhythm features in singing voice synthesis (SVS) offers a convenient method for non-professional musicians to generate target voice. However, due to the diversity and variability of voice, timber characteristics are challenging to describe accurately with natural language. To overcome this challenge, reference audio is introduced as input for controlling the timber style of synthesized voice. To effectively extract detailed stylistic features from short audio clips, a Cross-Fusion Encoder is proposed. Additionally, advanced GPT-4 analyzes rhythm changes from text prompts to directly modify rhythm features, enabling text control of synthesized singing rhythm. We propose Fitted-Singer, which allows users to control the rhythm of singing voice through natural language and control the style of voice using reference audio, meeting the demand for more personalized singing voice generation. Experimental results demonstrate that Fitted-Singer excels in generating high-quality singing voice that meets target requirements.
Shiguang Liu
ICME3
2025 Visual Localization with Offline Google Satellite Map-Assisted for Ground Vehicles in GNSS-Denied Environment
abstract
Vehicle localization is a critical component in the planning and navigation of autonomous driving system. Generally, traditional vehicle localization methods rely on the Global Navigation Satellite System (GNSS) for self-localization. Unfortunately, GNSS can become unreliable and may fail in urban canyons, under trees, and beneath overpasses. To address this problem, we propose a visual localization framework assisted by offline Google satellite maps in GNSS-weak or GNSS-denied environments. And we introduce learning-based ground-to-satellite map feature matching method to mitigate the long-term cumulative drift of visual odometry. To reduce the negative impact of cross-view matching errors on localization accuracy, we propose a novel cross-view pose selection method to build two pose uncertainty models. Moreover, we combine the proposed method with classical SLAM methods to develop a vehicle localization framework. To verify the performance of the proposed method, we carried out the accuracy comparison experiment with state-of-the-art fusion localization methods and feature matching methods. Experimental results indicate that the proposed method achieves the best localization performance compared with the state-of-the-art methods, and our method achieves the root mean square error of 0.290m and 0.014rad in KITTI-05. The implementation code of this paper will be open-source at https://github.com/NEU-REAL/visualLocalization-with-satelliteMap.
Jibo Wang, Bairen Mao, Chenglin Pang, Shiguang Liu, Jindi Guo, Zheng Fang 0001
IROS4
2025 FocalFormer: Leveraging focal modulation for efficient action segmentation in egocentric videos
Jialu Xi, Shiguang Liu
Comput. Graph.2
2025 Texture dominated no-reference quality assessment for high resolution image by multi-scale mechanism
Ziqing Huang, Shiguang Liu
Neurocomputing6
2025 Botanical-Based Simulation of Fruit Shape Change During Growth
abstract
ABSTRACT Fruit growth is an interesting time‐lapse process. The simulation of this process using computer graphics technology can have many applications in areas such as films, games, agriculture, etc. Although there are some methods to model the shape of the fruit, it is challenging to accurately simulate its growth process and include shape changes. We propose a botanical‐based framework to address this problem. By combining the growth pattern function and the exponential model in botany, we propose a mesh scaling method that can accurately simulate the fruit volume increase. Specifically, the RGR (relative growth rate) in the exponential model is automatically calculated according to the user's input growth pattern function or real size data. In addition, we model and simulate fruit shape changes by integrating axial, longitudinal, and latitudinal shape parameters into the RGR function. Various defective fruits can be simulated by adjusting these parameters. Inspired by the principle of root curvature, we propose a deformation technique‐based approach in conjunction with our volume increase approach to simulate the bending growth of fruits such as cucumber. Various experiments show that our framework can effectively simulate the growth process of a wide range of fruits with shape change or bending.
Shiguang Liu
Comput. Animat. Virtual Worlds2
2025 PredVSD: Video saliency prediction based on conditional diffusion model
Shiguang Liu
Knowl. Based Syst.2
2025 TM2SP: A Transformer-Based Multi-Level Spatiotemporal Feature Pyramid Network for Video Saliency Prediction
abstract
This paper proposes an end-to-end video saliency prediction network model, termed TM2SP-Net (Transformer-based Multi-level Spatiotemporal Feature Pyramid Network). Leveraging the strong encoding learning capability of Video Swin Transformer for video data, we design a Multi-level Spatiotemporal Feature Pyramid Network (MLSTFPN) that effectively detects and enriches salient regions and spatial details across different scales. In particular, a pre-trained image saliency detection encoder is employed to extract salient features from each frame, serving as prior knowledge to guide the multi-scale spatiotemporal feature fusion and decoding processes. Additionally, we introduce an Inception Gate-Controlled Fusion (IGCF) and Layered Self-Attention Aggregation Fusion (LSAF) mechanisms to efficiently merge spatiotemporal features across various stages. Finally, extensive experiments conducted on the DHF1K, Hollywood-2, UCF-Sports, and six audio-visual saliency datasets demonstrate the superiority of our method over existing state-of-the-art approaches.
Shiguang Liu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Structure-Aware Style Transfer Based on Multi-Scale and Edge Texture
Hanadi Al-Mekhlafi, Shiguang Liu
ADMA (5)2
2024 MTIQA360: An Easily Trainable Multitasking Network for Blind Omnidirectional Image Quality Assessment
Qinghai Wang, Shiguang Liu
ICPR (20)2
2024 Artistic Style Transfer Based on Attention with Knowledge Distillation
abstract
Abstract Artistic style transfer involves the adaption of an input image to reflect the style of a reference image while maintaining its original content. This technique, now a prominent focus due to its prospective use in creative fields like digital art and graphic design, typically applies normalization techniques and attention mechanisms. While these methods yield decent results, they often fall short due to distortion of content image details and non‐artefact styles. In this paper, we introduce a novel approach that synergizes adaptive instance normalization (AdaIN), attention mechanisms, knowledge distillation (KD) and strategically placed internal layers, and new enhancements designed to preserve content details and provide a nuanced control over the style transfer process. We introduce a Detail Enhancement Module to amplify high‐frequency details in the content image, enhancing edge and texture preservation. A Multi‐scale Strategy is implemented to ensure uniform style application across various detail levels, leading to more coherent stylization. The Content Feature Refinement process refines content features, sharpening and emphasizing details to preserve structural and textural integrity. AdaIN's distinctive feature of efficiently collecting style data is exploited in our approach, coupled with attention mechanisms' inherent ability to conserve content information. We supplement this blend with KD for the enhancement of model accuracy and efficiency. Additionally, the introduction of internal layers acts as a conduit to further improve the style transfer process, increasing the transfer level of features and fostering better stylized results. The cornerstone of our technique lies in preserving the content structure amidst complex style transfers. Experimental results affirm the superior performance of our method over existing techniques in both quantitative and qualitative evaluations.
Hanadi Al-Mekhlafi, Shiguang Liu
Comput. Graph. Forum2
2024 Single image super-resolution: a comprehensive review and recent insight
Hanadi Al-Mekhlafi, Shiguang Liu
Frontiers Comput. Sci.2
2024 A multi-species material point method with a mixture model
abstract
Abstract The material point method (MPM) has attracted more and more attention in computer graphics. It is very successful in simulating both fluid flow and solid deformation, but may fail in simulating multiple fluids and solids coupling. We propose a unified MPM solver for multi‐species simulations. Compared to traditional MPM, we extend the degree of freedom on background grid to store information of multiple materials, so that our framework is able to deal with multiple materials well. The proposed method leverages the advantages of MPM as a hybrid method. We introduce the mixture model into the framework, which was the most widely used for grid‐based multi‐fluid flows. This enables MPM to capture the interaction and relative motion, and animates complex and coupled fluids and solids in a unified manner. A series of experiments are presented to demonstrate effectiveness of our method.
Shiguang Liu
Comput. Animat. Virtual Worlds2
2024 Multi-scale edge aggregation mesh-graph-network for character secondary motion
abstract
Abstract As an enhancement to skinning‐based animations, light‐weight secondary motion method for 3D characters are widely demanded in many application scenarios. To address the dependence of data‐driven methods on ground truth data, we propose a self‐supervised training strategy that is free of ground truth data for the first time in this domain. Specifically, we construct a self‐supervised training framework by modeling the implicit integration problem with steps as an optimization problem based on physical energy terms. Furthermore, we introduce a multi‐scale edge aggregation mesh‐graph block (MSEA‐MG Block), which significantly enhances the network performance. This enables our model to make vivid predictions of secondary motion for 3D characters with arbitrary structures. Empirical experiments indicate that our method, without requiring ground truth data for model training, achieves comparable or even superior performance quantitatively and qualitatively compared to state‐of‐the‐art data‐driven approaches in the field.
Shiguang Liu
Comput. Animat. Virtual Worlds2
2024 Botanical-based simulation of color change in fruit ripening: Taking tomato as an example
abstract
Abstract The color change of plant fruit in ripening is a typical time‐varying phenomenon involving various factors. Due to its complexity and biodiversity, it is challenging to model this phenomenon. To address this issue, we take the tomato as an example and propose a botanical‐based framework considering variety, environment, phytohormone, and genes to simulate fruit color change during the ripening process. Specifically, we propose a first‐order kinetic model that integrates varietal, environmental, and phytohormonal factors to represent the variation of pigment concentrations in the pericarp. Moreover, we introduce a logistic model to describe the change in pigment concentration in the epidermis. Based on the gene expression pathway of tomato color in botany, we propose a genotype‐to‐phenotype simulation method to represent its biodiversity. An improved method is proposed to convert pigment concentrations into color accurately. Furthermore, we propose a gradient descent‐based method to assist the user in quickly setting pigment concentration parameters. Experiments verified that the proposed framework can simulate a wide range of tomato colors. Both qualitative and quantitative experiments validated the proposed method. Furthermore, our framework can be applied to more fruits.
Shiguang Liu
Comput. Animat. Virtual Worlds2
2024 UnderwaterImage2IR: Underwater impulse response generation via dual-path pre-trained networks and conditional generative adversarial networks
abstract
Abstract In the field of acoustic simulation, methods that are widely applied and have been proven to be highly effective rely on accurately capturing the impulse response (IR) and its convolution relationship. This article introduces a novel approach, named as UnderwaterImage2IR, that generates acoustic IRs from underwater images using dual‐path pre‐trained networks. This technique aims to achieve cross‐modal conversion from underwater visual images to acoustic information with high accuracy at a low cost. Our method utilizes deep learning technology by integrating dual‐path pre‐trained networks and conditional generative adversarial networks conditional generative adversarial networks (CGANs) to generate acoustic IRs that match the observed scenes. One branch of the network focuses on the extraction of spatial features from images, while the other is dedicated to recognizing underwater characteristics. These features are fed into the CGAN network, which is trained to generate acoustic IRs corresponding to the observed scenes, thereby achieving high‐accuracy acoustic simulation in an efficient manner. Experimental results, compared with the ground truth and evaluated by human experts, demonstrate the significant advantages of our method in generating underwater acoustic IRs, further proving its potential application in underwater acoustic simulation.
Yisheng Zhang, Shiguang Liu
Comput. Animat. Virtual Worlds2
2024 Music conditioned 2D hand gesture dance generation with HGS
abstract
Abstract In recent years, the short video industry is booming. However, there are still many difficulties in the action generation of virtual characters. We observed that on the short video social platform, “hand gesture dance” is a very popular short video form. However, its development is limited by the professionalism of choreography. In order to solve these problems, we propose an intelligent choreography framework, which can generate new gesture sequences for unseen audio based on pairing data in the database. Our framework adopts multimodal method and obtains excellent results. In additional, we collected and produced the first and largest pair labeled hand gesture dance data set. Various experiments showed that our results not only generate smooth and rich action sequences, but also collect some semantic information contained in the audio.
Dian Zhou, Shiguang Liu, Qing Xu 0002
Comput. Animat. Virtual Worlds2
2024 Hierarchical Equalization Loss for Long-Tailed Instance Segmentation
abstract
Multimedia data has the characteristics of large scale and skewed distribution with a long-tailed shape, which is a challenging imbalance problem faced by deep learning. In long-tailed image instance segmentation, the existing methods deal with this imbalance problem from a single perspective, ignoring the presence of multiple imbalance factors, which results in the limitation of performance. Considering that imbalances exist not only between positive and negative classes, but also between foreground and background subclasses, as well as between hard and easy examples, we argue that the losses of samples should be hierarchically equalized at multi-levels (HEL). In line with this idea, we first propose a focus based hierarchical-equalization loss (FHEL), which employs a class gradient ratio based reweighting mechanism to achieve the balance between classes, and uses a subclass-balance term and a sample-balance term to separately deal with the inter-subclass and inter-sample imbalances. FHEL can improve the performance of long-tailed instance segmentation in an end-to-end manner, avoiding the overfitting risk and manual hard division in the traditional methods. On the basis of FHEL, we further explore the relationship between inter-subclass imbalance and inter-sample imbalance, and propose a constrained-focus based hierarchical-equalization loss (CFHEL) that copes with the imbalances at multi-levels simultaneously with fewer hyperparameters. CFHEL is effective and easy to tune hyperparameters. We conduct extensive experiments on LVIS v1.0 and COCO-LT datasets with different benchmarks. Both FHEL and CFHEL are superior to the existing methods. On LVIS v1.0, with ResNet50 Mask R-CNN, ResNet101Mask R-CNN, ResNeXt101 Mask R-CNN and ResNet101 Cascade Mask R-CNN, CFHEL outperforms its baselines respectively with 19.8%, 18.5%, 21.6% and 21.2% AP% gains, and with 6.7%, 6.6% and 6.5% AP gains, achieving the new state-of-the-arts. On COCO-LT, our CFHEL outperforms the baseline with 13.2% tail AP gains and 3.3% whole AP gains, also achieving the new best performances.
Yaochi Zhao, Shiguang Liu, Zhuhua Hu, Jingwen Xia
IEEE Trans. Multim.3
2023 SemiRefiner: Learning to Refine Semi-realistic Paintings
Keyue Fan, Shiguang Liu
CGI2
2023 MFSFFuse: Multi-receptive Field Feature Extraction for Infrared and Visible Image Fusion Using Self-supervised Learning
Xueyan Gao, Shiguang Liu
ICONIP (6)2
2023 Lightweight Scene-aware Rain Sound Simulation for Interactive Virtual Environments
abstract
We present a lightweight and efficient rain sound synthesis method for interactive virtual environments. Existing rain sound simulation methods require massive superposition of scene-specific precomputed rain sounds, which is excessive memory consumption for virtual reality systems (e.g. video games) with limited audio memory budgets. Facing this issue, we reduce the audio memory budgets by introducing a lightweight rain sound synthesis method which is only based on eight physically-inspired basic rain sounds. First, in order to generate sufficiently various rain sounds with limited sound data, we propose an exponential moving average based frequency domain additive (FDA) synthesis method to extend and modify the pre-computed basic rain sounds. Each rain sound is generated in the frequency domain before conversion back to the time domain, allowing us to extend the rain sound which is free of temporal distortions and discontinuities. Next, we introduce an efficient binaural rendering method to simulate the 3D perception that coheres with the visual scene based on a set of Near-Field Transfer Functions (NFTF). Various results demonstrate that the proposed method drastically decreases the memory cost (77 times compressed) and overcomes the limitations of existing methods in terms of interaction.
Haonan Cheng, Shiguang Liu, Jiawan Zhang
VR2
2023 Realistic simulation of fruit mildew diseases: Skin discoloration, fungus growth and volume shrinkage
abstract
Time-varying effects simulation plays a critical role in computer graphics. Fruit diseases are typical time-varying phenomena. Due to the biological complexity, the existing methods fail to represent the biodiversity and biological law of symptoms. To this end, this paper proposes a biology-aware, physically-based framework that respects biological knowledge for realistic simulation of fruit mildew diseases. The simulated symptoms include skin discoloration, fungus growth, and volume shrinkage. Specifically, we take advantage of both the zero-order kinetic model and reaction–diffusion model to represent the complex fruit skin discoloration related to skin biological characteristics. To reproduce 3D mildew growth, we employ the Poisson-disk sampling technique and propose a template model instancing method. One can flexibly change hyphal template models to characterize the fungal biological diversity. To model the fruit’s biological structure, we fill the fruit mesh interior with particles in a biologically-based arrangement. Based on this structure, we propose a turgor pressure and a Lennard-Jones force-based adaptive mass–spring system to simulate the fruit shrinkage in a biological manner. Experiments verified that the proposed framework can effectively simulate mildew diseases, including gray mold, powdery mildew, and downy mildew. Our results are visually compelling and close to the ground truth. Both quantitative and qualitative experiments validated the proposed method.
Shiguang Liu
Graph. Model.2
2023 Talking Face Generation via Facial Anatomy
abstract
To generate the corresponding talking face from a speech audio and a face image, it is essential to match the variations in the facial appearance with the speech audio in subtle movements of different face regions. Nevertheless, the facial movements generated by the existing methods lack detail and vividness, or the methods are only oriented toward a specific person. In this article, we propose a novel two-stage network to generate talking faces for any target identity through annotations of the action units (AUs). In the first stage, the relationship between the audio and the AUs in the audio-to-AU network is learned. The audio-to-AU network needs to produce the consistent AU group for the input audio. In the second stage, the AU group in the first stage and a face image are fed into the generation network to output the resulting talking face image. Various results confirm that, compared to state-of-the-art methods, our approach is able to produce more realistic and vivid talking faces for arbitrary targets with richer details of facial movements, such as the cheek motion and eyebrow motion.
Shiguang Liu, Huixin Wang
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Generating Talking Face With Controllable Eye Movements by Disentangled Blinking Feature
abstract
In virtual reality, talking face generation is committed to using voice and face images to generate real face speech videos to improve the communication experience in the case of limited user information exchange. In a real video, blinking is an action often accompanied by speech, and it is also one of the indispensable actions in real face speech videos. However, the current methods either do not pay attention to the generation of eye movements, or cannot control the blinking in the generated results. To this end, this article proposes a novel system which produces vivid talking face with controllable eye blinks driven by the joint features including identity feature, audio feature, and blink feature. In order to disentangle the blinking action, we designed three independent features to individually drive the main components in the generated frame, namely the facial appearance, mouth movements, and eye movements. Through the adversarial training of the identity encoder, we filter out the information of the eye state from the identity feature, thereby strengthening the independence of the blinking feature. We introduced the blink score as the leading information of the blink feature, and through training, the value can be consistent with human perception to form a complete and independent control of the eyes. Experimental results on multiple datasets show that our method can not only reproduce real talking faces, but also ensure that the blinking pattern and time are fully controllable.
Shiguang Liu, Jiaqi Hao
IEEE Trans. Vis. Comput. Graph.1
2022 SemiPainter: Learning to Draw Semi-realistic Paintings from the Manga Line Drawings and Flat Shadow
Keyue Fan, Shiguang Liu, Wenhuan Lu
CGI2
2022 No-reference Omnidirectional Image Quality Assessment Based on Joint Network
abstract
In panoramic multimedia applications, the perception quality of the omnidirectional content often comes from the observer's perception of the viewports and the overall impression after browsing. Starting from this hypothesis, this paper proposes a deep-learning based joint network to model the no-reference quality assessment of omnidirectional images. On the one hand, motivated by different scenarios that lead to different human understandings, a convolutional neural network (CNN) is devised to simultaneously encode the local quality features and the latent perception rules of different viewports, which are more likely to be noticed by the viewers. On the other hand, a recurrent neural network (RNN) is designed to capture the interdependence between viewports from their sequence representation, and then predict the impact of each viewport on the observer's overall perception. Experiments on two popular omnidirectional image quality databases demonstrate that the proposed method outperforms the state-of-the-art omnidirectional image quality metrics.
Chaofan Zhang, Shiguang Liu
ACM Multimedia2
2022 Line drawing via saliency map and ETF
Shiguang Liu
Frontiers Comput. Sci.1
2022 Focal learning on stranger for imbalanced image segmentation
abstract
Abstract It is an open issue to train effective deep network models on class imbalance datasets. In the widely used cost‐sensitive imbalanced learning methods, the costs are based on the losses or class probabilities of samples. In this paper, it is discovered that these traditional cost‐sensitive methods discard the clustering feature, and introduce the errors of annotations into costs, leading to sub‐optimal models. It is further investigated that the feature magnitude of sample, which is computed before probability and loss, not only is independent of the annotation, but also represents the familiarity degree of model with the sample. These characteristics of feature magnitude are used to guide the training and inference of model. First, the concept of stranger is proposed, which is the sample with small feature magnitude value, and the idea of focal learning on strangers (FLS) is proposed. By adding the idea of FLS into two existing cost‐sensitive methods, two novel losses are put forward: instance‐level focal stranger loss (IFSL) and class‐level focal stranger loss (CFSL). The losses can improve the aggregation features of samples within class, and reduce the negative influences of annotation errors on imbalanced learning. Second, considering the large difference of feature magnitude means between minority class and majority class in case of extreme class‐imbalance dataset, a bias determination (BD) strategy is put forward to improve classification performance during inference. The methods are applied to the tasks of image semantic segmentation and salient‐instance segmentation. The experimental results on four public semantic segmentation datasets demonstrate that IFSL can reduce the over‐fitting of model, improve the classification accuracy of rare samples, and alleviate the reliance of performance on the annotation quality. The experimental results on two public salient‐instance segmentation datasets show that CFSL makes the model have better scoring ability for salient object. Besides, the BD strategy can reduce the wrong classification caused by bias model. Therefore, the proposed methods can significantly advance image segmentation.
Yaochi Zhao, Shiguang Liu, Zhuhua Hu
IET Image Process.2
2022 Style enhanced line drawings based on multi-feature
Shiguang Liu
Multim. Tools Appl.1
2022 QL-IQA: Learning distance distribution from quality levels for blind image quality assessment
Ziqing Huang, Shiguang Liu
Signal Process. Image Commun.3
2022 Towards an End-to-End Visual-to-Raw-Audio Generation With GAN
abstract
Automatically synthesizing sounds for different visual contents poses a challenge and there is a strong need to facilitate the direct creation of realistic sounds. Different from previous works, in this paper, we propose a novel deep learning based approach, which formulates sound simulation as a regression problem. This allows us to circumvent the complexity of the acoustic theory by a novel, general-purpose neural sound synthesis (V2RA) network. Moreover, the end-to-end architecture of V2RA ensures full training without any extra inputs, which thereby greatly improves the scalability and reusability over previous works. In contrast to conventional visual-to-audio generation methods, the V2RA problem is established and solved by generative adversarial networks (GANs). Furthermore, our network architecture can directly predict synchronized raw audio signals (unlike most existing approaches that handle the audio through spectrograms) and generate sound in real time. To evaluate the performance of the neural network generator, we specifically introduce two quantitative scores. Various experiments demonstrate that our V2RA network can produce compelling sound results, which thus provides a viable solution for applications such as sound design and dubbing.
Shiguang Liu, Haonan Cheng
IEEE Trans. Circuits Syst. Video Technol.1
2022 Dual-Channel Multi-Task CNN for No-Reference Screen Content Image Quality Assessment
abstract
Nowadays the problem of image quality assessment (IQA) for screen content images (SCIs) has become a research hotspot as they are ubiquitous in multimedia applications. Although the quality assessment of natural images (NIs) has been continuously developed in the past few decades, few NI-oriented IQA methods can be directly applied on SCIs due to different visual characteristics between them. In this paper, we present a no-reference quality prediction approach considering the content information of SCIs, which is based on dual-channel multi-task convolutional neural network. First, we segment a SCI into small patches and classify them as the textual patches and the pictorial patches. Then, we devise a novel dual-channel convolutional neural network (CNN) to predict the quality of textual patches and pictorial patches. Finally, we propose an effective adaptive weighting strategy for quality score aggregation. The proposed CNN is built on an end-to-end multi-task learning framework, which assists the SCI quality prediction task through the histogram of oriented gradient (HOG) feature prediction task to learn a better mapping between the input patch and its quality score. The adaptive weighting strategy further improves the representation ability of each SCI patch. Experimental results on two largest SCI-oriented databases demonstrate that the proposed method outperforms most of the state-of-the-art no-reference IQA methods and the full-reference IQA methods.
Chaofan Zhang, Ziqing Huang, Shiguang Liu, Jian Xiao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Structure-Guided Arbitrary Style Transfer for Artistic Image and Video
abstract
Recently, neural style transfer has become a popular task in both academic research and industrial applications. Although the existing methods made great progress in terms of quality and efficiency, most of them mainly focus on extracting high-level features. Therefore, it is still challenging to display the hierarchical structure of the content image due to lack of texture information, which causes blurred boundaries and distortion of the stylized image. In this paper, a novel neural image and video style transfer scheme is proposed to suppress distortion and preserve the semantic content of the content image, which is capable of yielding satisfactory stylized images and videos of a variety of scenarios. We first propose to assemble a refine network into an auto-encoder framework to guide style transfer, which can ensure that the stylized image have diverse levels of details. Then, we introduce the global content loss and the local region structure loss to train the model and enhance the robustness of the model. In addition, in order to produce a high-quality stylized video, our method not only preserves the image structure, but also introduces a temporal consistency loss and a cycle-temporal loss to avoid temporal incoherence and motion blur as far as possible. Our approach is also friendly for photographic and exposed image and video style transfer. Both quantitative and qualitative evaluation demonstrated the effectiveness of our method.
Shiguang Liu
IEEE Trans. Multim.1
2022 Facial-expression-aware Emotional Color Transfer Based on Convolutional Neural Network
abstract
Emotional color transfer aims to change the evoked emotion of a source image to that of a target image by adjusting color distribution. Most of existing emotional color transfer methods only consider the low-level visual features of an image and ignore the facial expression features when the image contains a human face, which would cause incorrect emotion evaluation for the given image. In addition, previous emotional color transfer methods may easily result in ambiguity between the emotion of resulting image and target image. For example, if the background of the target image is dark while the facial expression is happiness, then previous methods would directly transfer dark color to the source image, neglecting the facial emotion in the image. To solve this problem, we propose a new facial-expression-aware emotional color transfer framework. Given a target image with facial expression features, we first predict the facial emotion label of the image through the emotion classification network. Then, facial emotion labels are matched with pre-trained emotional color transfer models. Finally, we use the matched emotion model to transfer the color of the target image to the source image. Considering none of the existing emotion image databases, which focus on images that contain face and background, we built an emotion database for our new emotional color transfer framework that is called “Face-Emotion database.” Experiments demonstrate that our method can successfully capture and transfer facial emotions, outperforming state-of-the-art methods.
Shiguang Liu, Huixin Wang, Min Pei
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Temporal-Consistency-Aware Video Color Transfer
Shiguang Liu
CGI1
2021 Multi-task Deep Learning for No-Reference Screen Content Image Quality Assessment
Ziqing Huang, Shiguang Liu
MMM (1)3
2021 Animating explosion with exploding sound and rigid-body sound
abstract
Abstract Visual simulation of explosion animation can be found many applications in virtual reality. In addition to visual simulation, sound synthesis of explosion scenes is also indispensable. However, little attention has been paid to this research area. To this end, this paper proposes an automatic sound synthesis method for explosion scenes. Explosion animation consists of two major events, namely explosive event and interactive rigid‐body fracture event. Correspondingly, the sound for an explosion scene can be viewed as the combination of two components, that is, the exploding sound and the rigid‐body sound. Among them, the exploding sound often shows similarity in terms of characteristics in frequency domain, while the rigid‐body sound is rich in variations due to the unpredictable motion state and different materials of debris. As a result, for the synthesis of the rigid‐body sound, firstly, we extract the rigid‐body sounds from recording examples by the empirical mode decomposition method. On this basis, a new spectral flux based approach is developed to segment the extracted sounds into a grain dictionary. Then, a cascade sound synthesis method based on impulse response is designed to synchronize the rigid‐body sound for fracture events. For the synthesis of exploding sound, due to its similar characteristics in frequency domain, an example‐based automatic sound synthesis method is adapted in this paper. Finally, we propose a priority‐aware sound blending method to merge these two types of completely different sounds reasonably. Various experimental results demonstrated the effectiveness of our new method.
Shiguang Liu, Siqi Xu
Comput. Animat. Virtual Worlds1
2021 Improved divergence-free smoothed particle hydrodynamics via priority of divergence-free solver and SOR
abstract
Abstract Fluid simulation plays an important role in movie special effects, computer games, etc. In recent years, the Smoothed Particle Hydrodynamics (SPH) has become a popular fluid simulation method due to its simpler implementation and better processing capabilities for various complex scenes. Although the computational efficiency of the current SPH method has been significantly improved, there is still much room for improvement in pressure solving. This paper proposes an improved pressure solution method based on DFSPH. Firstly, we eliminate the velocity divergence of the fluid by changing the solution order of the nonpressure force and divergence solvers, thereby reducing the number of iterations required by the density solver. Secondly, the SOR method is applied to the pressure solver, furthering reducing the number of iterations of the density solver. Various experiments demonstrated that our method improves the efficiency of the solver with little decrease in stability compared to state‐of‐the‐art methods. The speedup factor reaches between 1.35 and 1.4 at a time step of 5ms, and the more the number of particles, the faster the performance increases.
Maolin Wu, Shiguang Liu, Qing Xu 0002
Comput. Animat. Virtual Worlds2
2021 Perceptual Hashing With Visual Content Understanding for Reduced-Reference Screen Content Image Quality Assessment
abstract
Numerous screen content images (SCIs) have been produced to meet the needs of virtual desktop and remote display, which put forward a very urgent requirement for security and management of SCIs. Perceptual hashing is an effective way to deal with this issue. However, since SCIs are generally composed of pictures, graphics and texts, their intrinsic characteristics are different from those of natural images. Thus the previous hashing methods for natural images are not suitable for SCIs. In this article, we propose a perceptual hashing method for SCIs from the perspective of visual content understanding. Specifically, considering that the visual content understanding of SCIs mainly comes from textual regions, while the contours of text always have thinner width and higher contrast, it is decided to generate hash in the gradient field. An input screen image is first performed by some joint preprocessing operations. Then the maximum gradient magnitude and corresponding orientation information are extracted from three color channels R, G and B. Normalized histogram and local frequency coefficient features are further obtained from the maximum gradient magnitude. Finally, a hash sequence is constructed by statistics that are derived from extracted features. Experiments validated on three SCIs databases were conducted to evaluate classification between robustness and discrimination. Receiver operating characteristics (ROC) results demonstrate that the proposed method is superior to the state-of-the-art algorithms. Besides, SIQAD and SCID databases were leveraged to present the application in reduced-reference screen content image quality assessment, and comparisons show that our hashing could provide accurate predictions than other metrics.
Ziqing Huang, Shiguang Liu
IEEE Trans. Circuits Syst. Video Technol.2
2021 Perceptual Image Hashing With Texture and Invariant Vector Distance for Copy Detection
abstract
Content-based image copy detection has become one of the important technologies in copyright protection, where two major processes, content-based feature extraction and matching are included. However, it is certainly true that enough storage space is required to establish feature database for matching, which greatly increases time and storage consumption, as well as lacks flexibility. Fortunately, perceptual image hashing is a good strategy to address these problems, in which content-based features are extracted and further encoded to hash codes. On the one hand, content-based features provide and ensure higher copy detection accuracy, while on the other hand, hash codes instead of feature database reduce storage space and improve time efficiency. Meanwhile, a better balance between robustness and discrimination is one of the most objectives of image hashing, which is conducive to its application in multimedia management and security. Consequently, we present an effective image hashing method for copy detection. Specifically, to obtain perceptual robustness against to copy attacks, we extract the global statistical characteristics in gray-level co-occurrence matrix (GLCM) to reveal texture changes. Then, to make up the discrimination limitation, we leverage the local dominant DCT coefficients from the first row/column in each sub-image to calculate vector distance. Finally, two kinds of complementary information (global feature via texture and local feature via vector distance) are simultaneously preserved to generate hash codes. Various experiments performed on benchmark database indicate that our proposed perceptual image hashing provides higher detection accuracy and better balance between robustness and discrimination than the state-of-the-art algorithms.
Ziqing Huang, Shiguang Liu
IEEE Trans. Multim.2
2021 Binaural audio generation via multi-task learning
abstract
We present a learning-based approach for generating binaural audio from mono audio using multi-task learning. Our formulation leverages additional information from two related tasks: the binaural audio generation task and the flipped audio classification task. Our learning model extracts spatialization features from the visual and audio input, predicts the left and right audio channels, and judges whether the left and right channels are flipped. First, we extract visual features using ResNet from the video frames. Next, we perform binaural audio generation and flipped audio classification using separate subnetworks based on visual features. Our learning method optimizes the overall loss based on the weighted sum of the losses of the two tasks. We train and evaluate our model on the FAIR-Play dataset and the YouTube-ASMR dataset. We perform quantitative and qualitative evaluations to demonstrate the benefits of our approach over prior techniques.
Shiguang Liu, Dinesh Manocha
ACM Trans. Graph.2
2021 Video Decolorization Based on the CNN and LSTM Neural Network
abstract
Video decolorization is the process of transferring three-channel color videos into single-channel grayscale videos, which is essentially the decolorization operation of video frames. Most existing video decolorization algorithms directly apply image decolorization methods to decolorize video frames. However, if we only take the single-frame decolorization result into account, it will inevitably cause temporal inconsistency and flicker phenomenon meaning that the same local content between continuous video frames may display different gray values. In addition, there are often similar local content features between video frames, which indicates redundant information. To solve the preceding problems, this article proposes a novel video decolorization algorithm based on the convolutional neural network and the long short-term memory neural network. First, we design a local semantic content encoder to learn and extract the same local content of continuous video frames, which can better preserve the contrast of video frames. Second, a temporal feature controller based on the bi-directional recurrent neural networks with Long short-term memory units is employed to refine the local semantic features, which can greatly maintain temporal consistency of the video sequence to eliminate the flicker phenomenon. Finally, we take advantages of deconvolution to decode the features to produce the grayscale video sequence. Experiments have indicated that our method can better preserve the local contrast of video frames and the temporal consistency over the state of the-art.
Shiguang Liu, Huixin Wang
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Underwater sound propagation for virtual environments
Shiguang Liu
Vis. Comput.2
2021 Helmholtz decomposition-based SPH
abstract
SPH method has been widely used in the simulation of water scenes. As a numerical method of partial differential equations, SPH can easily deal with the distorted and complex boundary. In addition, the implementation of SPH is relatively simple, and the results are stable and not easy to diverge. However, SPH method also has its own limitations. In order to further improve the performance of SPH method and expand its application scope, a series of key and difficult problems restricting the development of SPH need to be improved. In this paper, we introduce the idea of Helmholtz decomposition into the framework of smoothed particle hydrodynamics (SPH) and propose a novel velocity projection scheme for three-dimensional water simulation. First, we apply Helmholtz decomposition to a three-dimensional velocity field and decompose it into three orthogonal subspaces. Then, our method combines the idea of spatial derivatives in SPH to obtain a discrete Poisson velocity equation. Finally, the conjugate gradient (CG) is utilized to efficiently solve the Poisson equation. The experimental results show that the proposed scheme is suitable for various situations and has higher efficiency than the current SPH projection scheme. Compared with the previous projection scheme, our solution does not need to modify the particle velocity indirectly by pressure projection, but directly by velocity field projection. The new scheme can be well integrated into the existing SPH framework, and can be applied to the interaction of water with static and dynamic obstacles, even for viscous fluid.
Zhongyao Yang, Maolin Wu, Shiguang Liu
Virtual Real. Intell. Hardw.3
2020 Flow Visualization with Density Control
Shiguang Liu, Hange Song
CGI1
2020 Physics-Guided Sound Synthesis for Rotating Blades
Siqi Xu, Shiguang Liu
CGI2
2020 Detail-Preserving Arbitrary Style Transfer
abstract
Style transfer has recently received considerable attention in the academic and industrial field. Recent work has achieved high quality stylized results by using the hidden activation. However, the hierarchy of content image has been mostly ignored. Artistic style factors including texture, stroke sizes or color, also distort, blur, add or remove content details. In this paper, we present a detail-preserving arbitrary style transfer network that can suppress structural distortion and preserve semantic content. We propose a refine network, which can not only spatially match the hierarchical structure but also preserve style patterns. Meanwhile, we introduce the global and local content losses to keep global structure and local detail information, respectively. In addition, our approach is not only effective against detail-preserving arbitrary style transfer, but also friendly with photographic and exposed style transfer. We utilize both quantitative and qualitative evaluations to demonstrate the effectiveness of our method.
Shiguang Liu
ICME2
2020 Outdoor Sound Propagation Based on Adaptive FDTD-PE
abstract
In outdoor scenes, the inhomogeneity of the atmosphere and the ground effect have a great impact on sound propagation, but these two effects are usually ignored in previous methods. We propose an adaptive FDTD-PE method to simulate sound propagation in 3D scenes taking into account atmospheric inhomogeneity and the ground effect to produce more realistic sound propagation results. In the simulation, the ground is considered as a porous medium with a certain thickness. The scene is categorized into a number of two-dimensional vertical ground planes in the three-dimensional cylindrical coordinate system. These planes are decomposed into the near-source complex regions and the far-source regions, which are solved by the FDTD solver and the parabolic equation (PE) solver, respectively. Furthermore, a novel encoding method was designed to process sound pressure data. In the far-source regions, the one-way sound propagation is only affected by the ground and atmosphere inhomogeneity, so we encode sound pressure data through function fitting. Finally, an efficient sound rendering method with this encoding representation is developed to perform auralization in the frequency-domain. We validated our method in various outdoor scenes, and the results indicate that our method can realistically simulate outdoor sound propagation, with quite higher speed and lower storage.
Shiguang Liu
VR1
2020 Multipath affinage stacked - hourglass networks for human pose estimation
Guoguang Hua, Shiguang Liu
Frontiers Comput. Sci.3
2020 Efficient Image Hashing with Geometric Invariant Vector Distance for Copy Detection
abstract
Hashing method is an efficient technique of multimedia security for content protection. It maps an image into a content-based compact code for denoting the image itself. While most existing algorithms focus on improving the classification between robustness and discrimination, little attention has been paid to geometric invariance under normal digital operations, and therefore results in quite fragile to geometric distortion when applied in image copy detection. In this article, a novel effective image hashing method is proposed based on geometric invariant vector distance in both spatial domain and frequency domain. First, the image is preprocessed by some joint operations to extract robust features. Then, the preprocessed image is randomly divided into several overlapping blocks under a secret key, and two different feature matrices are separately obtained in the spatial domain and frequency domain through invariant moment and low frequency discrete cosine transform coefficients. Furthermore, the invariant distances between vectors in feature matrices are calculated and quantified to form a compact hash code. We conduct various experiments to demonstrate that the proposed hashing not only reaches good classification between robustness and discrimination, but also resists most geometric distortion in image copy detection. In addition, both receiver operating characteristics curve comparisons and mean average precision in copy detection clearly illustrate that the proposed hashing method outperforms state-of-the-art algorithms.
Shiguang Liu, Ziqing Huang
ACM Trans. Multim. Comput. Commun. Appl.1
2019 Haptic Force Guided Sound Synthesis in Multisensory Virtual Reality (VR) Simulation for Rigid-Fluid Interaction
abstract
This paper tackles a challenging problem for interactive rigid-fluid interaction sound synthesis. One core issue of the rigid-fluid interaction in multisensory VR system is how to balance the algorithm efficiency, result authenticity and result synchronization. Since the sampling rate of audio is far greater than visual and haptic modalities, sound synthesis for a multisensory VR system is more difficult than visual simulation and haptic rendering, which still remains an open challenge until now. Therefore, this paper focuses on developing an efficient sound synthesis method tailored for a multisensory system. To improve the result authenticity while ensuring real time performance and result synchronization, we propose a novel haptic force guided granular sound synthesis method tailored for sounding in multisensory VR systems. To the best of our knowledge, this is the first step that exploits haptic force feedback from the tactile channel for guiding sound synthesis in a multisensory VR system. Specifically, we propose a modified spectral granular sound synthesis method, which can ensure real time simulation and improve the result authenticity as well. Then, to balance the algorithm efficiency and result synchronization, we design a multi-force (MF) granulation algorithm which avoids repeated analysis of fluid particle motion and thereby improves the synchronization performance. Various results show that the proposed sound synthesis method effectively overcomes the limitations of existing methods in terms of audio modality, which has great potential to provide powerful technological support for building a more immersive multisensory VR system.
Haonan Cheng, Shiguang Liu
VR2
2019 Liquid-solid interaction sound synthesis
Haonan Cheng, Shiguang Liu
Graph. Model.2
2019 Inside Cover Image
abstract
The cover image is based on the Inside Cover Image Robust Simultaneous Localization and Mapping in Low-light Environment, by Shiguang Liu et al., https://doi.org/10.1002/cav.1895.
Jiawei Huang 0002, Shiguang Liu
Comput. Animat. Virtual Worlds2
2019 Robust simultaneous localization and mapping in low-light environment
abstract
Abstract Complex and varied illumination makes computer vision research studies difficult. This research field pays much attention to scenes with weak illumination, especially in visual simultaneous localization and mapping (SLAM). Although the current feature‐based algorithm is mature, the existing SLAM method often fails because it cannot extract enough feature information in the low‐light environment. In this paper, we propose a new solution to this problem, which allows our system to work in environments with the majority of lighting. We propose a multifeature extraction algorithm to extract two kinds of image features simultaneously. With such a solution, our system can work when the single‐feature algorithm fails to extract enough feature points. We also add an image preprocessing step before tracking thread to cope with extremely dark conditions. Finally, we fully evaluate our approach on existing public data sets. Experiments show that the method combining multiple features can improve the robustness of the state‐of‐the‐art algorithm under weak illumination without affecting the real‐time performance.
Jiawei Huang 0002, Shiguang Liu
Comput. Animat. Virtual Worlds2
2019 An optimization on water wave diffraction approximation based on wave packets
abstract
Abstract This paper proposes a novel approach for diffraction approximation in two‐dimensional water wave simulation. Our method is based on wave packets, which ensures unconditionally stability. We introduce the Green function in order to generate diffraction patterns that are used to judge whether our method can produce realistic diffraction effects. This method also implements propagation direction optimization for diffractive wave packets to handle the efficiency problem existed in the state‐of‐the‐art approach. Various results indicate that our new method can produce more appealing diffractive effects than current particle‐based approaches. Besides, from the efficiency perspective, compared to current approaches, our implementation greatly enhances the computational efficiency for scenes with high curvature obstacle boundaries.
Zhongyao Yang, Shiguang Liu
Comput. Animat. Virtual Worlds2
2019 Progressive complex illumination image appearance transfer based on CNN
Shiguang Liu, Zhichao Song
J. Vis. Commun. Image Represent.1
2019 Human Pose Estimation in Video via Structured Space Learning and Halfway Temporal Evaluation
abstract
Human pose estimation from image or video is a basic issue in computer graphics and computer vision. The challenge of human pose estimation in video lies in the temporal coherency issue. The temporal consistency in video is the contents' similarity shown in the video frames. In video, temporal consistency maintenance of human pose estimation is to obtain better long-term consistency. Great major methods for the long-term consistency are using the whole video optimization method, which makes very large computation and the absence of consistency before and after the articulated limbs. In this paper, a novel method for the maintenance of temporal consistency is proposed. We maintain the temporal consistency of the video by the structured space learning and halfway temporal evaluation methods. We adopt a three-stage multi-feature deep convolution network framework to generate the initial posture joints position data, and a long-term temporal coherence is propagated to the overall video at each stage. The long-term consistency is more appealing since it produces stable results over larger periods of time. Our method can achieve good temporal consistency and get accurate and stable human pose estimation results. Various experimental results demonstrated the superiority of our method.
Shiguang Liu, Yang Li 0067, Guoguang Hua
IEEE Trans. Circuits Syst. Video Technol.1
2019 Shape-Optimizing and Illumination-Smoothing Image Stitching
abstract
Image stitching is usually subject to projective distortion and color difference problems. This paper presents an illumination-smoothing image stitching method based on the shape-optimizing hybrid transformation. An automatic mesh generation strategy is especially designed to reduce the calculation of the hybrid transformation within a reasonable range and guarantee the accuracy of image alignment. We consider stitching multiple images while constraining the distortion with the hybrid transformation and alleviate the color difference. In this paper, the method of color correction is adapted according to the characteristics of image stitching so as to achieve illumination-smoothing results. Moreover, the triangulation of the matching feature points is used to partition the color regions of the image and then more abundant color information is obtained for the calculation of the color transformation model. For general horizontal images, our method can achieve better results with less distortion in comparison with state-of-the-art methods. Various experimental results validate our new method.
Shiguang Liu, Qingpeng Chai
IEEE Trans. Multim.1
2019 Image Decolorization Combining Local Features and Exposure Features
abstract
Image decolorization is a task aiming to transform a color image to a grayscale one and is a dimension reduction process which inevitably suffers from information loss. The general goal of image decolorization is to preserve the color contrast of the color image. According to human visual study, exposure affects the human visual perception, and low-exposure areas or over-exposure areas will first attract the sense of sight. In addition, exposure also affects the contrast of the image, the contrast of low-exposure areas and over-exposure areas often cannot be well shown. Thus, the exposure should be taken into account in the process of image decolorization. Traditional local methods are not accurate enough to process local pixel blocks which may tend to cause local artifacts, while traditional global methods cannot greatly deal with local color blocks, which are usually time consuming too. Besides, the traditional image decolorization method usually uses the low-level features of an image. In this paper, the convolutional neural network is used to learn high-level abstract features of the image. We design a new convolutional neural-network framework with a local feature network and a rough classifier, which can learn the local semantic features and distinguish the different exposure conditions of color images. It is possible to learn the mapping model between input-output image pairs, which can generate better results in terms of color contrast preservation and exposure adjustment. Experiments indicate that our method does better in terms of color contrast preservation and exposure adjustment than the state of the art.
Shiguang Liu
IEEE Trans. Multim.1
2019 Physically-based statistical simulation of rain sound
abstract
A typical rainfall scenario contains tens of thousands of dynamic sound sources. A characteristic of the large-scale scene is the strong randomness in raindrop distribution, which makes it notoriously expensive to synthesize such sounds with purely physical methods. Moreover, the raindrops hitting different surfaces (liquid or various solids) can emit distinct sounds, for which prior methods with unified impact sound models are ill-suited. In this paper, we present a physically-based statistical simulation method to synthesize realistic rain sound, which respects surface materials. We first model the raindrop sound with two mechanisms, namely the initial impact and the subsequent pulsation of entrained bubbles. Then we generate material sound textures (MSTs) based on a specially designed signal decomposition and reconstruction model. This allows us to distinguish liquid surface with bubble sound and different solid surfaces with MSTs. Furthermore, we build a basic rain sound (BR-sound) bank with the proposed raindrop sound clustering method based on a statistical model, and design a sound source activator for simulating spatial propagation in an efficient manner. This novel method drastically decreases the computational cost while producing convincing sound results. Various experiments demonstrate the effectiveness of our sound simulation model.
Shiguang Liu, Haonan Cheng, Yiying Tong
ACM Trans. Graph.1
2018 Non-uniform Illumination Video Enhancement Based on Zone System and Fusion
abstract
In daily life, the acquisition of digital video will introduce non-uniform exposure due to the set of shooting equipment or scene illumination. Due to the low visibility, details are hidden and these videos usually fail to present visually pleasing browsing. Previous work typically relies on a single heuristic tone mapping curve to expand the dynamic range, which inevitably leads to uneven exposure. To solve this problem, we present a new video enhancement method based on zone system and image fusion. Given an input non-uniform illumination video, we first apply the zone system for exposure evaluation. We then remap each region using a series of tone mapping curves to generate multi-exposure regions which contain different exposed versions. Guided by some visual perception quality measures, we locate all the best exposed regions and then integrate them into a well-exposed video frame. Finally, in order to keep temporal consistency, we temporally propagate the zone regions from the former frame to the current frame. Experimental results have shown that the enhanced video exhibits uniform exposure, and preserves temporal consistency.
Shiguang Liu
ICPR2
2018 Robustness and Discrimination Oriented Hashing Combining Texture and Invariant Vector Distance
abstract
Image hashing is a novel technology of multimedia processing with wide applications. Robustness and discrimination are two of the most important objectives of image hashing. Different from existing hashing methods without a good balance with respect to robustness and discrimination, which largely restrict the application in image retrieval and copy detection, i.e., seriously reducing the retrieval accuracy of similar images, we propose a new hashing method which can preserve two kinds of complementary features (global feature via texture and local feature via DCT coefficients) to achieve a good balance between robustness and discrimination. Specifically, the statistical characteristics in gray-level co-occurrence matrix (GLCM) are extracted to well reveal the texture changes of an image, which is of great benefit to improve the perceptual robustness. Then, the normalized image is divided into image blocks, and the dominant DCT coefficients in the first row/column are selected to form a feature matrix. The Euclidean distance between vectors of the feature matrix is invariant to commonly-used digital operations, which helps make hash more compact. Various experiments show that our approach achieves a better balance between robustness and discrimination than the state-of-the-art algorithms.
Ziqing Huang, Shiguang Liu
ACM Multimedia2
2018 Detail-preserving SPH fluid control with deformation constraints
abstract
Abstract It is challenging to drive particle‐based smoothed‐particle hydrodynamics fluid to match the target shape and the deforming fluid shape between different models smoothly, especially when the natural fluid motion must be preserved. To achieve the desired behavior, we first generate control particles by sampling the target shapes and then apply a deformation constraint to each control particle, with its neighboring fluid particles keeping details within its influence region. For the generation of control particles, we classify input models into source object and target object, then separately sample them by voxelization method, and generate source control particles and target control particles, respectively. Our deformation constraint includes two parts. In the first part, we deform the source control model to the target control model according to specific space point correspondence between source control particles and target control particles; then, fluid particles are attracted by control particles and complete deformation between different shapes. In the second part, to reduce the lacking of fluid details when fluid deforms, we introduce a new control energy transfer mechanism for control particles. This deformation constraint is solved under smoothed‐particle hydrodynamics‐based fluid simulation framework, which makes our simulation fast, robust, and well suitable for interactive applications. Various experiments demonstrated the effectiveness of our method.
Shiguang Liu
Comput. Animat. Virtual Worlds2
2018 Example-based synthesis for sound of ocean waves caused by bubble dynamics
abstract
Abstract We present an automatic approach for the semantic modeling of indoor scenes based on a single photograph, instead of relying on depth sensors. Without using handcrafted features, we guide indoor scene modeling with feature maps extracted by fully convolutional networks. Three parallel fully convolutional networks are adopted to generate object instance masks, a depth map, and an edge map of the room layout. Based on these high‐level features, support relationships between indoor objects can be efficiently inferred in a data‐driven manner. Constrained by the support context, a global‐to‐local model matching strategy is followed to retrieve the whole indoor scene. We demonstrate that the proposed method can efficiently retrieve indoor objects including situations where the objects are badly occluded. This approach enables efficient semantic‐based scene editing.
Kai Wang 0015, Shiguang Liu
Comput. Animat. Virtual Worlds2
2018 Example-based synthesis for sound of ocean waves caused by bubble dynamics
abstract
Subsequent to publication, the abstract of the article by Wang et al.1 was found to be incorrect. The correct abstract is given below. This has also been corrected in the online version of the article. Abstract Water sound synthesis has attracted more and more attention in the recent years. Most current methods generate realistic water sounds based on simulation of bubble dynamics. However, for a large-scale dynamic water scene, this type of simulation is prohibitively expensive. This paper proposes a new approach to generate sound caused by bubble dynamics for ocean waves. In addition, the synthesis method is based on sound examples. The sound synthesis process includes the following steps: bubble particle extraction, wave attribute generation, and an example-based sound synthesis process. We propose a grid-based clustering method to improve the efficiency of our algorithm so that we can synthesize sound in real time. We also exploit a mapping method to properly generate the ocean sound according to the wave attributes. Various experiments demonstrated that our method can achieve competitive sounding results for ocean waves in terms of sound quality.
Kai Wang 0015, Shiguang Liu
Comput. Animat. Virtual Worlds2
2018 Parallel SPH fluid control with dynamic details
abstract
Abstract Real‐time fluid control is indispensable in computer animation, games, virtual reality, etc. In the field of Smoothed Particle Hydrodynamics (SPH) fluid control, the strategy of control force is often employed to control fluid particles; however, the artificial viscosity introduced by the control force would frequently lead to the loss of fine‐scale details. Although the introduction of the low‐pass filter can add back details, it may easily destroy the control target, and the control force method itself cannot make SPH fluid follow the fast‐moving control target. Meanwhile, this type of method is computation intensive and time consuming. To remedy the above problems, this paper proposed a novel, interactive SPH fluid control framework with turbulent details. We run SPH fluid simulation on Compute Unified Device Architecture (CUDA) and greatly improve the efficiency of fluid control. The control particle with curvature framework was adapted in this paper. We specially designed spring forces to make the fluid match a fast‐moving control target. Moreover, fine fluid details were preserved separately by calculating the fluid turbulence under control and the free fluid turbulence. This improved SPH fluid control can run in real time, which can enhance the visual quality of fluid animation as well. Our novel method can be applied in fluid animation with special control effects to guide fluid to form a target shape while greatly preserving the dynamic details of fluid. Various experiment results demonstrated the ability of our novel method.
Xiaoyong Zhang 0004, Shiguang Liu
Comput. Animat. Virtual Worlds2
2018 Emotional image color transfer via deep learning
Yaxi Jiang, Min Pei, Shiguang Liu
Pattern Recognit. Lett.4
2018 Stain Formation on Deforming Inelastic Cloth
abstract
We propose a novel approach to simulating the formation and evolution of stains on cloths in motion. We accurately capture the diffusion of a pigmented solution over a complex knitted or woven fabric through homogenization of its inhomogeneous and/or anisotropic properties into bulk anisotropic diffusion tensors. Secondary effects such as absorption, adsorption and evaporation are also accounted for through physically-based modeling. Finally, the influence of the cloth motion on the shape and evolution of the stain is captured by evaluating the inertial (e.g., centrifugal and Coriolis) forces experienced by the solution. The governing equations of motion are integrated in time directly on a deforming triangle mesh discretizing the inelastic cloth for efficiency and robustness. The deformation of the cloth can be precomputed or integrated through simplified two-way coupling, by using off-the-shell cloth simulations. Finally, numerical experiments demonstrate the plausibility of our results in practical applications by reproducing the usual shape and behavior of stains on various fabrics.
Shiguang Liu, Yiying Tong
IEEE Trans. Vis. Comput. Graph.2
2018 Sounding Solid Combustibles: Non-Premixed Flame Sound Synthesis for Different Solid Combustibles
abstract
With the rapidly growing VR industry, in recent years, more and more attention has been paid for fire sound synthesis. However, previous methods usually ignore the influences of the different solid combustibles, leading to unrealistic sounding results. This paper proposes SSC (sounding solid combustibles), which is a new recording-driven non-premixed flame sound synthesis framework accounting for different solid combustibles. SSC consists of three components: combustion noise, vortex noise and popping sounds. The popping sounds are the keys to distinguish the differences of solid combustibles. To improve the quality of fire sound, we extract the features of popping sounds from the real fire sound examples based on modified Empirical Mode Decomposition (EMD) method. Unlike previous methods, we take both direct combustion noise and vortex noise into account because the fire model is non-premixed flame. In our method, we also greatly resolve the synchronization problem during blending the three components of SSC. Due to the introduction of the popping sounds, it is easy to distinguish the fire sounds of different solid combustibles by our method, with great potential in practical applications such as games, VR system, etc. Various experiments and comparisons are presented to validate our method.
Qiang Yin 0003, Shiguang Liu
IEEE Trans. Vis. Comput. Graph.2
2018 Contrast preserving image decolorization combining global features and local semantic features
Shiguang Liu
Vis. Comput.2
2017 Efficient Decolorization via Perceptual Group Difference Enhancement
Shiguang Liu
ICIG (2)2
2017 Efficient sound synthesis for natural scenes
abstract
This paper presents a novel framework to generate the sound of outdoor natural scenes, such as waterfall, ocean, etc. Our method firstly simulates liquid with a grid-based method. Then combined with the movement of liquid, we generate seed-particles which represent bubbles, foams or splashes. Next, we assign each seed-particles a radius with a new radius distribution model. By calculating the bubbles' pressure wave we generate the sound. Experiments demonstrated that our novel framework can efficiently synthesize the sounds for natural scenes.
Kai Wang 0015, Haonan Cheng, Shiguang Liu
VR3
2016 Dynamic Fluid Visualization Based on Multi-level Density
abstract
This paper presents a novel method for streamline generation and selection for 2D and 3D fluid fields based on multi-level density. In the process, the outcome of the lowest-level density that includes the most streamlines is provided by a hybrid algorithm, which includes the entropy-based seeding strategy and grid-based filling method. Then the other levels are generated by using the selection algorithm based on the distance of the streamline and the fluid field, which is convenient for acquiring the results with different density. In 3D fields, the users can move viewpoint freely at the higher level so to obtain the general structure of fluid field and observe more details in the lower level.
Hange Song, Shiguang Liu
CASA2
2016 Shape-optimizing hybrid warping for image stitching
abstract
Projective distortion remains an open problem for image stitching. This paper proposed a novel weight-based shape-optimizing warping framework, which combines a projective transformation and a similarity transformation so as to reduce the projective distortion. This method aligned images with the projective transformation, and optimized images' shape under the constraint condition of similarity transformation. By automatically locating the overlapping and non-overlapping regions of input images, the weight of the constraint condition was determined. In addition, to ensure the smoothness of the final stitching result, we specially designed a continuous variation of the weight according to the location information. Since the proposed warping method joints the advantages of both projective transformation and similarity transformation, it can preserve the accuracy of image alignment and alleviate the projective distortion as small as possible. Various experiments demonstrated the effectiveness of our method.
Qingpeng Chai, Shiguang Liu
ICME2
2016 A computational approach to digital hand-painted printing patterns on cloth
Shiguang Liu
Multim. Tools Appl.1
2015 SPH fluid control with self-adaptive turbulent details
abstract
Abstract Smoothed particle hydrodynamics (SPH)‐based fluid control is often involved in fluid animation. Because most of the existing SPH fluid control methods employ the strategy of control force to control fluid particles, the artificial viscosity introduced by control force would lead to the loss of fine‐scale details. Although the introduction of the low‐pass filter can add details, it may easily destroy the target shape. To remedy the previous problems, we sample the control particles with curvature information to represent the shape complexity. Because of the shape's complexity, we suppress the generation of turbulence in the high‐curvature areas and promote turbulence in the low‐curvature regions. Our self‐adaptive way to randomly generate turbulence can effectively prevent the lack of fluid dynamics caused by the artificial viscosity. Our new method can improve the visual quality of the fluid animation, and the shape control result is consistent to the target shape. Copyright © 2015 John Wiley & Sons, Ltd.
Xiaoyong Zhang 0004, Shiguang Liu
Comput. Animat. Virtual Worlds2
2014 Visual fluid animation via lifting wavelet transform
abstract
ABSTRACT While small‐scale fluid details are crucial elements for the creation of visually pleasing fluid animations, their synthesis often requires heavy computation with traditional grid‐based fluid simulation methods. This paper proposes a novel method for enhancing the appearance of small‐scale details through frequency‐domain analysis. Different from previous work, our method detects and improves fluid details in the frequency‐domain via lifting wavelet decomposition. Based on a coarse‐to‐fine mechanism, the lifting wavelet composition first transforms the velocity in a fine grid into the frequency domain. Next, the velocity field is enhanced separately for different frequency bands. A novel velocity fusion method is developed for the enhancement of low‐frequency parts. On the other hand, high‐frequency parts are enhanced using a specially designed vorticity confinement method. Finally, the application of the inverse lifting wavelet transform determines the final velocity field with increased fine details. Our method can generate perceptually interesting fluid details, matching human visual perception theory. The results of various experiments validate the effectiveness and efficiency of our method. Copyright © 2014 John Wiley & Sons, Ltd.
Shiguang Liu, Jun-yong Noh, Yiying Tong
Comput. Animat. Virtual Worlds1
2013 Dynamic Fluids Mixed with Local-Control Effects
abstract
In dynamic fluid simulation, the user often need to control the local regions of the entire fluid scene for the special design effects, such as flowing marks and icons in dynamic fluids. Currently, the research aimed at such an issue is rare. It is rather difficult to realize the local feature effects using the global control methods. This paper proposes a novel local-control method for dynamic fluid simulation. Specifically, our method enables the generation vivid local features according to the pattern image specified by the user. We first search the whole fluid space to find the local areas best matched to the input pattern image. The special driving force is calculated to shape the local areas into the target pattern. In order to ensure the continuity of the whole velocity field, we mix the velocity fields of the local area and the one of its corresponding position in the global zone. Our method can provide the user or designers with the local-control tool for editing the dynamic fluid scenes intuitively. The dynamic fluid simulations with the local features have been experimented, which showed the validity and effectiveness of our method.
Liqun Wu, Shiguang Liu, Zhuojun Yu, Hanqiu Sun
ICIG2
2012 Selective color transferring via ellipsoid color mixture map
Shiguang Liu, Hanqiu Sun
J. Vis. Commun. Image Represent.1
2012 Automatic grayscale image colorization using histogram regression
Shiguang Liu
Pattern Recognit. Lett.1
2011 Fast Patch-Based Image Hybrids Synthesis
abstract
To synthesize a variety of texture images from examples in common PCs is an important topic in computer graphics. This paper proposed a novel patch-based image hybrids synthesis method, which can generate plausible results with visual diversity. First, image analysis was performed to build an effective data structure for synthesizing samples in a fast way which can maintain the texture structure. In the image analysis phase, the candidate patch set was built for each position. Then, various hybrid images were generated using patch-based texture synthesis method. Poisson editing algorithm and deblurring techniques were further adopted to optimize the results. Various experiments showed that this new method can hybrid sample images with large color difference, which can also overcome the less smooth boundary problem of the conventional synthesis methods.
Shiguang Liu, Jingting Wu
CAD/Graphics1
2011 Texture Transfer in Frequency Domain
abstract
Texture transfer has been extensively studied in literature recently. Most of the techniques are either structural or statistical, with an emphasis on utilizing texture synthesis algorithms. In this paper we address texture transfer in an alternative way, that is, to extract texture information in frequency domain, and then apply it to the target image. This algorithm is based on our intuitive discovery that texture transfer consists of color transfer and spatial-variation transfer. This algorithm possesses astonishing performance because of various FFT (fast Fourier transform) algorithms existing in literature for transforming images to frequency domain. The method can utilize any of the texture synthesis algorithms to generate an texture of the proper size, and then transfer this texture to the target image. We also propose a set of tools and parameters to control the process of this method.
Shiguang Liu
ICIG2
2011 Interactive soft-fabrics watering simulation on GPU
abstract
Abstract Physics‐based simulation is usually complex and time consuming, and consequently not suitable for real‐time applications. In this paper, we propose the efficient dynamics models for the real‐time simulation of soft fabrics interacting with water, including multi‐soaking effect and underwater dynamics. The multi‐soaking effect of soft fabrics is modeled based on the physics processes. Further, we develop the optimized mass–spring model that supports the large flow forces interacting with the fabrics underwater. The fabric spring forces are linearly derived and integrated with GPU–CUDA acceleration, feasible for real‐time VR applications with large set of fabric particles. Copyright © 2011 John Wiley & Sons, Ltd.
Hanqiu Sun, Shiguang Liu, Ping Li 0016
Comput. Animat. Virtual Worlds3
2011 Realistic simulation of mixing fluids
Shiguang Liu, Qiguang Liu, Qunsheng Peng 0001
Vis. Comput.1
2009 Physically based simulation of thin-shell objects' burning
Shiguang Liu, Qiguang Liu, Tai An, Qunsheng Peng 0001
Vis. Comput.1
2008 Simulation of atmospheric binary mixtures based on two-fluid model
Shiguang Liu, Zhangye Wang, Qunsheng Peng 0001
Graph. Model.1
2007 Texture Advection Based Simulation of Dynamic Cloud Scene
abstract
Fast display of realistic dynamic cloud scene is a challenging task for researchers in computer graphics. This paper describes a novel method of simulating animations of cloud. We first model the motion of cloud based on the physical theory of cloud formation. Then, the physical model of cloud is solved in discrete sparse grids. Thus, we can get approximate motion of cloud. Next, we advect texture image using the calculated velocity field to add the local detail of dynamic cloud scene. Here the texture image we used is Perlin noise map. The color of arbitrary point in the simulation space is got by calculating the product of the density texture and the regenerated Perlin noise texture. Finally, by changing different noise maps, realistic dynamic cloud scenes in different light environment are generated at high rendering rates.
Shiguang Liu, Ruoguan Huang, Zhangye Wang, Qunsheng Peng 0001, Jiawan Zhang
CAD/Graphics1
2007 Physically based animation of sandstorm
abstract
Abstract This paper describes a physically based method for modeling and animating sandstorm, a type of disastrous natural phenomenon. The method adopts a relatively stable incompressible multiple fluid model to simulate the motion of air, sand, and dust particles. The wind field of sandstorm is established based on Reynold‐average Navier‐Stokes equations. The sand and dust particle flow is therefore computed taking interaction among the wind, sand, and dust particles into account. To accelerate the modeling process of a dynamic sandstorm scene, a special Multi‐Fluid Solver is designed and implemented on GPU. Various illumination effects of sandstorm scenes can be simulated by spectral sampling scattering calculation. Animations of realistic sandstorms occurring in desert and urban areas based on our model are demonstrated. Compared with the real sandstorm photos, our simulated results are satisfactory. Copyright © 2007 John Wiley & Sons, Ltd.
Shiguang Liu, Zhangye Wang, Qunsheng Peng 0001
Comput. Animat. Virtual Worlds1
2007 Real time simulation of a tornado
Shiguang Liu, Zhangye Wang, Qunsheng Peng 0001
Vis. Comput.1