EDBT 2026 Demo / reviewers in the wild / expert
Ronghua Liang
dblp:49/6414
· DBLP profile ↗
159ranked-venue papers
13as first author
116since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 4 first-author · 61 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 26 · 3 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 18 · 2 first-author · 11 since 2021Security and privacy · 11 · 11 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ClassAid: A Real-time Instructor-AI-Student Orchestration System for Classroom Programming ActivitiesabstractGenerative AI is reshaping education, but it also raises concerns about instability and overreliance. In programming classrooms, we aim to leverage its feedback capabilities while reinforcing the educator’s role in guiding student–AI interactions. We developed ClassAid, a real-time orchestration system that integrates TA Agents to provide personalized support and an AI-driven dashboard that visualizes student–AI interactions, enabling instructors to dynamically adjust TA Agent modes. Instructors can configure the Agent to provide technical feedback (direct coding solutions), heuristic feedback (hint-based guidance), automatic feedback (autonomously selecting technical or heuristic support), or silent operation (no AI support). We evaluated ClassAid through three aspects: (1) the TA Agents’ performance, (2) feedback from 54 students and one instructor during a classroom deployment, and (3) interviews with eight educators. Results demonstrate that dynamic instructor control over AI supports effective real-time personalized feedback and provides design implications for integrating AI into authentic educational settings. Gefei Zhang 0002, Guodao Sun, Meng Xia 0002, Ronghua Liang |
CHI | 4 |
| 2026 | A Scene-Aware Meta-learning Framework for Robust Photovoltaic Power Forecasting
Yihan Yu, Yuanjie Dang, Peng Chen 0008, Yilong Zhang 0001, Ronghua Liang |
ICIC (16) | 6 |
| 2026 | Human-computer interaction and visualization in natural language generation models: applications, challenges, and opportunities
Yunchao Wang, Guodao Sun, Zihang Fu, Ronghua Liang |
Frontiers Comput. Sci. | 4 |
| 2026 | GCFL: Gray-Modality Conversion and Feature Learning for Visible-Infrared Person Re-Identification
Ruohong Huan, Mingzhen Wu, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2026 | A Dual Asynchronous-Synchronous Relation Graph Method for Sensor-Based Group Activity Recognition
Ai Bo 0001, Ruohong Huan, Peng Chen 0008, Ronghua Liang |
IEEE Internet Things J. | 5 |
| 2026 | Cloud-aided practical multi-keyword medical data retrieval based on ABE
Baichengjia Sun, Leyou Zhang, Ronghua Liang |
J. Syst. Archit. | 4 |
| 2026 | FreTransLS: Frequency Transformer based large-scale group activity recognition model for sensor data
Ruohong Huan, Meijiao Cao, Yantong Zhou, Peng Chen 0008, Guodao Sun, Ronghua Liang |
Pervasive Mob. Comput. | 7 |
| 2026 | RoCA: Robust Contrastive Adaptation for unsupervised anomaly detection
Chaohui Chen, Haidong Gao, Ronghua Liang |
Pattern Recognit. | 5 |
| 2026 | Learning Decoupled Features With Perceptual Distillation for Blind Image Quality AssessmentabstractExisting Blind Image Quality Assessment (BIQA) approaches typically employ subjective scores as optimization targets to train the model, aiming for results consistent with human judgments. Such judgments are derived from a comprehensive analysis of complex distortions and diverse semantics from images, whereas subjective scores represent the overall quality. This poses a significant challenge for a single model to learn diverse perceptual cues under weak supervision. To address this, we propose a Decoupled Feature Learning (DFL) framework that learns compact global content-aware and local distortion-aware features in a disentangled modeling for BIQA. Our key insight is to leverage global-local input pairs to decompose content-aware and distortion-aware cues entangled in distorted images, and aggregate decoupled perceptual features into a single network. We design a perceptual knowledge distillation strategy that progressively guides the student from fragmented representations to build local-to-global correspondences by distilling self-supervised semantic knowledge, while incorporating the Just-Noticeable-Difference (JND) model to highlight the transfer of perceptually sensitive content features. Finally, we introduce a local distortion-guided attention module to model synergistic effects of different perceptual features from the student for quality evaluation. Extensive experiments on eight benchmark datasets demonstrate the superior performance of the proposed model over the state-of-the-arts. In addition, the DFL framework is flexibly used to improve the perception ability of other Transformer variants. The code is released at https://github.com/JianjunXiang/DFT. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2026 | Performance Optimization Strategies for Data Transmission From Edge to Cloud: A ReviewabstractWith the rapid proliferation of IoT devices, the volume of generated data is growing at an unprecedented pace. Due to the limited resources of edge devices, a significant portion of this data must be transmitted to the cloud for in-depth processing, large-scale analysis, long-term storage, and archival purposes. Consequently, the performance has become a critical concern. While identifying prevailing challenges and research gaps in this domain requires a systematic review, such efforts remain largely absent from existing survey literature. This article addresses this gap by offering a structured review of recent optimization approaches. It begins by categorizing the literature into three main strategies: lossless transmission, lossy transmission, and hybrid approaches. In the context of lossless transmission, we analyze techniques such as data compression algorithms and incremental versus full synchronization mechanisms. For lossy strategies, we analyze approaches including lossy compression and predictive methods. In addition, we investigate hybrid strategies that integrate both lossless and lossy techniques to leverage their complementary advantages. Finally, we discuss the limitations of existing studies and highlight promising directions for future research in optimizing edge-to-cloud data transmission. Jian Liu 0053, Yangyang Lin, Ziguang Fu, Gexi Lin, Guodao Sun, Zhu Xiao, Yilong Zhang 0001, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Knowl. Data Eng. | 10 |
| 2026 | CompoVis: Is Cross-Modal Semantic Alignment of CLIP Optimal? A Visual Analysis AttemptabstractVision-language pre-trained models (VLMs) have shown impressive cross-modal understanding, yet their “compositional understanding” ability remains under investigation. We introduce CompoVis, a framework for visually probing cross-modal gaps in VLMs. CompoVis optimizes the grid layout to highlight alignment clusters and boundaries, visually interprets multi-head attention and semantic drift, and enables interactive fine-tuning unconstrained by closed datasets or offline models. Quantitative experiments and case studies explore key insights: VLMs rely on entity shortcuts rather than comprehension-driven; stubborn global modality isolation and suboptimal fine-grained alignment remain; fine-tuning with negative samples does not fundamentally alleviate the gaps. Approximately 89% of participants ($n=27$) found that, compared to methods relying solely on data metrics, CompoVis offers a more innovative and effective approach for investigating modality gaps in VLMs. Guodao Sun, Xueqian Zheng, Haidong Gao, Haixia Wang 0002, Ronghua Liang |
IEEE Trans. Multim. | 9 |
| 2026 | PorceVis: An interactive visual analytics system for exploring the history and culture of ancient Chinese porcelainabstractPorcelain, as a significant component of traditional Chinese culture, carries a profound historical legacy and rich cultural connotations. Its study involves a complex knowledge system spanning multiple dynasties and regions. Traditional research methods often rely on documentary analysis and artifact examination, which may not fully reveal the artistic and cultural characteristics of porcelain. In recent years, the rapid development of digital technologies has presented new opportunities for the research, preservation, and dissemination of cultural heritage. Therefore, this paper leverages image processing techniques and large language models to conduct a multidimensional quantitative analysis of the artistic features of porcelain, employing scientific methods to investigate its artistic value. Additionally, we developed an interactive visualization system that enables users to comprehend the development of porcelain from a spatiotemporal perspective and engage in interactive exploration of its artistic features at both macro and micro levels. Case studies and user evaluations demonstrate the system’s high usability and efficiency, providing a novel academic perspective and tools for the in-depth research and digital dissemination of Chinese cultural heritage. Xiaojie Pan, Jinhui Chu, Jian Liu 0053, Guodao Sun, Ronghua Liang |
Vis. Informatics | 9 |
| 2025 | CPVis: Evidence-based Multimodal Learning Analytics for Evaluation in Collaborative Programming
Gefei Zhang 0002, Shenming Ji, Yicao Li, Jingwei Tang, Jihong Ding, Meng Xia 0002, Guodao Sun, Ronghua Liang |
CHI | 8 |
| 2025 | DiffGen: Optimizing I/O Trace Generation with Differentiated Modeling Techniques
Jian Liu 0053, Zhiyang Feng, Ziguang Fu, Guodao Sun, Yilong Zhang 0001, Nan Gao 0001, Ronghua Liang, Peng Chen 0008 |
ICA3PP (5) | 7 |
| 2025 | Multimodal Sentiment Analysis with Parallel Attention and Correlation Fusion
Jie Lei 0002, Zunlei Feng, Ronghua Liang |
ICANN (3) | 5 |
| 2025 | Improving Sidescan Sonar Performance Using Array Upsampling Beamforming Synthetic ApertureabstractSidescan sonar technology plays an important role in seabed topography detection and mapping, but the performance of mainstream matched filter technology in long-distance image and towing speed limits its application scenarios. In addition, due to the slow propagation speed of sound waves in water, the synthetic aperture method in radar is difficult to adapt directly to sonar. To address these problems, a synthetic aperture algorithm framework for sidescan sonar with array upsampling beamforming is proposed (AUB-SAS). Firstly, the array element spacing is reduced by upsampling the array element to facilitate beamforming, and the long-distance information acquisition capability is improved by beam focusing. Secondly, the calculated multi-beam is used for synthetic aperture processing to make up for the low pulse repetition rate caused by the slow sound speed, thereby increasing the towing speed. In addition, the short-time Fourier transform with high resolution in the lateral direction and the motion compensation in the azimuth direction are derived, which further improve the imaging quality. The imaging results of lake test data prove the effectiveness of the proposed algorithm, achieving a lateral detection distance of 250 meters while the towing speed reaches 9 knots, and the imaging quality is significantly better than the traditional method. Weibo Mao, Peng Chen 0008, Shihui Liang, Ronghua Liang |
ICASSP | 6 |
| 2025 | Dynamic Routing and Calibration for Few-Shot Object DetectionabstractFew-shot object detection (FSOD), aiming to enhance the performance of novel object detection with limited labeled samples, has recently gained significant attention. Recent researches primarily focus on improving the generalization of novel classes and enhancing detector performance. However, the diversity of samples is often overlooked, and object proposals with inaccurate classifications or locations remain uncorrected. In this paper, we propose Dynamic Routing and Calibration for Few-Shot Object Detection (DRC-FSOD). Our approach includes a dynamic backbone routing that adapts to various samples by selecting appropriate backbones dynamically. Meanwhile, we construct a dynamic calibration module, which dynamically perform individual calibration for proposals based on their scores. Experimental results on MS COCO and Pascal VOC datasets show superiority over state-of-the-art methods. Jie Lei 0002, Zunlei Feng, Ronghua Liang |
ICASSP | 6 |
| 2025 | Spatial Continuity-Aware OCT Fingerprint Reconstruction Using Iterative Feature EnhancementabstractOptical coherence tomography (OCT) is a non-invasive imaging technique capable of acquiring depth information up to 1-3mm beneath the skin surface, including the stratum corneum and viable epidermis regions. This technique allows for the reconstruction of internal and external fingerprint images from grayscale data. However, existing fingerprint extraction methods heavily rely on contour features and current 2D approaches overlook the spatial continuity of biometric features in OCT images. Therefore, this paper proposes a novel iterative algorithm for internal and external fingerprint extraction from OCT images. This algorithm incorporates the spatial continuity of OCT slice images and an iterative feature enhancement module during the prediction phase to improve segmentation continuity. Additionally, a soft label technique is employed to reduce contour dependence and mitigate interference from noise and anomaly interference. Qualitative and quantitative experiments demonstrate significant improvements in segmentation accuracy with higher fingerprint quality, validating the effectiveness of the proposed approach. Yilong Zhang 0001, Xuanbing Chen, Shengming Zhu, Haohao Sun, Haixia Wang 0002, Jian Liu 0053, Yuanjie Dang, Ronghua Liang, Peng Chen 0008 |
IJCB | 8 |
| 2025 | What is the Role of Dataset Size and Fine-Tuning Method in Optimizing Small Language Models for Story Generation?
Yunchao Wang, Guodao Sun, Zihang Fu, Ronghua Liang |
ICIC (9) | 4 |
| 2025 | Dual Teacher with Dempster-Shafer Guidance for Decision Making in Semi-Supervised Small Object DetectionabstractSmall-scale object detection remains a major challenge in semi-supervised object detection (SSOD), particularly in medical image analysis. Conventional teacher models often struggle to accurately capture the features of low-contrast small lesions, leading to noisy pseudo-labels in both localization and classification, which introduces severe uncertainty and degrades detection performance. To address this issue, we propose Dual Teacher, a novel multimodal semi-supervised detection framework designed to enhance pseudo-label reliability and improve small-scale lesion detection. Specifically, we introduce two complementary teacher models: Hybrid-Scale Teacher, which exploits downsampled views to strengthen multi-scale feature learning, and Entropy-Based Multi-Modal Teacher, which leverages entropy maps to refine the quality of small-scale pseudo-labels. To effectively fuse predictions from both teachers and resolve conflicts, we propose a Dempster-Shafer-based Dual-Teacher pseudo-label fusion strategy that explicitly models uncertainty and optimizes classification confidence. Additionally, we introduce a class-adaptive threshold mechanism that dynamically adjusts pseudo-label selection based on dual-teacher predictions, further boosting the recall of small-scale lesions. Extensive experiments on the Dental Disease Dataset, ChestX-Det and M3FD demonstrate that our method consistently surpasses state-of-the-art SSOD approaches. Code is available at: https://github.com/z316910/Dual-Teacher.git. Nan Gao 0001, Junchao Zhu, Yilong Zhang 0001, Ronghua Liang, Guodao Sun, Peng Chen 0008 |
ACM Multimedia | 4 |
| 2025 | DOMR: Establishing Cross-View Segmentation via Dense Object MatchingabstractCross-view object correspondence involves matching objects between egocentric (first-person) and exocentric (third-person) views. It is a critical yet challenging task for visual understanding. In this work, we propose the Dense Object Matching and Refinement (DOMR) framework to establish dense object correspondences across views. The framework centers around the Dense Object Matcher (DOM) module, which jointly models multiple objects. Unlike methods that directly match individual object masks to image features, DOM leverages both positional and semantic relationships among objects to find correspondences. DOM integrates a proposal generation module with a dense matching module that jointly encodes visual, spatial, and semantic cues, explicitly constructing inter-object relationships to achieve dense matching among objects. Furthermore, we combine DOM with a mask refinement head designed to improve the completeness and accuracy of the predicted masks, forming the complete DOMR framework. Extensive evaluations on the Ego-Exo4D benchmark demonstrate that our approach achieves state-of-the-art performance with a mean IoU of 49.7% on Ego→Exo and 55.2% on Exo→Ego. These results outperform those of previous methods by 5.8% and 4.3%, respectively, validating the effectiveness of our integrated approach for cross-view understanding. Jitong Liao, Yulu Gao, Shaofei Huang 0001, Jialin Gao, Jie Lei 0002, Ronghua Liang, Si Liu 0001 |
ACM Multimedia | 6 |
| 2025 | SCLSTE: Semi-supervised Contrastive Learning-Guided Scene Text EditingabstractThe objective of scene text editing is to substitute the existing text with desirable text, while preserving the background and text styles intact. However, existing methods struggle to effectively replicate textual styles due to the complexity of real scenarios, and most are limited to training exclusively on labeled synthetic datasets. One alternative semi-supervised method incorporates unlabeled real-scene text images for training by using the original image as supervision and editing it with the same text. However, this approach risks degrading the model into an identity mapping network. To address these problems, we introduce a novel semi-supervised training strategy incorporating contrastive learning. It allows for editing real-scene text images with any text, circumventing the identity mapping issue while ensuring the accuracy of both text content and style. Moreover, we propose a robust Style-Aware Text Editing Module to address complexity of real scenarios and enhance the imitation of text styles. To the best of our knowledge, our work is the first to apply contrastive learning to the scene text editing. Extensive experiments demonstrate that our method outperforms existing models in terms of quality and quantity. Especially for the HEval metrics on both real scene (Tamper-Scene) and synthetic scene (Temper-Syn2k), we get both 12% improvement compared to state-of-the-art method. Min Yin, Liang Xie 0003, Haoran Liang 0001, Xing Zhao 0001, Ben Chen 0004, Ronghua Liang |
MMM (3) | 6 |
| 2025 | AutoMA: Automated Generation of Multi-level Annotations for Time Series VisualizationabstractTime series data is ubiquitous in people’s daily production and life, and visualizations augmented with annotations can significantly facilitate the understanding of such data and promote downstream tasks. Consequently, numerous annotation tools have been developed to detect and narrate useful patterns within time series visualizations. However, most existing tools can only identify basic factual insights (e.g., increasing or decreasing trends) that are already present in the charts. When users need deeper insights (e.g., predicting future trends) and richer contextual information (e.g., associative patterns between dimensions within and beyond the chart), these tools often fall short. To address this challenge, we present AutoMA, a system that automatically generates multi-level annotations for time series visualizations. We introduce an LLM-based pattern extraction method that supports the identification of seven distinct temporal patterns. Furthermore, we present a multi-level annotation design space that encompasses six specific annotation tasks, aimed at delivering richer contextual information. The generated annotations span a spectrum of information, ranging from directly observable temporal patterns to deeper insights obtained through further computation, and ultimately to advanced inter-dimensional association patterns. Finally, we demonstrate the effectiveness of our approach through experiments and user evaluations. The results indicate that AutoMA significantly enhances users’ ability to comprehend and explore time series data. Guodao Sun, Jingwei Tang, Yunchao Wang, Ronghua Liang |
PacificVis | 8 |
| 2025 | A Reflection on Leveraging Vision Language Model for Visual Analysis in Image-Based Person Re-IdentificationabstractImage-based person re-identification (Re-ID) aims to identify and track individuals across multiple camera views using query images. While machine learning methods have made progress, their real-world performance remains limited. A key challenge lies in the nature of person retrieval, which requires users to perform fine-grained matching and filtering. This process often involves manually comparing and evaluating a large number of candidate images, resulting in low retrieval efficiency and being time-consuming. To address these issues, we introduce textual information to assist users in performing fine-grained retrieval tasks. Specifically, we utilize vision-language models fine-tuned with domain knowledge to generate hierarchical textual descriptions as retrieval cues. We also provide a visual analysis tool, which adopts multi-view and adjustable visual encodings to aggregate and present image data, supporting interactive browsing and retrieval. Finally, we conduct a case study and visual analysis experiments to evaluate the effectiveness of the textual retrieval cues. The evaluation results reveal the potential of textual information in optimizing person retrieval and offers insights for future work. Guodao Sun, Ronghua Liang |
PacificVis | 5 |
| 2025 | Multi-granularity semantic relational mapping for image caption
Nan Gao 0001, Renyuan Yao, Peng Chen 0008, Ronghua Liang, Guodao Sun, Jijun Tang |
Expert Syst. Appl. | 4 |
| 2025 | FactExplorer: Fact Embedding-Based Exploratory Data Analysis for Tabular DataabstractDespite exploratory data analysis (EDA) is a powerful approach for uncovering insights from unfamiliar datasets, existing EDA tools face challenges in assisting users to assess the progress of exploration and synthesize coherent insights from isolated findings. To address these challenges, we present FactExplorer, a novel fact-based EDA system that shifts the analysis focus from raw data to data facts. FactExplorer employs a hybrid logical-visual representation, providing users with a comprehensive overview of all potential facts at the outset of their exploration. Moreover, FactExplorer introduces fact-mining techniques, including topic-based drill-down and transition path search capabilities. These features facilitate in-depth analysis of facts and enhance the understanding of interconnections between specific facts. Finally, we present a usage scenario and conduct a user study to assess the effectiveness of FactExplorer. The results indicate that FactExplorer facilitates the understanding of isolated findings and enables users to steer a thorough and effective EDA. Guodao Sun, Lvhan Pan, Baofeng Chang, Haoran Liang 0001, Ronghua Liang |
Int. J. Hum. Comput. Interact. | 8 |
| 2025 | Towards Better Utilization of Haptic Interaction in Visualization: Design Space and Knob PrototypeabstractHumans encounter a vast array of sensory stimuli in their everyday lives. However, many visualization techniques primarily utilize visual feedback, which may disregard certain intricate details. Relying on a single visual channel may overlook complex layouts. However, how haptic force feedback can be used to assist visualization remained under-explored. In this work, we initially conducted a literature review to identify potential problems in the visualization of large datasets and engaged in discussions with domain experts to explore the potential of haptic force feedback and visual collision representation. Subsequently, we designed an innovative haptic force feedback knob, which included 3 primary modules and 29 elements. To evaluate the clarity and usefulness of this design space, we conducted a workshop and devised “recommended solutions” for the identified visualization problems. Finally, we implemented a prototype of the haptic force feedback knob and assessed its performance on scatterplot and parallel coordinate plot tasks using large datasets. The results indicated that the knob prototype could reduce visual strain and enhance the efficiency of visualization tasks. Gefei Zhang 0002, Guodao Sun, Zifeng Sun, Jingwei Tang, Ronghua Liang |
Int. J. Hum. Comput. Interact. | 6 |
| 2025 | A non-interactive Online Medical Pre-Diagnosis system on encrypted vertically partitioned data
Ronghua Liang, Guoqiang Deng |
J. Biomed. Informatics | 3 |
| 2025 | Semantic-aware representations for unsupervised Camouflaged Object Detection
Zelin Lu, Xing Zhao 0001, Liang Xie 0003, Haoran Liang 0001, Ronghua Liang |
J. Vis. Commun. Image Represent. | 5 |
| 2025 | TWDT: Training-free word-level controllable diffusion model for text generation
Nan Gao 0001, Yangjie Lu, Peng Chen 0008, Guodao Sun, Ronghua Liang, Yilong Zhang 0001 |
Knowl. Based Syst. | 5 |
| 2025 | DBNetVizor: Visual Analysis of Dynamic Basketball Player NetworksabstractVisual analysis has been increasingly integrated into the exploration of temporal networks, as visualization methods have the capability to present time-varying attributes and relationships of entities in an easy-to-read manner. Visualization techniques have been employed in a variety of dynamic network datasets, including social media networks, academic citation networks, and financial transaction networks. However, effectively visualizing dynamic basketball player network data, which consists of numerical networks, intensive timestamps, and subtle changes, remains a challenge for analysts. To address this issue, we propose a snapshot extraction algorithm that involves human-in-the-loop methodology to help users divide a series of networks into hierarchical snapshots for subsequent network analysis tasks, such as node exploration and network pattern analysis. Furthermore, we design and implement a prototype system, called DBNetVizor, for dynamic basketball player network data visualization. DBNetVizor integrates a graphical user interface to help users extract snapshots visually and interactively, as well as multiple linked visualization charts to display macro- and micro-level information of dynamic basketball player network data. To demonstrate the usability and efficiency of our proposed methods, we present two case studies based on dynamic basketball player network data in a competition. Additionally, we conduct an evaluation and receive positive feedback. Baofeng Chang, Guodao Sun, Sujia Zhu, Jingwei Tang, Ronghua Liang |
IEEE Trans. Big Data | 7 |
| 2025 | Towards Enhancing Inter-Domain Routing Security With Visualization and Visual AnalyticsabstractIn the complex landscape of the Internet, inter-domain routing systems are essential for ensuring seamless connectivity and reachability across autonomous systems. However, the lack of dependable security validation mechanisms in these systems poses persistent challenges. Vulnerabilities such as prefix hijacking, path forgery, and route leakage not only compromise network operators and users, but also threaten the stability and accessibility of the Internet’s core infrastructure. To address this, visualization and visual analytics techniques are adept at identifying and detecting security threats, offering network administrators effective methods to monitor and maintain network operations. This paper presents a comprehensive survey of the state-of-the-art research in visualization and visual analytics for inter-domain routing security. We delineate four scenarios for tasks analysis in network visualization: monitoring, detection, verification, and discovery. Each category is explored in detail, focusing on the employed data sources and visualization techniques. Several key findings are presented at the end of each category, aimed at providing researchers and practitioners with research inspiration. Furthermore, we examine the trends of academic interest observed in recent decades and propose potential directions for future research in visual analytics pertaining to Internet infrastructure security. Jingwei Tang, Guodao Sun, Gefei Zhang 0002, Yanbiao Li 0001, Guangxing Zhang, Jian Liu 0053, Haixia Wang 0002, Ronghua Liang |
IEEE Trans. Big Data | 10 |
| 2025 | A Weakly-Supervised Cross-Domain Query Framework for Video Camouflage Object DetectionabstractVCOD (Video Camouflage Object Detection) is a crucial security technology that identifies camouflaged objects in videos, bolstering security measures across diverse applications. On one hand, appearance-based VCOD methods face challenges because camouflaged appearances cause objects to blend into their surroundings, and current VCOD methods typically utilize optical flow to represent motion information. However, over-reliance on accurate estimation renders the model overly fragile. On the other hand, there is a shortage of effectively annotated camouflaged video datasets, coupled with the time-consuming and labor-intensive annotation process, severely constraining the development of this field. To address this, we propose a novel weakly-supervised framework for VCOD based on cross-domain querying of preceding and succeeding frames. Specifically, we propose a time-efficient and labor-saving manual annotation approach based on large visual models to rapidly generate pseudo-labels. Furthermore, we design a network based on Spatio-Temporal Memory (STM) that performs cross-modal feature querying with the current frame against preceding and succeeding frames to acquire useful information, thereby enhancing the focus on temporal information. Extensive experiments conducted on two common VCOD datasets have proven the effectiveness of our method, achieving state-of-the-art performance on the challenging camouflaged video data. Zelin Lu, Liang Xie 0003, Xing Zhao 0001, Binwei Xu, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Multidimensional Exploration of Segment Anything Model for Weakly Supervised Video Salient Object DetectionabstractFully supervised video salient object detection (VSOD) has made considerable breakthroughs using costly and time-consuming pixel-wise annotations. Recently, to achieve a trade-off between the annotation burden and the model performance, scribble-based VSOD tasks have attracted increasing attention. However, learning the complete object structure and precise boundary details from sparse scribble annotations remains challenging. In this paper, we propose a series of strategies to effectively explore valid information from the recently proposed segmentation foundation model “Segment Anything Model (SAM)” in various perspectives to address these challenges. Specifically, due to the limited performance of SAM on videos, we propose a SAM-guided label enhancement method instead of directly using the results of SAM, which can introduce edge information while reducing the interference of erroneous information. Moreover, we propose a SAM-driven spatiotemporal network guided by general semantic features from the SAM encoder to help the model be aware of global connections. Additionally, we propose a SAM-based global-aware loss, which further considers the affinity constraint between predicted results and foreground labels or background labels from a global perspective, guiding the model to perceive the complete salient objects. Experimental results demonstrate that our method outperforms state-of-the-art weakly supervised VSOD methods and is comparable to fully supervised VSOD methods. Binwei Xu, Qiuping Jiang, Xing Zhao 0001, Chenyang Lu 0002, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | SonarPoint: Weak-Heterogeneity Awareness Object Detection Network for 3D Sonar Point CloudabstractUnderwater target detection is primarily achieved through two methods: optical imaging and underwater sonar. 3D sonar, as the most advanced underwater detection technology, is characterized by strong penetration and long scanning distance, making it more suitable for tasks such as deep-sea exploration, murky water detection, and long-distance target identification. However, acquiring underwater sonar images is challenging, and there is no open-source 3D sonar dataset. Traditional three-dimensional target detection methods typically require highquality data and face significant challenges when dealing with weak heterogeneous sonar point clouds caused by high noise, low resolution, and occlusions. To address the aforementioned issues, we first propose a novel fuzzy decoupling module that differs from traditional foreground-background segmentation. This module simultaneously extracts valuable information about the target and its surrounding environment, mitigating the reduction in heterogeneity caused by noise and sonar side lobes. To achieve efficient fusion and capture global information after fuzzy decoupling, a multi-hop Mamba seamless adaptive decoupling point is introduced. It effectively enhances the connection between the two decoupled parts. To address missing and occlusion problems, a second-stage refinement based on Markov prediction is proposed. This low-cost approach, in contrast to using the original point cloud for contour and detail completion, enriches target boundary information. To validate our method, we have designed a practical 3D sonar imaging system and tested it through lake-based experiments. We have collected extensive raw data from Qiandao Lake and conducted annotation work. Through qualitative and quantitative experiments, our method outperforms the most advanced methods by 11.4%. Tiancheng Cai, Peng Chen 0008, Weibo Mao, Yingtian Hu, Yilong Zhang 0001, Yuanjie Dang, Ronghua Liang, Xiang Tian 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Orthogonal View-Based Attention Network for Layer Segmentation of 3D OCT FingerprintsabstractRecently, optical coherence tomography (OCT) has been used to noninvasively image the 3D structure of fingertip skin at high resolution. Unlike traditional 2D sensors (e.g., infrared light or capacitive technologies), the friction ridge information in 3D OCT fingerprint measurements requires reconstruction through layer segmentation. Accurate layer segmentation is helpful for fingerprint recognition and antispoofing applications. OCT volumes contain information corresponding to different directions that naturally provide complementary views. Inspired by this fact, we propose a novel orthogonal view-based attention network called OVA-Net, which exploits orthogonal views to learn the complementary information implied in the 3D fingerprint structure. Specifically, 3D convolutions and an A-line-based attention module are proposed in the B-scan view to model the long short-term intraslice correlations, whereas their counterparts in the C-scan view aim to model interslice correlations. An optical flow-based attention module is also proposed in the B-scan view to extract correlations between B-scans, which complements the interslice correlation learned in the C-scan view. Features from orthogonal views are progressively incorporated into a fusion pipeline for 3D layer segmentation. The effectiveness of OVA-Net is comprehensively evaluated in terms of layer segmentation accuracy, fingerprint reconstruction quality, and recognition performance. Yipeng Liu 0002, Zhanqing Li, Jiajin Qi, Hangtao Yu, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | Test-Time Image Reconstruction for Cross-Device OCT Fingerprint ExtractionabstractOptical coherence tomography (OCT) technology enables imaging of 3D fingerprint structures. Extracting surface and internal fingerprints for identity recognition is possible by processing OCT images with layer segmentation and contour extraction. However, due to domain shift effects, OCT fingerprint extraction models often struggle to perform well across different devices. In this paper, a cross-device OCT fingerprint extraction method based on test-time image reconstruction is proposed. This method simultaneously trains layer segmentation and image reconstruction tasks during training. Additionally, a contour classification task is integrated to ensure the continuity and robustness of the contour extraction results. During the testing phase, image reconstruction is performed on test images, and the shared modules are updated to adapt the layer segmentation and contour classification network to the test domain. The result with the minimum inconsistency during the testing phase is selected as the final prediction. Experiments and comparisons are performed in terms of the distance between the ground truth and the extracted contours. Yipeng Liu 0002, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Supervised Enhancement for Fingertip OCT Images Based on Paired Dataset Generation StrategyabstractOptical Coherence Tomography (OCT) is a high-resolution, non-invasive imaging technology increasingly used for biometric data collection from fingertips. OCT captures volume data up to 3mm below the skin surface in the form of a series of B-scan images, enabling the reconstruction of internal fingerprints (IF) and internal sweat pores (ISP), thereby enhancing the security of biometric recognition. Despite the advantages, OCT images suffer from speckle noise and tissue discontinuity, making the extraction of subcutaneous biometric features challenging. Traditional hardware and software-based enhancement methods often result in over-smoothing and structural loss. Recent advancements in deep learning (DL) offer promising alternatives, with supervised DL methods showing efficacy when trained with high-quality paired datasets. However, the absence of ground-truth (GT) data makes it impossible to apply these models. This study proposes a novel supervised enhancement method for fingertip OCT images, with a paired dataset generation strategy. An OCT few-shot GAN and a Quality Estimation Module are proposed and incorporated into the strategy to realize translation from minimal GT manual augmentation to high-quality paired dataset, effectively addressing the challenge of data scarcity. A Fast Supervised Enhancement GAN (FSE-GAN) is proposed thereafter to perform simultaneous speckle noise reduction and tissue structure restoration, facilitating accurate extraction of internal fingerprints and sweat pores. Experiments demonstrate that the enhanced images significantly simplify IF and ISP extraction while achieving outstanding result quality. Qingran Miao, Haixia Wang 0002, Jianru Zhou, Yilong Zhang 0001, Peng Chen 0008, Ronghua Liang, Yuanjing Feng |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | CRM-NAS: A Structure-Adaptive and Attention-Based Approach for Fingerprint Reconstruction From Noisy OCT DataabstractAs essential biometric features, fingerprints have been widely utilized in various security domains. However, the performance of conventional Automated Fingerprint Identification Systems (AFISs) is limited by the quality of the external fingerprint (EF), particularly in cases involving damaged or deformed prints. Using the internal fingerprint (IF) acquired by Optical Coherence Tomography (OCT) to address these limitations has emerged as a promising method. IFs can compensate for and restore missing ridge pattern features in degraded EFs, thereby improving the overall recognition accuracy of AFIS. However, the reconstruction of IF was significantly constrained by speckle noise in OCT images, making the accurate extraction of finger tissue contours complex and computationally intensive. To improve the applicability of OCT fingerprint, this paper proposes a Neural Architecture Search (NAS)-based OCT fingertip internal contour regression network, denoted as CRM-NAS. The CRM-NAS employs a NAS-based internal feature extraction module (NAS-IEM) to adaptively optimize the network architecture and complexity with noisy OCT fingertip data, facilitating the effective capture of global internal contour features. Furthermore, an attention-based contour regression module (Att-CRM) is introduced to refine local contour details by leveraging multi-scale intermediate features extracted from different network layers and to enable the generation of continuous and accurate internal contours. Experimental results demonstrate that CRM-NAS not only outperforms existing methods in terms of contour extraction accuracy, fingerprint reconstruction quality, and verification performance, but also maintains a relatively compact parameter size. Haohao Sun, Sihan Lan, Haixia Wang 0002, Yilong Zhang 0001, Yipeng Liu 0002, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | A Fingerprint Quality Driven Transformer-CNN Hybrid Model for External and Internal Fingerprint FusionabstractAdvancements in internal fingerprint extraction technology have made the fusion of external and internal fingerprints possible. It offers a viable solution to the problem of degraded performance in Automatic Fingerprint Identification System (AFIS) caused by epidermal abrasion and aging. Traditional fusion methods focus on information maximization. But for fingerprint, features like wrinkles and scars often yield high gradient variation information. It is detrimental to generating high-quality fingerprint. To address this, we propose a novel quality driven fusion method for external and internal fingerprints. It comprises several components. Firstly, there is a lightweight and efficient Transformer-CNN hybrid model. Secondly, it includes a closed-loop quality driven fusion mechanism. This mechanism is equipped with a quality prediction module, Weighted Complementary Fusion (WCF), and quality feedback. Thirdly, there is a jointly optimized combined loss function, which is accompanied by an asynchronous cross-training strategy. Unlike traditional paradigms, we change the optimization objective. It is shifted from information maximization to quality maximization, which is more appropriate for fingerprint. Experimental evaluations have been conducted, covering aspects such as fingerprint quality, matching performance, and network model ablation. The method we proposed demonstrates superiority in terms of quality score and matching performance. It outperforms both traditional and state-of-the-art approaches. It gives a new research path to boost fingerprint identification performance in identity security authentication. Haixia Wang 0002, Yilong Zhang 0001, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | A Spatial-Aware Temporal Modeling Network for Imitation Learning-Based Drone NavigationabstractImitation learning-based drone autonomous navigation has attracted significant attention due to the ability of leveraging deep neural networks to learn the control policy from human pilot demonstrations. However, most current studies generate the control command using only a single image, overlooking the semantic information embedded in the sequential input images. While some reinforcement learning-based methods have explored the temporal modeling of sequential input images, they often overlook the spatial relations between frames and vectorize 2D information of each image into a 1D feature. In this paper, we propose a novel imitation learning-based method, termed the spatial-aware temporal modeling network (SATMN), for autonomous drone navigation using sequential images as input. Specifically, we introduce a spatial-temporal-separated modeling mechanism to extract low-resolution spatial features from original images and then perceive spatial-temporal relations among these 2D features. SATMN preserves the spatial information of each 2D image feature during temporal modeling and enables real-time onboard computing on a drone. To validate the effectiveness of the proposed method, we design a compact quadrotor platform capable of autonomous navigation using SATMN, entirely powered by onboard computing devices. Comprehensive and reproducible experiments on public datasets demonstrate the superior performance of our method compared to existing approaches. Tianwei Yu, Yuanjie Dang, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | MulDeF: A Model-Agnostic Debiasing Framework for Robust Multimodal Sentiment AnalysisabstractIn recent years, multimodal sentiment analysis (MSA) has gained prominence with the proliferation of social media. However, prior studies have often disregarded the possibility of spurious correlations between multimodal data and sentiment labels. Neglecting these factors often results in significant performance degradation, hampering the model's ability to generalize in out-of-distribution (OOD) scenarios. To gain a comprehensive understanding of multimodal knowledge and enhance the model's generalization across diverse distribution scenarios, we present the Multimodal Debiasing Framework (MulDeF). This model-agnostic framework addresses label bias through causal intervention and tackles multimodal biases using counterfactual reasoning. During the training phase, MulDeF rectifies multimodal representations through frontdoor adjustment in causal intervention, effectively eliminating label bias. In order to model conditional expectation calculations within the context of frontdoor adjustment, we introduce multimodal causal attention (MCA). In the inference phase, it employs counterfactual reasoning to eliminate multimodal biases. To further refine our debiasing strategy, we categorize multimodal biases into two distinct types: nonverbal bias and verbal bias. Nonverbal bias is addressed at the utterance level, involving the establishment of unimodal models for audio and visual modalities to estimate their biases concerning sentiment labels. Conversely, verbal bias mitigation occurs at the word level. Here, we mask “harmless” words to generate corresponding counterfactual texts, which are then assessed by the text model to identify word-level bias. Experimental results validate the effectiveness of MulDeF, showcasing its superior performance in OOD settings compared to state-of-the-art methods, while also achieving competitive results in independent and identically distributed (IID) settings. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Multim. | 4 |
| 2025 | Learning Video Salient Object Detection Progressively From Unlabeled VideosabstractRecently, deep learning-based video salient object detection (VSOD) has achieved some breakthroughs, but these methods rely on expensive annotated videos with pixel-wise annotations or weak annotations. In this paper, based on the similarities and differences between VSOD and image salient object detection (SOD), we propose a novel VSOD method via a progressive framework that locates and segments salient objects in sequence without utilizing any video annotation. To efficiently use the knowledge learned in the SOD dataset for VSOD efficiently, we introduce dynamic saliency to compensate for the lack of motion information of SOD during the locating process while maintaining the same fine segmenting process. Specifically, we utilize the coarse locating model trained on the image dataset, to identify frames with both static and dynamic saliency. Locating results of these frames are selected as spatiotemporal location labels. Moreover, by tracking salient objects in adjacent frames, the number of spatiotemporal location labels is increased. On the basis of these location labels, a two-stream locating network with an optical flow branch is proposed to capture salient objects in videos. The results with respect to five public benchmarks demonstrate that our method outperforms the state-of-the-art weakly and unsupervised methods. Binwei Xu, Qiuping Jiang, Haoran Liang 0001, Dingwen Zhang, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 5 |
| 2025 | Position Fusing and Refining for Clear Salient Object DetectionabstractMultilevel feature fusion plays a pivotal role in salient object detection (SOD). High-level features present rich semantic information but lack object position information, whereas low-level features contain object position information but are mixed with noises such as backgrounds. Appropriately addressing the gap between low- and high-level features is important in SOD. We first propose a global position embedding attention (GPEA) module to minimize the discrepancy between multilevel features in this article. We extract the position information by utilizing the semantic information at high-level features to resist noises at low-level features. Object refine attention (ORA) module is introduced to refine features used to predict saliency maps further without any additional supervision and heighten discriminative regions near the salient object, such as boundaries. Moreover, we find that the saliency maps generated by the previous methods contain some blurry regions, and we design a pixel value (PV) loss to help the model generate saliency maps with improved clarity. Experimental results on five commonly used SOD datasets demonstrated that the proposed method is effective and outperforms the state-of-the-art approaches on multiple metrics. Xing Zhao 0001, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | SOEDiff: Efficient Distillation for Small Object EditingabstractIn this article, we delve into a new task known as Small Object Editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion . These failures stem from the limited use of small-sized objects in training datasets and the down-sampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA , which efficiently fine-tunes low-rank matrices, and Cross-scale score distillation , which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in SOE. In particular, when comparing SOEDiff with SD-I model on the OpenImage-small-val dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID. Yiming Wu 0005, Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Ronghua Liang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | VAC$^{2}$2: Visual Analysis of Combined Causality in Event SequencesabstractIdentifying causality behind complex systems plays a significant role in different domains, such as decision-making, policy implementations, and management recommendations. However, existing causality studies on temporal event sequence data mainly focus on individual causal discovery, which is incapable of capturing combined causality. To address the gap in combined causality discovery on temporal event sequence data, eliminating and recruiting principles are defined to balance the effectiveness and controllability of cause combinations. We also leverage the Granger causality algorithm based on the Reactive point processes to describe impelling or inhibiting behavior patterns among entities. In addition, we design an informative and aesthetic visual metaphor of "electrocircuit" to encode aggregated causality for ensuring that our causality visualization exhibits no node-overlap, no edge-intersection, and no link-ambiguity. Aggregation layout, diverse sorting strategies, and smooth interactions are also integrated into our directed, weighted, and parallel-based hypergraph for illustrating combined causality. Our developed combined causality visual analysis system, namely VAC$^{2}$2, can help users effectively explore combined causes as well as individual causes. This interactive system supports multi-level causality exploration with diverse ordering strategies and a focus+context technique to help users obtain different levels of information abstraction. The usefulness and effectiveness of our work are further evaluated by conducting two case studies and a controlled user study on event sequence data. Sujia Zhu, Guodao Sun, Baofeng Chang, Jingwei Tang, Ronghua Liang |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | ACPNet: Enhancing Small-Scale Dieases Detection in Panoramic X-raysabstractDeep learning-based disease detection can automatically identify dental diseases in panoramic X-rays and improve the accuracy and efficiency of doctors’ diagnoses. However, due to the complex data distribution of panoramic oral X-rays, significant scale differences among lesions, and the presence of many small-scale diseases, automated disease detection in panoramic oral X-rays faces considerable challenges. To alleviate the aforementioned issues, we propose ACPNet, which introduces a novel two-stage approach for detecting small-scale dental diseases in panoramic X-rays using the Contextual Attention Alignment Network (CAAN) and the Point-to-Patch Module (PTPM). To the best of our knowledge, we are the first to explore the detection of small-scale dental diseases in panoramic X-rays under limited sample conditions. Specifically, CAAN integrates deformable convolution with the global attention mechanism of transformer attention, enabling the model to more accurately extract small target foreground features in the complex background of panoramic X-rays. PTPM employs key point detection and cascade dynamic patches to adjust the bounding boxes of lesions, ensuring that small-scale diseases have sufficient high-quality proposals, thereby enhancing detector performance. Additionally, we collected a dataset containing 1157 instances of dental diseases to validate the effectiveness of our algorithm. Extensive experiments demonstrate that ACPNet achieves state-of-the-art performance, highlighting its superiority over baseline and other detection methods. Nan Gao 0001, Junchao Zhu, Peng Chen 0008, Jijun Tang, Ronghua Liang |
BIBM | 6 |
| 2024 | Self-Distilled Dynamic Fusion Network for Language-Based Fashion RetrievalabstractIn the domain of language-based fashion image retrieval, pinpointing the desired fashion item using both a reference image and its accompanying textual description is an intriguing challenge. Existing approaches lean heavily on static fusion techniques, intertwining image and text. Despite their commendable advancements, these approaches are still limited by a deficiency in flexibility. In response, we propose a Self-distilled Dynamic Fusion Network to compose the multi-granularity features dynamically by considering the consistency of routing path and modality-specific information simultaneously. Two new modules are included in our proposed method: (1) Dynamic Fusion Network with Modality Specific Routers. The dynamic network enables a flexible determination of the routing for each reference image and modification text, taking into account their distinct semantics and distributions. (2) Self Path Distillation Loss. A stable path decision for queries benefits the optimization of feature extraction as well as routing, and we approach this by progressively refine the path decision with previous path information. Extensive experiments demonstrate the effectiveness of our proposed model compared to existing methods. Yiming Wu 0005, Hangfei Li, Yilong Zhang 0001, Ronghua Liang |
ICASSP | 5 |
| 2024 | Visual-guided Query with Temporal Interaction for Video Object SegementationabstractThe task of referring video object segmentation (RVOS) involves segmenting objects in video frames based on a given text description. However, most existing approaches treat the text directly as a query, neglecting the valuable visual and temporal information from the video. This limitation may cause the query unable to accurately perceive the target object. To address this issue, we introduce a visual-guided query with temporal interaction for referring video object segmentation (VQTI) approach. Our method capitalizes on frame-level features and video-level features to guide the query generation process, resulting in an enhanced perception of the target object. In addition, we introduce a spectral-guided segmentation optimizer module to enhance the fine-grained information, leading to more precise segmentation masks. Extensive experiments shows competitive performance against state-of-the-art approaches. Jiaxin Qiu, Guoyu Yang, Jie Lei 0002, Zunlei Feng, Ronghua Liang |
ICME | 5 |
| 2024 | MFCA: Multimodal Object Detection Based on Feature Calibration and Aggregation
Jie Lei 0002, Guoyu Yang, Zunlei Feng, Ronghua Liang |
ICONIP (8) | 7 |
| 2024 | Task-Agnostic Self-Distillation for Few-Shot Action Recognition
Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Nan Gao 0001, Ruohong Huan, Xiaofei He 0001 |
IJCAI | 4 |
| 2024 | RE-IDVIS: Person Re-Identification System based on Interactive Visualizationabstractpixel-based visual encoding attribute-based visual encoding image-based visual encoding Figure 1: The interface of the system.(A) the probe panel which allows users to select the person-of-interest as a probe and set up the visual parameter of the search space.(B) the ranking list composed of the pixel-based visual encoding which allows users to quickly retrieve strong negative samples.(C) the search space view which supports visual exploration and enables users to provide feedback on the samples.(D) the spatiotemporal view which summarizes the spatiotemporal information of the retrieval results.(E) a cluster sample at three different visualization scales. Guodao Sun, Pan Liang, Sujia Zhu, Yiming Wu 0005, Haoran Liang 0001, Ronghua Liang |
ICMR | 8 |
| 2024 | Towards Small Object Editing: A Benchmark Dataset and A Training-Free ApproachabstractA plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in producing high-quality images, their application to small object generation has been limited due to difficulties in aligning cross-modal attention maps between text and these objects. Our approach offers a training-free method that significantly mitigates this alignment issue with local and global attention guidance, enhancing the model's ability to accurately render small objects in accordance with textual descriptions. We detail the methodology in our approach, emphasizing its divergence from traditional generation techniques and highlighting its advantages. What's more important is that we also provide SOEBench (Small Object Editing), a standardized benchmark for quantitatively evaluating text-based small object generation collected from MSCOCO[22] and OpenImage[18]. Preliminary results demonstrate the effectiveness of our method, showing marked improvements in the fidelity and accuracy of small object generation compared to existing models. This advancement not only contributes to the field of AI and computer vision but also opens up new possibilities for applications in various industries where precise image generation is critical.We will release our dataset on our project page: https://soebench.github.io/ Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Yiming Wu 0005, Wei Ji 0008, Haoran Liang 0001, Ronghua Liang |
ACM Multimedia | 8 |
| 2024 | Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality AssessmentabstractCurrent state-of-the-art video quality assessment (VQA) models typically integrate various perceptual features to comprehensively represent video quality degradation. These models either directly concatenate features or fuse different perceptual scores while ignoring the domain gaps between cross-aware features, thus failing to adequately learn the correlations and interactions between different perceptual features. To this end, we analyze the independent effects and information gaps of quality-and semantic-aware features on video quality. Based on an analysis of the spatial and temporal differences between two aware features, we propose a semantic-Aware and quality-Aware Interaction Network (A2INet) for blind VQA. For spatial gaps, we introduce a cross-aware guided interaction module to enhance the interaction between semantic-and quality-aware features in a local-to-global manner. Considering temporal discrepancies, we design a cross-aware temporal modeling module to further perceive temporal content variation and quality saliency information, and perceptual features are regressed into quality score by a temporal network and a temporal pooling. Extensive experiments on six benchmark VQA datasets show that our model achieves state-of-the-art performance, and ablation studies further validate the effectiveness of each module. We also present a simple video sampling strategy to balance the effectiveness and efficiency of the model. The code for the proposed method will be released at https://github.com/JianjunXiang/A2INet. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan, Nan Gao 0001 |
ACM Multimedia | 4 |
| 2024 | Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionabstractTemporal relation modeling is one of the core aspects of few-shot action recognition. Most previous works mainly focus on temporal relation modeling based on coarse-level actions, without considering the atomic action details and fine-grained temporal information. This oversight represents a significant limitation in this task. Specifically, coarse-level temporal relation modeling can make the few-shot models overfit in high-discrepancy temporal context, and ignore the low-discrepancy but high-semantic relevance action details in the video. To address these issues, we propose a saliency-guided fine-grained temporal mask learning method that models the temporal atomic action relation for few-shot action recognition in a finer manner. First, to model the comprehensive temporal relations of video instances, we design a temporal mask learning architecture to automatically search for the best matching of each atomic action snippet. Next, to exploit the low-discrepancy atomic action features, we introduce a saliency-guided temporal mask module to adaptively locate and excavate the atomic action information. After that, the few-shot predictions can be obtained by feeding the embedded rich temporal-relation features to a common feature matcher. Extensive experimental results on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods. Yuanjie Dang, Peng Chen 0008, Ruohong Huan, Ronghua Liang |
ACM Multimedia | 6 |
| 2024 | Contextual Augmentation with Bias Adaptive for Few-Shot Video Object Segmentation
Shuaiwei Wang, Jie Lei 0002, Zunlei Feng, Ronghua Liang |
MMM (1) | 7 |
| 2024 | Dynamic-Static Graph Convolutional Network for Video-Based Facial Expression Recognition
Fahong Wang, Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang |
MMM (2) | 9 |
| 2024 | Hierarchical multi-modal video summarization with dynamic samplingabstractAbstract Previous video summarization methods often neglected inter‐frame variations during the preprocessing stage. Sampling repeated frames can lead to information redundancy, while missing key frames can result in deviations in semantic comprehension and inaccuracies in the generated summaries. This work proposes a dynamic sampling module that leverages frame‐level motion information to alleviate these issues. The module conducts high‐frequency sampling during intervals with significant changes, allowing for a finer capture of details. Combined with a hierarchical multi‐modal structure, it integrates shot‐level visual and textual information to enhance the semantic understanding of video clips and improve the accuracy of the summarized content. Extensive experiments on benchmark datasets SumMe and TVSum demonstrate the effectiveness of the proposed method. Lingjian Yu, Xing Zhao 0001, Liang Xie 0003, Haoran Liang 0001, Ronghua Liang |
IET Image Process. | 5 |
| 2024 | Salient object detection in egocentric videosabstractAbstract In the realm of video salient object detection (VSOD), the majority of research has traditionally been centered on third‐person perspective videos. However, this focus overlooks the unique requirements of certain first‐person tasks, such as autonomous driving or robot vision. To bridge this gap, a novel dataset and a camera‐based VSOD model, CaMSD , specifically designed for egocentric videos, is introduced. First, the SalEgo dataset, comprising 17,400 fully annotated frames for video salient object detection, is presented. Second, a computational model that incorporates a camera movement module is proposed, designed to emulate the patterns observed when humans view videos. Additionally, to achieve precise segmentation of a single salient object during switches between salient objects, as opposed to simultaneously segmenting two objects, a saliency enhancement module based on the Squeeze and Excitation Block is incorporated. Experimental results show that the approach outperforms other state‐of‐the‐art methods in egocentric video salient object detection tasks. Dataset and codes can be found at https://github.com/hzhang1999/SalEgo . Haoran Liang 0001, Xing Zhao 0001, Jian Liu 0053, Ronghua Liang |
IET Image Process. | 5 |
| 2024 | TLCSFI: A Pose-Guided Person Re-Identification Method with Two-Level Channel-Spatial Feature IntegrationabstractPerson re-identification methods currently encounter challenges in feature learning, primarily due to difficulties in expressing the correlation between local features and integrating global and local features effectively. To address these issues, a pose-guided person re-identification method with Two-Level Channel–Spatial Feature Integration (TLCSFI) is proposed. In TLCSFI, a two-level integration mechanism is implemented. At the first level, TLCSFI integrates the spatial information from local features to generate fine-grained spatial features. At the second level, the fine-grained spatial feature and the coarse-grained channel feature are integrated together to complete channel–spatial feature integration. In the method, a Pose-based Spatial Feature Integration (PSFI) module is introduced to generate the pose union feature, which calculates intra-body affinity to guide the integration of spatial information among local pose feature maps. Then, a Channel and Spatial Union Feature Integration (CSUFI) module is proposed to efficiently integrate the channel information of the global feature and the spatial information of the pose union feature. Two individual networks are designed to extract channel and spatial information, respectively, in CSUFI, which are then weighted and integrated. Experiments are conducted on three publicly available datasets to evaluate TLCSFI, and the experimental results demonstrate its competitive performance. Ruohong Huan, Nan Gao 0001, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2024 | Learning Reliable Dense Pseudo-Labels for Point-Level Weakly-Supervised Action LocalizationabstractAbstract Point-level weakly-supervised temporal action localization aims to accurately recognize and localize action segments in untrimmed videos, using only point-level annotations during training. Current methods primarily focus on mining sparse pseudo-labels and generating dense pseudo-labels. However, due to the sparsity of point-level labels and the impact of scene information on action representations, the reliability of dense pseudo-label methods still remains an issue. In this paper, we propose a point-level weakly-supervised temporal action localization method based on local representation enhancement and global temporal optimization. This method comprises two modules that enhance the representation capacity of action features and improve the reliability of class activation sequence classification, thereby enhancing the reliability of dense pseudo-labels and strengthening the model’s capability for completeness learning. Specifically, we first generate representative features of actions using pseudo-label feature and calculate weights based on the feature similarity between representative features of actions and segments features to adjust class activation sequence. Additionally, we maintain the fixed-length queues for annotated segments and design a action contrastive learning framework between videos. The experimental results demonstrate that our modules indeed enhance the model’s capability for comprehensive learning, particularly achieving state-of-the-art results at high IoU thresholds. Yuanjie Dang, Guozhu Zheng, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Ronghua Liang |
Neural Process. Lett. | 7 |
| 2024 | ZJUT-EIFD: A Synchronously Collected External and Internal Fingerprint DatabaseabstractExternal fingerprints (EFs) based only on epidermal information are vulnerable to spoofing attacks and non-ideal skin conditions. To solve such shortcomings, internal fingerprints (IFs) collected using optical coherence tomography (OCT) have been proposed and widely researched. However, the development of IF is limited by the lack of in-depth researches on the IF and the EF-IF interoperability, which is partially caused by the lack of public OCT database. The obvious gap in the applications of EF and IF recognition motivated us to design and publish a comprehensive fingerprint database containing both traditional EFs and OCT IFs, denoted as ZJUT-EIFD. To the best of our knowledge, ZJUT-EIFD is the first public database that combines OCT and total internal reflection (TIR) via synchronous acquisition, with 399 different fingers from 60 subjects. In this article, the composition of the database, the quality of EFs and IFs, and the verification performance of different types of fingerprints were detailed. In addition, potential application directions of ZJUT-EIFD were demonstrated. ZJUT-EIFD can serve benchmarks and interoperability tests for EF-IF research, which will promote the research and development of EF and IF. Haohao Sun, Haixia Wang 0002, Yilong Zhang 0001, Ronghua Liang, Peng Chen 0008, Jianjiang Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | TriSAT: Trimodal Representation Learning for Multimodal Sentiment AnalysisabstractTransformer-based multimodal sentiment analysis frameworks commonly facilitate cross-modal interactions between two modalities through the attention mechanism. However, such interactions prove inadequate when dealing with three or more modalities, leading to increased computational complexity and network redundancy. To address this challenge, this paper introduces a novel framework, Trimodal representations for Sentiment Analysis from Transformers (TriSAT), tailored for multimodal sentiment analysis. TriSAT incorporates a trimodal transformer featuring a module called Trimodal Multi-Head Attention (TMHA). TMHA considers language as the primary modality, combines information from language, video, and audio using a single computation, and analyzes sentiment from a trimodal perspective. This approach significantly reduces the computational complexity while delivering high performance. Moreover, we propose Attraction-Repulsion (AR) loss and Trimodal Supervised Contrastive (TSC) loss to further enhance sentiment analysis performance. We conduct experiments on three public datasets to evaluate TriSAT's performance, which consistently demonstrates its competitiveness compared to state-of-the-art approaches. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | A Distributed and Parallel Accelerator Design for 3-D Acoustic Imaging on FPGA-Based Systemsabstract3-D imaging sonar is crucial in the exploration of marine resources, and the development of portable device with high imaging quality and high real-time performance is the general trend. However, traditional framework methods are limited by the huge amount of computation brought by high-quality imaging, making it difficult to implement in engineering. To address this issue, we develop 3-D real-time sonar system in an algorithm-hardware co-designed way. An ultrawideband distributed and parallel subarray beamforming algorithm (UWBDPS) is proposed for 3-D acoustic imaging. This is a multi-stage array time-frequency beamforming method under a distributed parallel computing architecture. Based on this, we propose field-programmable gate array (FPGA)-based accelerator. It divides a large sonar receiving planar array into multiple parallel subarrays, and complete the beamforming in two stages, which can reduce the calculation load and speeds up 3-D imaging. For engineering implementation, we optimized the sparseness of the planar transducer array, with a sparse rate as high as 97.7%. The experimental results show that the calculation amount of the proposed UWB-DPS algorithm is reduced to 1/5.7 of the traditional framework algorithm, the imaging performance is effectively improved, and the FPGA-based accelerator outperforms the CPU software implementation by 935×. Weibo Mao, Peng Chen 0008, Yingtian Hu, Haoran Liang 0001, Yuanjie Dang, Ronghua Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | Video Visualization and Visual Analytics: A Task-Based and Application- Driven InvestigationabstractVideo data refers to digital information in the form of a series of frames or images representing continuous motion captured by a video recording device. In various domains such as security, sports, education, and entertainment, a significant amount of video data is generated and stored daily. However, analyzing these videos manually is challenging due to their intrinsic characteristics, including large-scale, redundancy, contextual dependencies, and multimodality. Consequently, researchers have extensively explored visualization techniques to address these complexities. In this investigation, we review the state-of-the-art techniques in video visualization and visual analysis. Initially, we provide an overview of the design space for video visualization and visual analysis techniques. Subsequently, we organize and classify these techniques based on visual analysis tasks and application scenarios, providing detailed descriptions within each category. Drawing upon a comprehensive review of existing research, we provide a critical evaluation and propose potential opportunities for future research. Additionally, we have developed a web-based survey browser for convenient exploration of our created classification framework and the associated scholarly articles (https://zjutvis.github.io/VOVideo/). Guodao Sun, Baofeng Chang, Jingwei Tang, Gefei Zhang 0002, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | A Visual Representation-Guided Framework With Global Affinity for Weakly Supervised Salient Object DetectionabstractFully supervised salient object detection (SOD) methods have made considerable progress in performance, yet these models rely heavily on expensive pixel-wise labels. Recently, to achieve a trade-off between labeling burden and performance, scribble-based SOD methods have attracted increasing attention. Previous scribble-based models directly implement the SOD task only based on SOD training data with limited information, it is extremely difficult for them to understand the image and further achieve a superior SOD task. In this paper, we propose a simple yet effective framework guided by general visual representations with rich contextual semantic knowledge for scribble-based SOD. These general visual representations are generated by self-supervised learning based on large-scale unlabeled datasets. Our framework consists of a task-related encoder, a general visual module, and an information integration module to efficiently combine the general visual representations with task-related features to perform the SOD task based on understanding the contextual connections of images. Meanwhile, we propose a novel global semantic affinity loss to guide the model to perceive the global structure of the salient objects. Experimental results on five public benchmark datasets demonstrate that our method, which only utilizes scribble annotations without introducing any extra label, outperforms the state-of-theart weakly supervised SOD methods. Specifically, it outperforms the previous best scribble-based method on all datasets with an average gain of 5.5% for max f-measure, 5.8% for mean f-measure, 24% for MAE, and 3.1% for E-measure. Moreover, our method achieves comparable or even superior performance to the state-of-the-art fully supervised models. Binwei Xu, Haoran Liang 0001, Weihua Gong, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Asymptotic Feature Pyramid Network for Labeling Pixels and RegionsabstractMulti-scale features are crucial in encoding objects with varying scales in vision tasks. The classic top-down and bottom-up feature pyramid networks are a common strategy for multi-scale feature extraction. However, these approaches suffer from the loss or degradation of feature information, which impairs the fusion effect of non-adjacent levels. In this paper, we propose an Asymptotic Feature Pyramid Network (AFPN) that supports direct interaction between non-adjacent levels. AFPN starts by fusing two adjacent low-level features and asymptotic incorporates higher-level features into the fusion process. This fusion way avoids the significant semantic gap between non-adjacent levels. Adaptive spatial fusion operation is further used to mitigate potential multi-object information conflicts during feature fusion at each spatial location. To reduce parameters, computational requirements, and inference speed, we propose a Lightweight Asymptotic Feature Pyramid Network (LightAFPN) that uses the concept of reparametrization. We evaluate the proposed method on the MS-COCO 2017, PASCAL VOC and Cityscapes datasets in both object detection and semantic segmentation frameworks. Experimental evaluation shows that our method achieves more competitive results than other state-of-the-art feature pyramid networks. The code is available at https://github.com/gyyang23/AFPN. Guoyu Yang, Jie Lei 0002, Zunlei Feng, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | LANDER: Visual Analysis of Activity and Uncertainty in Surveillance VideoabstractVision algorithms face challenges of limited visual presentation and unreliability in pedestrian activity assessment. In this article, we introduce LANDER, an interactive analysis system for visual exploration of pedestrian activity and uncertainty in surveillance videos. This visual analytics system focuses on three common categories of uncertainties in object tracking and action recognition. LANDER offers an overview visualization of activity and uncertainty, along with spatio-temporal exploration views closely associated with the scene. Expert evaluation and user study indicate that LANDER outperforms traditional video exploration in data presentation and analysis workflow. Specifically, compared to the baseline method, it excels in reducing retrieval time ($p< $0.01), enhancing uncertainty identification ($p< $0.05), and improving the user experience ($p< $0.05). Guodao Sun, Baofeng Chang, Yunchao Wang, Yuanzhong Ying, Haixia Wang 0002, Ronghua Liang |
IEEE Trans. Hum. Mach. Syst. | 9 |
| 2024 | A Wavelet-Based Memory Autoencoder for Noncontact Fingerprint Presentation Attack DetectionabstractFingerprint presentation attack detection (FPAD) is essential in fingerprint identification systems. Noncontact methods such as fingerprint biometrics are becoming popular because they are not affected by skin conditions and there are no hygiene issues. However, most of the existing noncontact FPAD methods are supervised methods with poor generalizability and poor performance during events such as unseen presentation attacks (PAs). Moreover, easily overlooked frequency domain information contributes to the fingerprint antispoofing task. Therefore, we propose a wavelet-based memory-augmented autoencoder that fully utilizes the frequency domain information. Specifically, the model first decomposes the input image into high- and low-frequency information and extracts features separately. Subsequently, we propose a frequency complementary connection (FCC) module to realize the fusion and complementation of frequency domain information at the feature level. Moreover, a memory distance expansion loss is proposed to keep the memory module diverse. Experiments are conducted to verify the effectiveness of the method. The code of our model is available onhttps://github.com/SuperIOyht/WaveMemAE. Yipeng Liu 0002, Hangtao Yu, Haonan Fang, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Motion-Aware Memory Network for Fast Video Salient Object DetectionabstractPrevious methods based on 3DCNN, convLSTM, or optical flow have achieved great success in video salient object detection (VSOD). However, these methods still suffer from high computational costs or poor quality of the generated saliency maps. To address this, we design a space-time memory (STM)-based network that employs a standard encoder-decoder architecture. During the encoding stage, we extract high-level temporal features from the current frame and its adjacent frames, which is more efficient and practical than methods reliant on optical flow. During the decoding stage, we introduce an effective fusion strategy for both spatial and temporal branches. The semantic information of the high-level features is used to improve the object details in the low-level features. Subsequently, spatiotemporal features are methodically derived step by step to reconstruct the saliency maps. Moreover, inspired by the boundary supervision prevalent in image salient object detection (ISOD), we design a motion-aware loss that predicts object boundary motion, and simultaneously perform multitask learning for VSOD and object motion prediction. This can further enhance the model's capability to accurately extract spatiotemporal features while maintaining object integrity. Extensive experiments on several datasets demonstrate the effectiveness of our method and can achieve state-of-the-art metrics on some datasets. Our proposed model does not require optical flow or additional preprocessing, and can reach an impressive inference speed of nearly 100 FPS. Xing Zhao 0001, Haoran Liang 0001, Guodao Sun, Ronghua Liang, Xiaofei He 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | E2Storyline: Visualizing the Relationship with Triplet Entities and Event DiscoveryabstractThe narrative progression of events, evolving into a cohesive story, relies on the entity-entity relationships. Among the plethora of visualization techniques, storyline visualization has gained significant recognition for its effectiveness in offering an overview of story trends, revealing entity relationships, and facilitating visual communication. However, existing methods for storyline visualization often fall short in accurately depicting the specific relationships between entities. In this study, we present E 2 Storyline, a novel approach that emphasizes simplicity and aesthetics of layout while effectively conveying entity-entity relationships to users. To achieve this, we begin by extracting entity-entity relationships from textual data and representing them as subject-predicate-object (SPO) triplets, thereby obtaining structured data. By considering three types of design requirements, we establish new optimization objectives and model the layout problem using multi-objective optimization (MOO) techniques. The aforementioned SPO triplets, together with time and event information, are incorporated into the optimization model to ensure a straightforward and easily comprehensible storyline layout. Through a qualitative user study, we determine that a pixel-based view is the most suitable method for displaying the relationships between entities. Finally, we apply E 2 Storyline to real-world data, including movie synopses and live text commentaries. Through comprehensive case studies, we demonstrate that E 2 Storyline enables users to better extract information from stories and comprehend the relationships between entities. Yunchao Wang, Guodao Sun, Ronghua Liang |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2024 | Multi-Level Objective Alignment Transformer for Fine-Grained Oral Panoramic X-Ray Report GenerationabstractAutomatically generated oral panoramic X-ray report is highly beneficial for improving the efficiency of dental diagnosis. However, recent solutions adopt holistic methods, resulting in a cursory description of the oral condition. This may lead to reports lacking details, such as specific sites or lesion contours. Therefore, we propose a Multi-Level objective Alignment Transformer(MLAT) network, which integrates all tooth and disease objects into a positional alignment graph to extract fine-grained object-level features. Specifically, we introduce a novel Object-Level Collaborative Encoder (OLCE) module, which uses a positional alignment graph to construct object relationships. OLCE enhances object-level feature extraction by eliminating interference information between pathologically unrelated objects. In addition, we build a high-quality panoramic X-ray image-report dataset consisting of 562 sets of images and reports labeled by 13 experienced dental specialists. Experiments on the collected dataset show that the proposed MLAT significantly outperforms the state-of-the-art baselines by more than 5% in 4 different metrics, including BLEUs, Meteor, Rouge, and BERTScore. Nan Gao 0001, Renyuan Yao, Ronghua Liang, Peng Chen 0008, Tianshuang Liu, Yuanjie Dang |
IEEE Trans. Multim. | 3 |
| 2024 | UniMF: A Unified Multimodal Framework for Multimodal Sentiment Analysis in Missing Modalities and Unaligned Multimodal SequencesabstractIn current multimodal sentiment analysis, aligned and complete multimodal sequences are often crucial. Obtaining complete multimodal data in the real world presents various challenges, and aligning multimodal sequences often requires a significant amount of effort. Unfortunately, most multimodal sentiment analysis methods fail when dealing with missing modalities or unaligned multimodal sequences. To tackle these two challenges simultaneously in a simple and lightweight manner, we present the Unified Multimodal Framework (UniMF). The primary components of UniMF comprise two distinct modules. The first module, Translation Module, translates missing modalities using information from existing modalities. The second module, Prediction Module, uses the attention mechanism to fuse the multimodal information and generate predictions. To enhance the translation performance of the Translation Module, we introduce the Multimodal Generation Mask (MGM) and utilize it to construct the Multimodal Generation Transformer (MGT). The MGT can generate the missing modality while focusing on information from existing modalities. Furthermore, we introduce the Multimodal Understanding Transformer (MUT) in the Prediction Module, which includes the Multimodal Understanding Mask (MUM) and a unique sequence,MultiModalSequence(MMSeq), representing a unified multimodality. To assess the performance of UniMF, we perform experiments on four multimodal sentiment datasets, and UniMF attains competitive or state-of-the-art outcomes with fewer learnable parameters. Furthermore, the experimental outcomes signify that UniMF, supported by MGT and MUT - two transformers utilizing special attention mechanisms, can efficiently manage both generating task of missing modalities and understanding task of unaligned multimodal sequences. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Multim. | 4 |
| 2024 | Pseudo Light Field Image and 4D Wavelet-Transform-Based Reduced-Reference Light Field Image Quality AssessmentabstractReduced-reference light field image (LFI) quality assessment (RR LFIQA) automatically assesses image quality with only partial information about the reference LFI is available. Existing RR LFIQA has difficulty extracting effective RR information and perceptual features to represent the LFI quality. In this article, we propose an RR LFIQA model based on pseudo LFI (PLFI) and four-dimensional (4D) wavelet transform. To extract RR information related to LFI perceptual quality, a PLFI is created as the RR information of the LFI using a view synthesis algorithm. Considering that the high-dimensional characteristics of the PLFI, 4D wavelet transform is used to decompose the original and distorted PLFIs. The 4D wavelet transform essentially performs a continuous 1D wavelet transform for the 4D signal to enable the local 4D structure of the PLFIs to be characterized effectively in the 4D wavelet domain. A novel spatial-angular weighting strategy is proposed to describe the importance of each location for quality evaluation, to further improve the performance of the proposed method. Experimental results on four benchmark datasets show that the proposed model performs better than the representative 2DIQA and LFIQA models. Jianjun Xiang, Peng Chen 0008, Yuanjie Dang, Ronghua Liang, Gangyi Jiang |
IEEE Trans. Multim. | 4 |
| 2024 | Synthesize Boundaries: A Boundary-Aware Self-Consistent Framework for Weakly Supervised Salient Object DetectionabstractFully supervised salient object detection (SOD) has made considerable progress based on expensive and time-consuming data with pixel-wise annotations. Recently, to relieve the labeling burden while maintaining performance, some scribble-based SOD methods have been proposed. However, learning precise boundary details from scribble annotations that lack edge information is still difficult. In this article, we propose to learn precise boundaries from our designed synthetic images and labels without introducing any extra auxiliary data. The synthetic image creates boundary information by inserting synthetic concave regions that simulate the real concave regions of salient objects. Furthermore, we propose a novel self-consistent framework that consists of a global integral branch (GIB) and a boundary-aware branch (BAB) to train a saliency detector. GIB aims to identify integral salient objects, whose input is the original image. BAB aims to help predict accurate boundaries, whose input is the synthetic image. These two branches are connected through a self-consistent loss to guide the saliency detector to predict precise boundaries while identifying salient objects. Experimental results on five benchmarks demonstrate that our method outperforms the state-of-the-art weakly supervised SOD methods and further narrows the gap with the fully supervised methods. Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 3 |
| 2024 | Discriminative Action Snippet Propagation Network for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to classify and localize actions in untrimmed videos with only video-level labels. Recent studies have attempted to obtain more accurate temporal boundaries by exploiting latent action instances in ambiguous snippets or propagating representative action features. However, empirically handcrafted ambiguous snippet extraction and the imprecise alignment of representative snippet propagation lead to challenges in modeling the completeness of actions for these methods. In this article, we propose a Discriminative Action Snippet Propagation Network (DASP-Net) to accurately discover ambiguous snippets in videos and propagate discriminative instance-level features throughout the video for improving action completeness. Specifically, we introduce a novel discriminative feature propagation module for capturing the global contextual attention and propagating the action concept across the whole video by perceiving the discriminative action snippets with instance information from the same video. Simultaneously, we incorporate denoised pseudo-labels as supervision, where we correct the controversial prediction based on the feature space distribution during training, thereby alleviating false detection caused by noise background features. Furthermore, we design an ambiguous feature mining module, which maximizes the feature affinity information of action and background in ambiguous snippets to generate more accurate latent action and background snippets and learns more precise action instance boundaries through contrastive learning of action and background snippets. Extensive experiments show that DASP-Net achieves state-of-the-art results on THUMOS14 and ActivityNet1.2 datasets. Yuanjie Dang, Chunxia Huang, Peng Chen 0008, Nan Gao 0001, Ronghua Liang, Ruohong Huan |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | BTCN: Bridging the Gap Between Pre-trained and Downstream Models for Endoscopic Caries DetectionabstractAlthough deep learning has been widely applied in the field of dental caries detection, there are still certain challenges that need to be addressed. The limitations of sharing the same backbone between the pre-trained model and the downstream model hinder the feature alignment capability of self-supervised learning (SSL) during the fine-tuning stage, leading to incomplete transfer from the pre-trained model to the downstream model. To address this challenge, we introduce an SSL pre-trained model called Bi-branches Transformer CNN Network (BTCN). BTCN adopts a parallel structure combining the CNN and Transformer branches. This parallel structure allows the pre-trained model to capture additional global representations, which helps alleviate feature differences during fine-tuning and better adapt to downstream detection models. Additionally, to further enhance the fusion quality of the bi-branches encoder, we introduced the Multi-layer Supervision Strategy (MSS) to increase the supervision on features at different layers. To validate the effectiveness of our approach, we collected a dedicated dataset for caries detection, comprising 1039 endoscopic images of dental caries. Through extensive experimental research, our results demonstrate the effectiveness of the proposed BTCN and MSS, showing significant improvements compared to the current state-of-the-art methods. Nan Gao 0001, Peng Chen 0008, Yukai Li, Jijun Tang, Ronghua Liang, Tianshuang Liu |
BIBM | 5 |
| 2023 | Spatial-angular Quality-aware Representation Learning for Blind Light Field Image Quality AssessmentabstractBlind light field image quality assessment (BLFIQA) remains a challenging task in deep learning due to the unique spatial-angular structure of light field images (LFIs) and the lack of large-scale labeled data for training. In this work, we propose a novel BLFIQA method using spatial-angular quality-aware representation learning in a self-supervised learning manner. Visual content and distortion type are important factors affecting the perceived quality of LFIs. In our observation, the band-pass transform maps of LFIs with the same distortion type exhibit similar Gaussian distributions. Thus, we learn spatial-angular quality-aware representations by minimizing the distance in the embedding space between the luminance map and the band-pass transform map of the same LFI. To implement spatial-angular quality-aware representations of LFI, we also build a large-scale unlabeled dataset containing 40k distorted LFIs with different distortion types and visual content. Further, we propose a fusion-separation-fusion network (FSFNet) to extract features for representing the intrinsic spatial-angular structure of the LFI. After pre-training on the unlabeled dataset using the proposed self-supervised learning, the FSFNet is employed for downstream BLFIQA tasks and achieves good performance. Experimental results show that our proposed method outperforms seventeen state-of-the-art models on the Win5-LID, NBU-LF1.0 and LFDD datasets, and achieves 3.78%, 6.61% and 4.06% SRCC improvements, respectively. The code and dataset will be publicly available in https://github.com/JianjunXiang/SSL_and_FSFNet. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan |
ACM Multimedia | 4 |
| 2023 | Multi-Speed Global Contextual Subspace Matching for Few-Shot Action RecognitionabstractFew-shot action recognition (FSAR) aims to classify unseen query actions into categories represented by a few labeled support videos. Most current FSAR methods adopt the frame-level matching mechanism that requires continuous actions to be represented by a fixed number of frame features. However, this could compromise the completeness of the contextual video information and make it difficult to handle video features of varying frame sampling speeds. In this paper, we propose a multi-speed global contextual subspace matching (MGCSM) method that generates global contextual action subspace representations from videos containing different numbers of frames to preserve contextual semantic information. Specifically, we propose to obtain the scale-agnostic information of embedding video features using a global contextual aggregation (GCA) module and then generate the discriminative action subspace representation with an action subspace generation (ASG) module. Furthermore, we introduce a multi-speed subspace matching (MSM) mechanism that generates a multi-speed classification score by integrating the similarities between query videos and support subspaces of varying sampling speeds. The proposed method is embedding-agnostic and can be combined with most mainstream embedding networks without model re-designs. Comprehensive and reproducible experiments on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods. Tianwei Yu, Peng Chen 0008, Yuanjie Dang, Ruohong Huan, Ronghua Liang |
ACM Multimedia | 5 |
| 2023 | HPAN: A Hybrid Pose Attention Network for Person Re-Identification
Ruohong Huan, Tianya Chen, Ziwei Zhan, Peng Chen 0008, Ronghua Liang |
PRCV (12) | 5 |
| 2023 | Temporal Aggregation with Context Focusing for Few-Shot Video Object DetectionabstractFew-shot video object detection focuses on finding all the objects in a given query video that belong to the same class, given only a few support images of the target object in an unseen class. Unfortunately, due to the object blur or occlusion in video frames, using single-frame object detection directly will greatly limit the accuracy. The issue is significantly worse in few-shot settings due to insufficient support and timedomain information. In this paper, we propose a temporal aggregation with context focusing framework (TACF) for few-shot video object detection, which aims to fully use the information between support images and adjacent video frames. The context focusing module effectively encodes the target object in adjacent frames according to the support images. Afterward, the temporal aggregation module implicitly extracts the most similar ROI features from these adjacent frames to obtain the target proposals. In the end, the matching network determines the category and bounding box by calculating the distance with the support images. Extensive experimental evaluations on FSVOD and FSYTV databases show that our method achieves more competitive results than image-based methods, naive video-based extensions, and the state-of-the-art few-shot video object detection method. Jie Lei 0002, Fahong Wang, Zunlei Feng, Ronghua Liang |
SMC | 5 |
| 2023 | AFPN: Asymptotic Feature Pyramid Network for Object DetectionabstractMulti-scale features are of great importance in encoding objects with scale variance in object detection tasks. A common strategy for multi-scale feature extraction is adopting the classic top-down and bottom-up feature pyramid networks. However, these approaches suffer from the loss or degradation of feature information, impairing the fusion effect of non-adjacent levels. This paper proposes an asymptotic feature pyramid network (AFPN) to support direct interaction at non-adjacent levels. AFPN is initiated by fusing two adjacent low-level features and asymptotically incorporates higher-level features into the fusion process. In this way, the larger semantic gap between non-adjacent levels can be avoided. Given the potential for multi-object information conflicts to arise during feature fusion at each spatial location, adaptive spatial fusion operation is further utilized to mitigate these inconsistencies. We incorporate the proposed AFPN into both two-stage and one-stage object detection frameworks and evaluate with the MS-COCO 2017 validation and test datasets. Experimental evaluation shows that our method achieves more competitive results than other state-of-the-art feature pyramid networks. The code is available at https://github.com/gyyang23/AFPN. Guoyu Yang, Jie Lei 0002, Zhikuan Zhu, Siyu Cheng, Zunlei Feng, Ronghua Liang |
SMC | 6 |
| 2023 | Anti-spoofing study on palm biometric features
Haixia Wang 0002, Lixun Su, Hongxiang Zeng, Peng Chen 0008, Ronghua Liang, Yilong Zhang 0001 |
Expert Syst. Appl. | 5 |
| 2023 | SS-Norm: Spectral-spatial normalization for single-domain generalization with application to retinal vessel segmentationabstractAbstract Retinal vessel segmentation is an important computer vision task for eye retinopathy diagnosis. In the real scenarios, most datasets of source domain and target domain have distribution deviation, and the model often fails to generate accurate segmentation results due to the lack of data variation in single‐source domain, which damages the generalization ability to unseen target domains and may mislead doctors or artificial intelligence model in the following diseases diagnosis. Feature normalization is one feasible solution which can standardize data into uniform and stable distribution without additional data. However, the existing methods like batch normalization, uniform the data by global parameters. This leads to insufficient representation of important semantic information in the local region. To address this problem, the authors propose the spectral‐spatial normalization (SS‐Norm) module to enhance the generalization ability of the model. More specifically, the authors perform a discrete cosine transform (DCT) to decompose the feature into multiple frequency components and to analyze the semantic contribution degree of each component. By learning a spectral vector, the authors reweight the frequency components of features and therefore normalize the distribution in the spectral domain. Extensive experiments on six datasets prove the effectiveness of the authors’ methods. Yipeng Liu 0002, Dongxu Zeng, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IET Image Process. | 5 |
| 2023 | A progressive segmentation with weight contrast label enhancement for weakly supervised video salient object detectionabstractAbstract Scribble labels have gained increasing attention in the field of weakly supervised video salient object detection (VSOD). Based on scribble labels, latest methods can spread labeled pixels to unlabeled regions using local coherence loss, but predicted objects often lose detail and boundary information. In this work, a novel method based on back‐foreground weight contrast is proposed that adds label enhancement points to facilitate the model to learn the edge, detail and location of salient object. Additionally, a new VSOD framework based on global structural localization is introduced. Enhanced scribble labels are used to assist the model for global localization, and then the located regions are finely segmented by the trained model. Extensive experiments demonstrate that the method achieves the state‐of‐the‐art performance on common VSOD datasets, with an improvement of 3.75%, 4.68%, and 0.88% in S‐measure, F‐measure, and MAE, respectively. Zelin Lu, Haoran Liang 0001, Binwei Xu, Ronghua Liang |
IET Image Process. | 4 |
| 2023 | DOMOPT: A Detection-Based Online Multi-Object Pedestrian Tracking Network for VideosabstractDue to the problem of low tracking accuracy and weak tracking stability of current multi-object pedestrian tracking algorithms in complex scenes for videos, a Detection-based Online Multi-Object Pedestrian Tracking (DOMOPT) network is proposed. First, a Multi-Level Feature Fusion (MLFF) pedestrian detection network is proposed based on the Center and Scale Prediction (CSP) algorithm. The pyramid convolutional neural network is used as the backbone to enhance the feature extraction capability for small objects. The shallow features and deep features at multiple levels are integrated to fully obtain the position and semantic information to further improve the detection performance for small objects. Then, on the basis of Joint Detection and Embedding (JDE) architecture, a Multi-Branch Pedestrian Appearance (MBPA) feature extraction network is proposed and added into the pedestrian detection network to extract the appearance feature vector corresponding to each pedestrian. The pedestrian appearance feature extraction is treated as a classification task jointly training with the pedestrian detection task, using the multi-task learning strategy. Experimental results show that the proposed network has better tracking accuracy and stability compared with state-of-the-art algorithms. Ruohong Huan, Shuaishuai Zheng, Chaojie Xie, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2023 | DCAM: Disturbed class activation maps for weakly supervised semantic segmentation
Jie Lei 0002, Guoyu Yang, Shuaiwei Wang, Zunlei Feng, Ronghua Liang |
J. Vis. Commun. Image Represent. | 5 |
| 2023 | Visual interactive image clustering: a target-independent approach for configuration optimization in machine vision measurementabstractMachine vision measurement (MVM) is an essential approach that measures the area or length of a target efficiently and non-destructively for product quality control. The result of MVM is determined by its configuration, especially the lighting scheme design in image acquisition and the algorithmic parameter optimization in image processing. In a traditional workflow, engineers constantly adjust and verify the configuration for an acceptable result, which is time-consuming and significantly depends on expertise. To address these challenges, we propose a target-independent approach, visual interactive image clustering, which facilitates configuration optimization by grouping images into different clusters to suggest lighting schemes with common parameters. Our approach has four steps: data preparation, data sampling, data processing, and visual analysis with our visualization system. During preparation, engineers design several candidate lighting schemes to acquire images and develop an algorithm to process images. Our approach samples engineer-defined parameters for each image and obtains results by executing the algorithm. The core of data processing is the explainable measurement of the relationships among images using the algorithmic parameters. Based on the image relationships, we develop VMExplorer, a visual analytics system that assists engineers in grouping images into clusters and exploring parameters. Finally, engineers can determine an appropriate lighting scheme with robust parameter combinations. To demonstrate the effectiveness and usability of our approach, we conduct a case study with engineers and obtain feedback from expert interviews. Lvhan Pan, Guodao Sun, Baofeng Chang, Jingwei Tang, Ronghua Liang |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2023 | MLFFCSP: a new anti-occlusion pedestrian detection network with multi-level feature fusion for small targets
Ruohong Huan, Chaojie Xie, Ronghua Liang, Peng Chen 0008 |
Multim. Tools Appl. | 4 |
| 2023 | Application of Mathematical Optimization in Data Visualization and Visual Analytics: A SurveyabstractMathematical optimization is the process of determining the set of globally or locally optimal parameters in a finite or infinite search space. It has been extensively employed in the research areas of computer science, engineering, operations research, and economics. The application of mathematical optimization has also been extended to data visualization, where it can enhance data processing, structure visualization, and facilitate exploration. However, the current state of summarization in the application of mathematical optimization in data visualization remains inadequate. In this article, we review and classify the existing techniques for advanced mathematical optimization in the fields of data visualization and visual analytics. The classification is conducted based on a classical visualization pipeline, including data enhancement and transformation, representation and rendering, as well as interactive exploration and analysis. We also discuss various mathematical optimization models and their solution methods to help readers gain a better understanding of the relationship among models, visualization, and application scenarios. We additionally provide an online exploration demo, which could enable users to interactively find relevant articles. Based on the limitations and potential trends revealed in the existing literature, we define future challenges in the cross-disciplinary of mathematical optimization and data visualization. Guodao Sun, Gefei Zhang 0002, Chaoqing Xu, Yunchao Wang, Sujia Zhu, Baofeng Chang, Ronghua Liang |
IEEE Trans. Big Data | 8 |
| 2023 | End-to-End Surface and Internal Fingerprint Reconstruction From Optical Coherence Tomography Based on Contour RegressionabstractOptical coherence tomography (OCT), as a non-destructive and high-resolution imaging technique, has been used to collect 3D fingertip data, which contains surface and internal fingerprints. Methods have been proposed for OCT fingerprint reconstruction. However, these methods have complex processing flow and are time consuming. In this paper, an end-to-end convolutional neural network based surface and internal fingerprint reconstruction method is proposed. A simple yet effective contour regression module is proposed and integrated in the network for direct estimation of contours of stratum corneum and viable epidermis junction from noisy OCT volume data, thus greatly simplify the processing flow. The proposed network further integrates multi-task learning with conventional segmentation task as auxiliary task and contour regression task as main task to facilitate the feature extraction and improve the robustness of the network. Depthwise separable convolution is adapted to a light-weight network for network computation complexity reduction. To the best of our knowledge, it is the first time that an end-to-end method is proposed for surface and internal fingerprint extraction from noisy OCT volume data. Experiments and comparisons are carried out in terms of contour estimation accuracy, fingerprint quality, fingerprint matching performance and computation efficiency. Compared with conventional method, the proposed method utilizes only 6% of original network parameters and 0.7% of original computation time, but achieves comparably results. Fingerprint by depth proves the accuracy and robustness of contour regression than pixel-wise layer segmentation. The proposed method is noise-insensitive, process-simple and time-efficient for OCT fingerprint reconstruction, which is significant for real time application in Automated Fingerprint Recognition Systems. Baojin Ding, Haixia Wang 0002, Ronghua Liang, Yilong Zhang 0001, Peng Chen 0008 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Prototype-Guided Autoencoder for OCT-Based Fingerprint Presentation Attack DetectionabstractAnti-spoofing ability is vital for fingerprint identification systems. Conventional fingerprint scanning devices can only obtain information from the fingertip surfaces, and their performance is susceptible to skin conditions and presentation attacks (PAs). However, optical coherence tomography (OCT) can scan subcutaneous tissue and obtain 3D fingerprint structures, naturally enhancing its PA detection (PAD) ability from the perspective of hardware. Existing unsupervised PAD methods are based on image reconstruction. However, the reconstruction error is easily affected by OCT noise and the rich details of OCT images. Therefore we propose feature-based reconstruction to alleviate this problem, called the prototype-guided autoencoder. The model consists of a memory module and a denoising autoencoder without the requirement of PA fingerprints. As only bona fide fingerprints are available during the training phase, the memory module contains the prototype features of the bona fide fingerprints. During the inference phase, as the prototype memory module is frozen, the reconstructed representation of the bona fide input is close to the bona fide fingerprint features. Calculating the distance between the original features and the prototype reconstructed representation of the sample can achieve PAD. To obtain a better decision making boundary, we propose a representation consistency constraint, which reduces the bona fide representation reconstruction distance closer, so that it is easier to differentiate between fingerprints and PAs. Yipeng Liu 0002, Wangyang Zuo, Ronghua Liang, Haohao Sun, Zhanqing Li |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | A New Approach in Automated Fingerprint Presentation Attack Detection Using Optical Coherence TomographyabstractPresentation attack detection (PAD) is a critical component of automated fingerprint recognition systems (AFRSs). However, existing PAD technologies based on optical coherence tomography (OCT) mainly rely on local information, ignoring the global continuity and correlation of physiological structures. Furthermore, the lack of appropriate presentation attack instruments (PAIs) that cater to the unique OCT characteristics leads to the insufficient evaluation of PAD. The identification features, including external fingerprint (EF), internal fingerprint (IF), and subcutaneous sweat pore (SSP), provide valuable information about the intrinsic connections of physiological structures. Such intrinsic connections hold potential clues for PAD. Building upon this premise, this paper proposed a novel PAD method based on three OCT hand-crafted features: EF-IF self-matching score (SMS), SSP number (SN), and SSP coincidence rate (SCR). These simple yet effective PAD features offer a more precise and detailed description of the internal physiological structure, enabling accurate distinction between presentation attack (PA) and bona-fide. The proposed method achieves a 4% Equal Error Rate (EER), significantly outperforming other existing PAD methods. Additionally, the cross-device experiment demonstrates the generalization capability of the proposed method on both our dataset and the public OCT dataset. Haohao Sun, Yilong Zhang 0001, Peng Chen 0008, Haixia Wang 0002, Yipeng Liu 0002, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2023 | Path-Analysis-Based Reinforcement Learning Algorithm for Imitation FilmingabstractImitation filming has been applied to autonomous filming by mimicking human operators. To imitate the operation of cameramen when filming multiple human actions, existing methods plan the camera motion through time series prediction or train multiple models to handle a particular style in a specific situation. As a result, these methods require various settings to adapt to different scenarios. In this work, we overcome such limitations and propose an end-to-end imitation learning framework for drone cinematography systems. The framework consists of two main components: (1) an efficient motion feature extraction module for generating a compact motion feature space, (2) a path-analysis-based reinforcement learning (PABRL) algorithm for imitating multiple filming styles from demonstrations and incorporating aesthetical features for improved perspective shots. Our PABRL method is based on the actor–critic network, which regards multiple human motion variables, camera translations, and image composition as inputs and then outputs an aesthetical filming strategy related to the subject motion. In addition, we propose an attention mechanism and a long–short-term rewarding function to enhance the motion feature space and the integrity of the generated trajectory, respectively. Extensive experimental results in simulated and real outdoor environments demonstrate that compared with state-of-the-art methods, our method can achieve 69.8% higher performance in terms of trajectory planning accuracy while successfully incorporating aesthetical features into the captured videos. Yuanjie Dang, Chong Huang 0005, Peng Chen 0008, Ronghua Liang, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Multim. | 4 |
| 2023 | MUSE: Visual Analysis of Musical Semantic SequenceabstractVisualization has the capacity of converting auditory perceptions of music into visual perceptions, which consequently opens the door to music visualization (e.g., exploring group style transitions and analyzing performance details). Current research either focuses on low-level analysis without constructing and comparing music group characteristics, or concentrates on high-level group analysis without analyzing and exploring detailed information. To fill this gap, integrating the high-level group analysis and low-level details exploration of music, we design a musical semantic sequence visualization analytics prototype system (MUSE) that mainly combines a distribution view and a semantic detail view, assisting analysts in obtaining the group characteristics and detailed interpretation. In the MUSE, we decompose the music into note sequences for modeling and abstracting music into three progressively fine-grained pieces of information (i.e., genres, instruments and notes). The distribution view integrates a new density contour, which considers sequence distance and semantic similarity, and helps analysts quickly identify the distribution features of the music group. The semantic detail view displays the music note sequences and combines the window moving to avoid visual clutter while ensuring the presentation of complete semantic details. To prove the usefulness and effectiveness of MUSE, we perform two case studies based on real-world music MIDI data. In addition, we conduct a quantitative user study and an expert evaluation. Baofeng Chang, Guodao Sun, Houchao Huang, Ronghua Liang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | A Predictive Visual Analytics System for Studying Neurodegenerative Disease Based on DTI Fiber TractsabstractDiffusion tensor imaging (DTI) has been used to study the effects of neurodegenerative diseases on neural pathways, which may lead to more reliable and early diagnosis of these diseases as well as a better understanding of how they affect the brain. We introduce a predictive visual analytics system for studying patient groups based on their labeled DTI fiber tract data and corresponding statistics. The system's machine-learning-augmented interface guides the user through an organized and holistic analysis space, including the statistical feature space, the physical space, and the space of patients over different groups. We use a custom machine learning pipeline to help narrow down this large analysis space and then explore it pragmatically through a range of linked visualizations. We conduct several case studies using DTI and T1-weighted images from the research database of Parkinson's Progression Markers Initiative. Chaoqing Xu, Tyson Neuroth, Takanori Fujiwara, Ronghua Liang, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Enhanced Dual-Level Representations for Facial Expression RecognitionabstractFacial expression is an essential factor in conveying human emotional states and intentions. A common strategy used for facial expression recognition (FER) is encoding expression representations from facial images. Although remarkable advancement has been made, challenges due to large variations of expression patterns and unavoidable hard samples still remain. In this paper, we propose dual-level representation enhancements (DLRE) addressing these issues. On one hand, mid-level representation enhancement (MRE) is introduced to avoid expression representation learning being dominated by a limited number of highly discriminative patterns. On the other hand, high-level representation enhancement (HRE) is introduced to alleviate the disturbance of misclassified representations especially for hard samples. The proposed method not only has stronger generalization capability to handle different variations of expression patterns but also greater discriminative power to capture the subtle distinctions of hard samples. Experimental evaluation on four popular databases, CK+, Oulu-CASIA, RAF-DB, and AffectNet, shows that our method achieves more competitive results than other state-of-the-art methods. Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang |
ICIP | 8 |
| 2022 | CFN: A coarse-to-fine network for eye fixation predictionabstractAbstract Many image‐to‐image computer vision approaches have made great progress by an end‐to‐end framework with the encoder–decoder architecture. However, the same image‐to‐image eye fixation prediction task is not the same as those computer vision tasks in that it focuses more on salient regions rather than precise predictions for every pixel. Thus, it is not appropriate to directly apply the end‐to‐end encoder–decoder to the eye fixation prediction task. In addition, although high‐level feature is important, the contribution of low‐level feature should also be kept and balanced in computational model. Nevertheless, some low‐level features that attract attention are easily neglected while transiting through the deep network. Therefore, the effective way to integrate low‐level and high‐level features for improving eye fixation prediction performance is still a challenging task. In this paper, a coarse‐to‐fine network (CFN) that encompasses two pathways with different training strategies are proposed: coarse perceiving network (CFN‐Coarse) can be a simple encoder network or any of the existing pretrained network to capture the distribution of salient regions and generate high‐quality feature maps; fine integrating network (CFN‐Fine) uses fixed parameters from the CFN‐Coarse and combines features from deep to shallow in the deconvolution process by adding skip connections between down‐sampling and up‐sampling paths to efficiently integrate deep and shallow features. The saliency map obtained by the method is evaluated over 6 standard benchmark datasets, namely SALICON, MIT1003, MIT300, Toronto, OSIE, and SUN500. The results demonstrate that the method can surpass the state‐of‐the‐art accuracy of eye fixation prediction and achieves the competitive performance to date under most evaluation metrics on SALICON Saliency Prediction Challenge (LSUN2017). Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
IET Image Process. | 3 |
| 2022 | EvoSets: Tracking the Sensitivity of Dimensionality Reduction Results Across SubspacesabstractDimensionality reduction is commonly used for identifying and analyzing patterns in the visual analysis of multi-dimensional datasets. The selection of subspaces is a core building block in projecting high-dimensional data to low-dimensional space, which is usually illustrated as a scatterplot for analysts to easily understand and explore. This process involves human prior knowledge and domain-specific requirements. Thus, quantifying and tracking the changes of dimensionality reduction results across subspaces remain challenging. Existing methods can neither quantify the subsets-based changes of dimensionality reduction results when switching subspaces, nor automatically and comprehensively display the overall and subtle differences among dimensionality reduction results. To address this, we developedEvoSets, a novel visual analytics system designed to help users understand how subspaces affect dimensionality reduction results. The effects are quantified based on the distribution of subsets within projections to tracking the sensitivity of dimensionality reduction results across subspaces. In addition, the system supports the exploration of the overall evolution of the dimensionality reduction results for helping users track the convergence and divergence behavior changes of subsets based on an extendedBubble Setsvisualization. Similarities are intuitively illustrated, and dissimilarities are highlighted among the generated dimensionality reduction results across subspaces based on different layout constraints. The usefulness and effectiveness of the system are further evaluated with a user study and two case studies on multi-dimensional datasets. Guodao Sun, Sujia Zhu, Ronghua Liang |
IEEE Trans. Big Data | 5 |
| 2022 | 3-D Instance Segmentation of MVS BuildingsabstractWe present a novel 3D instance segmentation framework for Multi-View Stereo (MVS) buildings in urban scenes. Unlike existing works focusing on semantic segmentation of urban scenes, the emphasis of this work lies in detecting and segmenting 3D building instances even if they are attached and embedded in a large and imprecise 3D surface model. Multi-view RGB images are first enhanced to RGBH images by adding a heightmap and are segmented to obtain all roof instances using a fine-tuned 2D instance segmentation neural network. Instance masks from different multi-view images are then clustered into global masks. Our mask clustering accounts for spatial occlusion and overlapping, which can eliminate segmentation ambiguities among multi-view images. Based on these global masks, 3D roof instances are segmented out by mask back-projections and extended to the entire building instances through a Markov random field optimization. A new dataset that contains instance-level annotation for both 3D urban scenes (roofs and buildings) and drone images (roofs) is provided. To the best of our knowledge, it is the first outdoor dataset dedicated for 3D instance segmentation with much more annotations of attached 3D buildings than existing datasets1. Quantitative evaluations and ablation studies have shown the effectiveness of all major steps and the advantages of our multi-view framework over the orthophoto-based method. Jiazhou Chen 0002, Yanghui Xu, Shufang Lu, Ronghua Liang, Liangliang Nan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | AFExplorer: Visual analysis and interactive selection of audio featuresabstractAcoustic quality detection is vital in the manufactured products quality control field since it represents the conditions of machines or products. Recent work employed machine learning models in manufactured audio data to detect anomalous patterns. A major challenge is how to select applicable audio features to meliorate model’s accuracy and precision. To relax this challenge, we extract and analyze three audio feature types including Time Domain Feature, Frequency Domain Feature, and Cepstrum Feature to help identify the potential linear and non-linear relationships. In addition, we design a visual analysis system, namely AFExplorer, to assist data scientists in extracting audio features and selecting potential feature combinations. AFExplorer integrates four main views to present detailed distribution and relevance of the audio features, which helps users observe the impact of features visually in the feature selection. We perform the case study with AFExplore according to the ToyADMOS and MIMII Dataset to demonstrate the usability and effectiveness of the proposed system. Lei Wang 0023, Guodao Sun, Yunchao Wang, Ronghua Liang |
Vis. Informatics | 6 |
| 2022 | Visualization and visual analysis of multimedia data in manufacturing: A surveyabstractWith the development of production technology and social needs, sectors of manufacturing are constantly improving. The use of sensors and computers has made it increasingly convenient to collect multimedia data in manufacturing. Targeted, rapid, and detailed analysis based on the type of multimedia data can make timely decisions at different stages of the entire manufacturing process. Visualization and visual analytics are frequently adopted in multimedia data analysis of manufacturing because of their powerful ability to understand, present, and analyze data intuitively and interactively. In this paper, we present a literature review of visualization and visual analytics specifically for manufacturing multimedia data. We classify existing research according to visualization techniques, interaction analysis methods, and application areas. We discuss the differences when visualization and visual analytics are applied to different types of multimedia data in the context of particular examples of manufacturing research projects. Finally, we summarize the existing challenges and prospective research directions. Yunchao Wang, Lei Wang 0023, Guodao Sun, Ronghua Liang |
Vis. Informatics | 5 |
| 2022 | Towards a better understanding of the role of visualization in online learning: A reviewabstractWith the popularity of online learning in recent decades, MOOCs (Massive Open Online Courses) are increasingly pervasive and widely used in many areas. Visualizing online learning is particularly important because it helps to analyze learner performance, evaluate the effectiveness of online learning platforms, and predict dropout risks. Due to the large-scale, high-dimensional, and heterogeneous characteristics of the data obtained from online learning, it is difficult to find hidden information. In this paper, we review and classify the existing literature for online learning to better understand the role of visualization in online learning. Our taxonomy is based on four categorizations of online learning tasks: behavior analysis, behavior prediction, learning pattern exploration and assisted learning. Based on our review of relevant literature over the past decade, we also identify several remaining research challenges and future research work. Gefei Zhang 0002, Sujia Zhu, Ronghua Liang, Guodao Sun |
Vis. Informatics | 4 |
| 2021 | Locate Globally, Segment Locally: A Progressive Architecture With Knowledge Review Network for Salient Object DetectionabstractSalient object location and segmentation are two different tasks in salient object detection (SOD). The former aims to globally find the most attractive objects in an image, whereas the latter can be achieved only using local regions that contain salient objects. However, previous methods mainly accomplish the two tasks simultaneously in a simple end-to-end manner, which leads to the ignorance of the differences between them. We assume that the human vision system orderly locates and segments objects, so we propose a novel progressive architecture with knowledge review network (PA-KRN) for SOD. It consists of three parts. (1) A coarse locating module (CLM) that uses body-attention label locates rough areas containing salient objects without boundary details. (2) An attention-based sampler highlights salient object regions with high resolution based on body-attention maps. (3) A fine segmenting module (FSM) finely segments salient objects. The networks applied in CLM and FSM are mainly based on our proposed knowledge review network (KRN) that utilizes the finest feature maps to reintegrate all previous layers, which can make up for the important information that is continuously diluted in the top-down path. Experiments on five benchmarks demonstrate that our single KRN can outperform state-of-the-art methods. Furthermore, our PA-KRN performs better and substantially surpasses the aforementioned methods. Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
AAAI | 3 |
| 2021 | Facial Expression Recognition by Expression-Specific Representation Swapping
Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang |
ICANN (2) | 7 |
| 2021 | Flexible Knowledge Distillation with an Evolutional Network PopulationabstractDeep neural networks have continually surpassed traditional methods on a variety of computer vision tasks. Though deep neural networks are very powerful, the large number of parameters and complex structures consume considerable storage and calculation time, making it hard to deploy with limited resources. To tackle this issue, many recently proposed knowledge distillation approaches are aimed at obtaining a small student network to imitate a large teacher network. However, the student network structure is pre-defined and may be hard to train. In this paper, we propose to distill knowledge with an evolutional student network population. The population is initialized with several basic structures and each network is evaluated by the imitation ability (i.e., fitness) to the teacher network. By reusing the weights, we provide five enhancement options to strengthen the networks with high fitness and abandon the weak ones. By changing the fitness criterion, we can select networks to meet different requirements, such as balancing size and accuracy. This allows one to find a superior student network structure that better imitates the teacher model from various aspects with easier training. The experimental results demonstrate the proposed method can achieve superior performance of knowledge distillation with flexible student structures. Jie Lei 0002, Mingli Song, Jianping Shen, Ronghua Liang |
ICME | 6 |
| 2021 | Subcutaneous sweat pore estimation from optical coherence tomographyabstractAbstract Abstract Sweat pore, one of the level 3 features of fingerprint, has attracted much attention in fingerprint recognition. Traditional sweat pores on surface fingerprint are unclear or blurred when fingers are stained or damaged. Subcutaneous sweat pores, as cross section of the sweat glands, are resistant to external interferences. With 3D fingertip information measured by optical coherence tomography (OCT), the subcutaneous sweat pore estimation from OCT volume data is investigated. First, an adaptive subcutaneous pore image reconstruction method is proposed. It utilizes the skin surface and viable epidermis junction as reference and realizes depth‐adaptive pore image reconstruction. Second, a dilated U‐Net combining the U‐Net with dilated convolution is proposed for subcutaneous sweat pore extraction, which can prevent information loss of sweat pores caused by downsampling. To the best knowledge, it is the first time that subcutaneous sweat pore extraction is investigated and proposed. Experiments on subcutaneous pore image reconstruction and sweat pore extraction are both conducted. The qualitative and quantitative results show that the proposed adaptive method performs better in subcutaneous pore image reconstruction compared with the fix‐depth method, and the dilated U‐Net outperforms other methods on subcutaneous sweat pore extraction. Baojin Ding, Haixia Wang 0002, Peng Chen 0008, Yilong Zhang 0001, Ronghua Liang, Yipeng Liu 0002 |
IET Image Process. | 5 |
| 2021 | Blood vessel and background separation for retinal image quality assessmentabstractAbstract Retinal image analysis has become an intuitive and standard aided diagnostic technique for eye diseases. The good image quality is essential support for doctors to provide timely and accurate disease diagnosis. This paper proposes an end‐to‐end learning based method for evaluating the retinal image quality. First, blood vessels of the input image are segmented by U‐Net, and the fundus image is divided into two parts: blood vessels and background. Then, we design a dual branch network module which extracts global features that influence the image quality and suppress the interference of blood vessels and local textures to achieve better performance. The proposed module can be embedded in various advanced network structures. The experimental results show the more efficient convergence rate for the network with the module. The best network accuracy rate is 85.83%, the AUC is 0.9296, and the F1‐score is 0.7967 on the collected local dataset. Additionally, the model generalization is tested on the public DRIMDB dataset. The accuracy, AUC, and F1‐score reach 97.89%, 0.9978, and 0.9688, respectively. Compared with the state‐of‐the‐art networks, the performance of the proposed method is proven to be accurate and effective for retinal image quality assessment. Yipeng Liu 0002, Yajun Lv, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IET Image Process. | 7 |
| 2021 | Feature pyramid U-Net for retinal vessel segmentationabstractAbstract The retinal vessel is the only microvascular network that can be directly and non‐invasively observed in humans. Cardiovascular and cerebrovascular diseases, such as diabetes, hypertension, can lead to structural changes of the retinal microvascular network. Therefore, it is of great significance to study effective retinal vessel segmentation methods and assist doctors in early diagnoses with quantitative results for vascular networks. In this study, we propose a novel convolutional neural network named feature pyramid U‐Net (FPU‐Net) that extracts multiscale representations by constructing two feature pyramids both on the encoder and the decoder of U‐Net. In this representation, objects features with different size like micro‐vessels and pathology will be fused for better vessel segmentation. The experimental results show that compared with state‐of‐the‐art methods, FPU‐Net is superior in terms of accuracy, sensitivity, F1‐score, and area under the curve and capable of stronger domain generalisation across different datasets. Yipeng Liu 0002, Xue Rui, Zhanqing Li, Dongxu Zeng, Peng Chen 0008, Ronghua Liang |
IET Image Process. | 7 |
| 2021 | Multiscale ensemble of convolutional neural networks for skin lesion classificationabstractAbstract Early detection and treatment of skin cancer can considerably reduce the patient mortality rates. Convolutional neural network (CNN) has been widely applied in the field of computer aided diagnosis. However, for skin lesions, the inconsistent size of lesion regions in dermatoscope images hinders the convolutional neural network precise discrimination. To solve this problem, multiscale ensemble of convolutional neural networks called MECNN is proposed, which involves three branches with different lesion scales as the model input. The first branch locates the lesion region outline by identifying the largest local response point. Then, MECNN reduces the search area of the lesion region and divides the outline into two scales used as the input for the other two branches. A global loss function is defined to control the learning objectives of the three branches and MECNN fuses the branches output as the final classification result. The proposed model is evaluated on the public HAM10000 dataset and achieves a higher classification accuracy than the comparative state‐of‐the‐art methods. Yipeng Liu 0002, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IET Image Process. | 7 |
| 2021 | Video multimodal emotion recognition based on Bi-GRU and attention fusion
Ruohong Huan, Jia Shu, Shenglin Bao, Ronghua Liang, Peng Chen 0008, Kaikai Chi |
Multim. Tools Appl. | 4 |
| 2021 | A hybrid CNN and BLSTM network for human complex activity recognition with multi-feature fusion
Ruohong Huan, Ziwei Zhan, Luoqi Ge, Kaikai Chi, Peng Chen 0008, Ronghua Liang |
Multim. Tools Appl. | 6 |
| 2021 | Scene Categorization by Deeply Learning Gaze Behavior in a Semisupervised ContextabstractAccurately recognizing different categories of sceneries with sophisticated spatial configurations is a useful technique in computer vision and intelligent systems, e.g., scene understanding and autonomous driving. Competitive accuracies have been observed by the deep recognition models recently. Nevertheless, these deep architectures cannot explicitly characterize human visual perception, that is, the sequence of gaze allocation and the subsequent cognitive processes when viewing each scenery. In this paper, a novel spatially aware aggregation network is proposed for scene categorization, where the human gaze behavior is discovered in a semisupervised setting. In particular, as semantically labeling a large quantity of scene images is labor-intensive, a semisupervised and structure-preserved non-negative matrix factorization (NMF) is proposed to detect a set of visually/semantically salient regions from each scenery. Afterward, the gaze shifting path (GSP) is engineered to characterize the process of humans perceiving each scene picture. To deeply describe each GSP, a novel spatially aware CNN termed SA-Net is developed. It accepts input regions with various shapes and statistically aggregates all the salient regions along each GSP. Finally, the learned deep GSP features from the entire scene images are fused into an image kernel, which is subsequently integrated into a kernel SVM to categorize different sceneries. Comparative experiments on six scene image sets have shown the advantage of our method. Ronghua Liang, Jianwei Yin, Dongxiang Zhang, Ling Shao 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Surface and Internal Fingerprint Reconstruction From Optical Coherence Tomography Through Convolutional Neural NetworkabstractOptical coherence tomography (OCT), as a non-destructive and high-resolution fingerprint acquisition technology, is robust against poor skin conditions and resistant to spoof attacks. It measures fingertip information on and beneath skin as 3D volume data, containing the surface fingerprint, internal fingerprint and sweat glands. Various methods have been proposed to extract internal fingerprints, which ignore the inter-slice dependence and often require manually selected parameters. In this article, a modified U-Net that combines residual learning, bidirectional convolutional long short-term memory and hybrid dilated convolution (denoted as BCL-U Net) for OCT volume data segmentation and two fingerprint reconstruction approaches are proposed. To the best of our knowledge, it is the first time that simultaneous and automatic extraction is performed for surface fingerprint, internal fingerprint and sweat gland. The proposed BCL-U Net utilizes the spatial dependence in OCT volume data and deals with segmentation of objects with diverse sizes to achieve accurate extraction. Comparisons have been performed to demonstrate the advantages of the proposed method. A thorough evaluation of the recognition abilities of internal and surface fingerprints is conducted using a dataset significantly larger than previous studies. Four databases containing internal and surface fingerprints are generated from 1572 OCT volume data by the proposed method. The internal fingerprint matching experiment has achieved a lowest equal error rate (EER) of 0.95%. Mixed internal and surface fingerprint matching experiment is also performed and achieves an EER of 3.67%, verifying the consistency of the internal and surface fingerprints. The matching experiments for fingers under poor skin conditions show a 2.47% EER of internal fingerprints that is much lower than that of surface fingerprints, which proves the advantage of internal fingerprints and indicates the potential of the internal fingerprints to supplement or replace the surface fingerprints for some specific applications. Baojin Ding, Haixia Wang 0002, Peng Chen 0008, Yilong Zhang 0001, Zhenhua Guo 0001, Jianjiang Feng, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2021 | VSumVis: Interactive Visual Understanding and Diagnosis of Video Summarization ModelabstractWith the rapid development of mobile Internet, the popularity of video capture devices has brought a surge in multimedia video resources. Utilizing machine learning methods combined with well-designed features, we could automatically obtain video summarization to relax video resource consumption and retrieval issues. However, there always exists a gap between the summarization obtained by the model and the ones annotated by users. How to help users understand the difference, provide insights in improving the model, and enhance the trust in the model remains challenging in the current study. To address these challenges, we propose VSumVis under a user-centered design methodology, a visual analysis system with multi-feature examination and multi-level exploration, which could help users explore and analyze video content, as well as the intrinsic relationship that existed in our video summarization model. The system contains multiple coordinated views, i.e., video view, projection view, detail view, and sequential frames view. A multi-level analysis process to integrate video events and frames are presented with clusters and nodes visualization in our system. Temporal patterns concerning the difference between the manual annotation score and the saliency score produced by our model are further investigated and distinguished with sequential frames view. Moreover, we propose a set of rich user interactions that enable an in-depth, multi-faceted analysis of the features in our video summarization model. We conduct case studies and interviews with domain experts to provide anecdotal evidence about the effectiveness of our approach. Quantitative feedback from a user study confirms the usefulness of our visual system for exploring the video summarization model. Guodao Sun, Chaoqing Xu, Haoran Liang 0001, Binwei Xu, Ronghua Liang |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2021 | A survey of volume visualization techniques for feature enhancementabstractVolume rendering techniques have been developed for decades and have been widely applied in many research fields, such as medical image visualization, geological exploration, scientific computing. etc. With the maturity of volume visualization techniques, one may have many choices to analyze volume data. However, facing different application requirements in specific cases, one may need pertinent methods to visualize volume data and highlight specific volume features. In this paper, we review and classify the existing literature on feature enhancement volume rendering. The classification is conducted based on the enhancement of four types of features (the external feature, the internal feature, the structure feature, and the ideographic feature) in volume data. Finally, we conclude this survey with future challenges in feature enhancement volume visualization. Chaoqing Xu, Guodao Sun, Ronghua Liang |
Vis. Informatics | 3 |
| 2020 | Video summarisation with visual and semantic cuesabstractVideo summarisation greatly improves the efficiency of people browsing videos and saves storage space. A good video summary should satisfy human visual interestingness and preserve the theme of the original video at the semantic level. Unlike many existing methods that consider only visual features to generate video summaries, this study proposes a method that combines visual and semantic cues to extract important information for dynamic video summarisation. The authors propose visual‐verbal saliency consistency to add semantic information and propose a novel attention motion, along with other visual features to fully represent visual interestingness. Based on the importance score of each frame calculated by combining these features, they select an optimal subset of segments to generate an important and interesting summary. They evaluate their method using the SumMe and TVSum datasets and experimental results show that their method generates high‐quality video summaries. Binwei Xu, Haoran Liang 0001, Ronghua Liang |
IET Image Process. | 3 |
| 2020 | A structure-guided approach to the prediction of natural image saliency
Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
Neurocomputing | 3 |
| 2020 | A survey on automatic infographics and visualization recommendationsabstractAutomatic infographics generators employ machine learning algorithms/user-defined rules and visual embellishments into the creation of infographics. It is an emerging topic in the field of information visualization that has requirements in many sectors, such as dashboard design, data analysis, and visualization recommendation. The growing popularity of visual analytics in recent years brings increased attention to automatic infographics. This creates the need for a broad survey that reviews and assesses the significant advances in this field. Automatic tools aim to lower the barrier for visually analyzing data by automatically generating visualizations for analysts to search and make a choice, instead of manually specifying. This survey reviews and classifies automatic tools and papers of visualization recommendations into a set of application categories including network-graph visualizations, annotation visualizations, and storytelling visualization. More importantly, this report presents several challenges and promising directions for future work in the field of automatic infographics and visualization recommendations. Sujia Zhu, Guodao Sun, Meng Zha, Ronghua Liang |
Vis. Informatics | 5 |
| 2019 | Referable diabetic retinopathy identification from eye fundus images with weighted path for convolutional neural network
Yipeng Liu 0002, Zhanqing Li, Ronghua Liang |
Artif. Intell. Medicine | 5 |
| 2019 | CapVis: Toward Better Understanding of Visual-Verbal Saliency ConsistencyabstractWhen looking at an image, humans shift their attention toward interesting regions, making sequences of eye fixations. When describing an image, they also come up with simple sentences that highlight the key elements in the scene. What is the correlation between where people look and what they describe in an image? To investigate this problem intuitively, we develop a visual analytics system, CapVis, to look into visual attention and image captioning, two types of subjective annotations that are relatively task-free and natural. Using these annotations, we propose a word-weighting scheme to extract visual and verbal saliency ranks to compare against each other. In our approach, a number of low-level and semantic-level features relevant to visual-verbal saliency consistency are proposed and visualized for a better understanding of image content. Our method also shows the different ways that a human and a computational model look at and describe images, which provides reliable information for a captioning model. Experiment also shows that the visualized feature can be integrated into a computational model to effectively predict the consistency between the two modalities on an image dataset with both types of annotations. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2018 | SocialWave: Visual Analysis of Spatio-temporal Diffusion of Information on Social MediaabstractRapid advancement of social media tremendously facilitates and accelerates the information diffusion among users around the world. How and to what extent will the information on social media achieve widespread diffusion across the world? How can we quantify the interaction between users from different geolocations in the diffusion process? How will the spatial patterns of information diffusion change over time? To address these questions, a dynamic social gravity model (SGM) is proposed to quantify the dynamic spatial interaction behavior among social media users in information diffusion. The dynamic SGM includes three factors that are theoretically significant to the spatial diffusion of information: geographic distance, cultural proximity, and linguistic similarity. Temporal dimension is also taken into account to help detect recency effect, and ground-truth data is integrated into the model to help measure the diffusion power. Furthermore, SocialWave, a visual analytic system, is developed to support both spatial and temporal investigative tasks. SocialWave provides a temporal visualization that allows users to quickly identify the overall temporal diffusion patterns, which reflect the spatial characteristics of the diffusion network. When a meaningful temporal pattern is identified, SocialWave utilizes a new occlusion-free spatial visualization, which integrates a node-link diagram into a circular cartogram for further analysis. Moreover, we propose a set of rich user interactions that enable in-depth, multi-faceted analysis of the diffusion on social media. The effectiveness and efficiency of the mathematical model and visualization system are evaluated with two datasets on social media, namely, Ebola Epidemics and Ferguson Unrest. Guodao Sun, Tan Tang, Tai-Quan Peng, Ronghua Liang, Yingcai Wu |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2018 | Real-Time Object Tracking on a Drone With Multi-Inertial Sensing DataabstractReal-time object tracking on a drone under a dynamic environment has been a challenging issue for many years, with existing approaches using off-line calculation or powerful computation units on board. This paper presents a new lightweight real-time onboard object tracking approach with multi-inertial sensing data, wherein a highly energy-efficient drone is built based on the Snapdragon flight board of Qualcomm. The flight board uses a digital signal processor core of the Snapdragon 801 processor to realize PX4 autopilot, an open-source autopilot system oriented toward inexpensive autonomous aircraft. It also uses an ARM core to realize Linux, robot operating systems, open-source computer vision library, and related algorithms. A lightweight moving object detection algorithm is proposed that extracts feature points in the video frame using the oriented FAST and rotated binary robust independent elementary features algorithm and adapts a local difference binary algorithm to construct the image binary descriptors. The K-nearest neighbor method is then used to match the image descriptors. Finally, an object tracking method is proposed that fuses inertial measurement unit data, global positioning system data, and the moving object detection results to calculate the relative position between coordinate systems of the object and the drone. All the algorithms are run on the Qualcomm platform in real time. Experimental results demonstrate the superior performance of our method over the state-of-the-art visual tracking method. Peng Chen 0008, Yuanjie Dang, Ronghua Liang, Wei Zhu 0006, Xiaofei He 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Saliency prediction with scene structural guidanceabstractPrevious works have suggested the role of scene information in directing gaze. The structure of a scene provides global contextual information that complements local object information in saliency prediction. In this study, we explore how scene envelopes such as openness, depth, and perspective affect visual attention in natural outdoor images. To facilitate this study, an eye tracking dataset is first built with 500 natural scene images and eye tracking data with 15 subjects free-viewing the images. We make observations on scene layout properties and propose a set of scene structural features relating to visual attention. We further integrate features from deep neural networks and use the set of complementary features for saliency prediction. Our features are independent of and can work together with many computational modules, and this work demonstrates the use of Multiple kernel learning (MKL) as an example to integrate the features at low- and high-levels. Experimental results demonstrate that our model outperforms existing methods and our scene structural features can improve the performance of other saliency models in outdoor scenes. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
SMC | 3 |
| 2017 | Visual-verbal consistency of image saliencyabstractWhen looking at an image, humans shift their attention towards interesting regions, making sequences of eye fixations. When describing an image, they also come up with simple sentences that highlight the key elements in the scene. What is the correlation between where people look and what they describe in an image? To investigate this problem, we look into eye fixations and image captions, two types of subjective annotations that are relatively task-free and natural. From the annotations, we extract visual and verbal saliency ranks to compare against each other. We then propose a number of low-level and semantic-level features relevant to the visualverbal consistency. Integrated into a computational model, the proposed features effectively predict the consistency between the two modalities on a large dataset with both types of annotations, namely SALICON [1]. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
SMC | 3 |
| 2017 | Exemplar-based image inpainting using structure consistent patch matching
Haixia Wang 0002, Ronghua Liang |
Neurocomputing | 3 |
| 2017 | Sparse Learning with Stochastic Composite OptimizationabstractIn this paper, we study Stochastic Composite Optimization (SCO) for sparse learning that aims to learn a sparse solution from a composite function. Most of the recent SCO algorithms have already reached the optimal expected convergence rate O(1/λT), but they often fail to deliver sparse solutions at the end either due to the limited sparsity regularization during stochastic optimization (SO) or due to the limitation in online-to-batch conversion. Even when the objective function is strongly convex, their high probability bounds can only attain O(√{log(1/δ)/T}) with δ is the failure probability, which is much worse than the expected convergence rate. To address these limitations, we propose a simple yet effective two-phase Stochastic Composite Optimization scheme by adding a novel powerful sparse online-to-batch conversion to the general Stochastic Optimization algorithms. We further develop three concrete algorithms, OptimalSL, LastSL and AverageSL, directly under our scheme to prove the effectiveness of the proposed scheme. Both the theoretical analysis and the experiment results show that our methods can really outperform the existing methods at the ability of sparse learning and at the meantime we can improve the high probability bound to approximately O(log(log(T)/δ)/λT). Lijun Zhang 0005, Zhongming Jin 0001, Rong Jin 0001, Deng Cai 0001, Xuelong Li 0001, Ronghua Liang, Xiaofei He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2017 | Embedding Spatio-Temporal Information into Maps by Route-ZoomingabstractAnalysis and exploration of spatio-temporal data such as traffic flow and vehicle trajectories have become important in urban planning and management. In this paper, we present a novel visualization technique called route-zooming that can embed spatio-temporal information into a map seamlessly for occlusion-free visualization of both spatial and temporal data. The proposed technique can broaden a selected route in a map by deforming the overall road network. We formulate the problem of route-zooming as a nonlinear least squares optimization problem by defining an energy function that ensures the route is broadened successfully on demand while the distortion caused to the road network is minimized. The spatio-temporal information can then be embedded into the route to reveal both spatial and temporal patterns without occluding the spatial context information. The route-zooming technique is applied in two instantiations including an interactive metro map for city tourism and illustrative maps to highlight information on the broadened roads to prove its applicability. We demonstrate the usability of our spatio-temporal visualization approach with case studies on real traffic flow data. We also study various design choices in our method, including the encoding of the time direction and choices of temporal display, and conduct a comprehensive user study to validate our embedded visualization design. Guodao Sun, Ronghua Liang, Huamin Qu, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | Scientific Ranking over Heterogeneous Academic HypernetworkabstractRanking is an important way of retrieving authoritative papers from a large scientific literature database. Current state-of-the-art exploits the flat structure of the heterogeneous academic network to achieve a better ranking of scientific articles, however, ignores the multinomial nature of the multidimensional relationships between different types of academic entities. This paper proposes a novel mutual ranking algorithm based on the multinomial heterogeneous academic hypernetwork, which serves as a generalized model of a scientific literature database. The proposed algorithm is demonstrated effective through extensive evaluation against well-known IR metrics on a well-established benchmarking environment based on the ACL Anthology Network. Ronghua Liang, Xiaorui Jiang |
AAAI | 1 |
| 2016 | TravelDiff: Visual comparison analytics for massive movement patterns derived from TwitterabstractGeo-tagged microblog data covers billions of movement patterns on a global and local scale. Understanding these patterns could guide urban and traffic planning or help coping with disaster situations. We present a visual analytics system to investigate travel trajectories of people reconstructed from microblog messages. To analyze seasonal changes and events and to validate movement patterns against other data sources, we contribute highly interactive visual comparison methods that normalize and contrast trajectories as well as density maps within a single view. We also compute an adaptive hierarchical graph from the trajectories to abstract individual movements into higher-level structures. Specific challenges that we tackle are, among others, the spatio-temporal sparsity of the data, the volume of data varying by region, and a diverse mix of means of transportation. The applicability of our approach is presented in three case studies. Robert Krüger, Guodao Sun, Fabian Beck 0001, Ronghua Liang, Thomas Ertl |
PacificVis | 4 |
| 2016 | Visual exploration of HARDI fibers with probabilistic tracking
Ronghua Liang, Zhengzhou Wang, Song Zhang 0004, Yuanjing Feng, Xiangyin Ma, Wei Chen 0001, David F. Tate |
Inf. Sci. | 1 |
| 2016 | Coupled Dictionary Learning for the Detail-Enhanced Synthesis of 3-D Facial ExpressionsabstractThe desire to reconstruct 3-D face models with expressions from 2-D face images fosters increasing interest in addressing the problem of face modeling. This task is important and challenging in the field of computer animation. Facial contours and wrinkles are essential to generate a face with a certain expression; however, these details are generally ignored or are not seriously considered in previous studies on face model reconstruction. Thus, we employ coupled radius basis function networks to derive an intermediate 3-D face model from a single 2-D face image. To optimize the 3-D face model further through landmarks, a coupled dictionary that is related to 3-D face models and their corresponding 3-D landmarks is learned from the given training set through local coordinate coding. Another coupled dictionary is then constructed to bridge the 2-D and 3-D landmarks for the transfer of vertices on the face model. As a result, the final 3-D face can be generated with the appropriate expression. In the testing phase, the 2-D input faces are converted into 3-D models that display different expressions. Experimental results indicate that the proposed approach to facial expression synthesis can obtain model details more effectively than previous methods can. Haoran Liang 0001, Ronghua Liang, Mingli Song, Xiaofei He 0001 |
IEEE Trans. Cybern. | 2 |
| 2016 | Looking Into Saliency Model via Space-Time VisualizationabstractWe introduce a visual analytics method to analyze eye-tracking data and saliency models for dynamic stimuli, such as video or animated graphics. The focus lies on the analysis of the different performance of saliency models in contrast to human observers to identify trends in the general viewing behavior, including time sequences of attentional synchrony and objects with a strong attentional focus. By using a space-time cube visualization in combination with clustering, the dynamic stimuli and associated eye gazes as well as the attention maps from saliency models can be analyzed in a static three-dimensional representation. We propose algorithms to keep the appearance of the computer's attention data in line with the human's eye-tracking data. The analytical process is supported by multiple coordinated views that allow the user to focus on different aspects of spatial and temporal information in eye gaze data and saliency map. By comparing attention data from both human and computer incorporated with the spatiotemporal characteristics, we are able to find the different patterns within human and computer algorithms. We list our key findings to help developing better saliency detection algorithms. Haoran Liang 0001, Ronghua Liang, Guodao Sun |
IEEE Trans. Multim. | 2 |
| 2015 | Mixed Error Coding for Face Recognition with Mixed Occlusions
Ronghua Liang, Xiaoxin Li 0001 |
IJCAI | 1 |
| 2015 | Bayesian multi-distribution-based discriminative feature extraction for 3D face recognition
Ronghua Liang, Wenjia Shen, Haixia Wang 0002 |
Inf. Sci. | 1 |
| 2015 | Motion recognition and synthesis based on 3D sparse representation
Ronghua Liang |
Signal Process. | 2 |
| 2015 | Uncertainty-Aware Multidimensional Ensemble Data Visualization and ExplorationabstractThis paper presents an efficient visualization and exploration approach for modeling and characterizing the relationships and uncertainties in the context of a multidimensional ensemble dataset. Its core is a novel dissimilarity-preserving projection technique that characterizes not only the relationships among the mean values of the ensemble data objects but also the relationships among the distributions of ensemble members. This uncertainty-aware projection scheme leads to an improved understanding of the intrinsic structure in an ensemble dataset. The analysis of the ensemble dataset is further augmented by a suite of visual encoding and exploration tools. Experimental results on both artificial and real-world datasets demonstrate the effectiveness of our approach. Haidong Chen, Song Zhang 0004, Wei Chen 0001, Honghui Mei, Andrew Mercer 0001, Ronghua Liang, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2014 | Embedding Temporal Display into Maps for Occlusion-Free Visualization of Spatio-temporal DataabstractIt is often necessary to analyze spatio-temporal data such as traffic flow, air pollution, and vehicle trajectories in a city. A map is often used to show the spatial context while various temporal displays like time series plots can be used to present the changes in the data over time. In this paper, we present a novel visualization that can seamlessly embed temporal displays into a map for occlusion-free visualization of both the spatial and temporal attributes of the data. We first extend the seam carving algorithm to broaden the roads of interest in a map with the least distortion to other areas, and then embed temporal displays into the roads to reveal temporal patterns without the occlusion of map information. We study various design choices in our method, including the encoding of the time direction and temporal display, and conduct two comprehensive user studies to validate our design decisions. We also demonstrate the usability of our approach with case studies on real traffic flow data in a major city. Guodao Sun, Ronghua Liang, Huamin Qu |
PacificVis | 4 |
| 2014 | Counting crowd flow based on feature points
Ronghua Liang, Yuge Zhu, Haixia Wang 0002 |
Neurocomputing | 1 |
| 2014 | Research on the conjugate gradient algorithm with a modified incomplete Cholesky preconditioner on GPU
Jiaquan Gao, Ronghua Liang, Jun Wang 0077 |
J. Parallel Distributed Comput. | 2 |
| 2014 | Oriented boundary padding for iterative and oriented fringe pattern denoising techniques
Haixia Wang 0002, Kemao Qian, Ronghua Liang, Huayin Wang, Xiaofei He 0001 |
Signal Process. | 3 |
| 2014 | Scaling Hop-Based Reachability Indexing for Fast Graph Pattern Query ProcessingabstractGraphs are becoming increasingly dominant in modeling real-life networked data including social and biological networks, the WWW and the Semantic Web, etc. Graph pattern queries are useful for gathering information with expressive semantics from these graph-structured data. Current methods for graph pattern query processing have performance deficiency caused by inefficiencies of the underlying reachability index and costly merge-join operations on huge amounts of tuple-formatted intermediate results. To overcome the above problems, this paper contributes in the following aspects to boost graph pattern query evaluation. First, we propose an improved hop-based reachability indexing scheme 3-Hop which gains faster reachability query evaluation, less indexing costs and better scalabilities than state-of-the-art hop-based methods. Second, we propose a two-stage node filtering algorithm based on 3-Hop to answer tree pattern queries more efficiently. Tree pattern queries serve as the underlying facility for graph pattern query evaluation. Furthermore, we use a graph representation of the intermediate results during node filtering and final results enumeration. Experiments on real-life and synthetic datasets demonstrate the effectiveness of the proposed methods. Ronghua Liang, Hai Zhuge, Xiaorui Jiang, Qiang Zeng 0002, Xiaofei He 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | EvoRiver: Visual Analysis of Topic Coopetition on Social MediaabstractCooperation and competition (jointly called "coopetition") are two modes of interactions among a set of concurrent topics on social media. How do topics cooperate or compete with each other to gain public attention? Which topics tend to cooperate or compete with one another? Who plays the key role in coopetition-related interactions? We answer these intricate questions by proposing a visual analytics system that facilitates the in-depth analysis of topic coopetition on social media. We model the complex interactions among topics as a combination of carry-over, coopetition recruitment, and coopetition distraction effects. This model provides a close functional approximation of the coopetition process by depicting how different groups of influential users (i.e., "topic leaders") affect coopetition. We also design EvoRiver, a time-based visualization, that allows users to explore coopetition-related interactions and to detect dynamically evolving patterns, as well as their major causes. We test our model and demonstrate the usefulness of our system based on two Twitter data sets (social topics data and business topics data). Guodao Sun, Yingcai Wu, Shixia Liu, Tai-Quan Peng, Jonathan J. H. Zhu, Ronghua Liang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2013 | A Web-based visual analytics system for real estate data
Guodao Sun, Ronghua Liang, Fuli Wu, Huamin Qu |
Sci. China Inf. Sci. | 2 |
| 2013 | Topic hypergraph: hierarchical visualization of thematic structures in long documents
Guizhen Wang, Chao-Kai Wen, Binghui Yan, Ronghua Liang, Wei Chen 0001 |
Sci. China Inf. Sci. | 5 |
| 2013 | Moplex orderings generated by the LexDFS algorithm
Shoujun Xu, Xianyue Li, Ronghua Liang |
Discret. Appl. Math. | 3 |
| 2013 | A Survey of Visual Analytics Techniques and Applications: State-of-the-Art Research and Future Challenges
Guodao Sun, Yingcai Wu, Ronghua Liang, Shixia Liu |
J. Comput. Sci. Technol. | 3 |
| 2012 | Accumulation of local maximum intensity for feature enhanced volume rendering
Ronghua Liang, Feng Dong 0005, Gordon Clapworthy |
Vis. Comput. | 1 |
| 2010 | A Distributed Workflow Modeling Method Based on Users' DemandabstractConstructing Web service workflow faces huge challenges in the volatile, heterogeneous, distributed environment: it is necessary to consider the dynamic changes in web services, but also take into account the rapid method of modeling workflow. Comparing workflow modeling and artificial intelligence planning process, if the Web service as a planned action (or activity), then the modern AI planning and workflow modeling to integrate, so we can use AI planning technology to solve the distributed workflow modeling. This paper presents a distributed service workflow model based on user demand (DSWMoUD), which is composed of the web services organizations model, the business concept model, business logic model, user demand model, business scheduling model and business enactment model. In order to improve the retrieval efficiency of distributed service, our proposed distributed service organizations model made up of web service registration system(WSRS) and web service spanning tree (WSST), and we give a building algorithm for WSST and a business logic spanning graph algorithm. Introducing the artificial intelligence planning techniques into the distributed workflow modeling, we implement the prototype system, which is Self-Adaptive Web Contractual Computing Management System (SAWCM). Property analysis of the new model is also made, which leads to the conclusion that the DSWMoUD may be applied as an optimized approach towards efficient and effective distributed service workflow modeling. Mingyuan Yu, Ronghua Liang |
APSCC | 4 |
| 2009 | Craniofacial model reconstruction from skull data based on feature pointsabstract3D craniofacial model reconstruction from skull data has been a challenging research topic for many years. In this paper we develop a craniofacial reconstruction system based on feature points, and our approach only requires the skull model from a standard 3D scanner and feature points superimposed on the skull data. Our method consists of three stages: hole repairing for the 3D skull model, facial reconstruction with Radial Basis Function(RBF) deformation and texture mapping. A new advancing layer-wise solution algorithm and a template matching method are proposed for repairing big and particular holes on the skull model. The reference face model is then transformed to reconstruct the face model base on marked feature points on the skull model, and we improve the RBF deformation.. The experiments demonstrate the high efficiency and visual realism achieved in our approach. Ronghua Liang, Yaolei Lin, Jiarui Bao, Xianping Huang |
CAD/Graphics | 1 |
| 2005 | Face animation with eigenface detection and FAP trackerabstractAn approach for face animation based on FDP definition in MPEG-4 is proposed by integrating eigenface detection and facial features tracker. Our method only requires simple device - a digital camera and a PC, and our system needs minimal user interaction. First, the algorithm makes full use of the eigenface to detect the position and the size of face the first frame, and facial features of the first frame are acquired automatically by Plessey corner detector according to FDP in MPEG-4 based on anatomical knowledge and general 3D model. Then feature motion data is obtained from maximizing the cross-correlation and Kalman filter, and 2D facial features motion data is converted to 3D FAPs in MPEG-4 by image normalization and features alignment. Finally face animation is obtained by deformation parameters in FAPs. Experimental results show that the approaches we present can be realized, and the results also show the high prospect of our method. Ronghua Liang, QingHong Zhou |
SMC | 2 |
| 2004 | New Algorithm for 3D Facial Model Reconstruction and Its Application in Virtual Reality
Ronghua Liang, Chun Chen 0001 |
J. Comput. Sci. Technol. | 1 |
| 2003 | Human Expressions Interaction Between Avatar and Virtual World
Ronghua Liang, Chun Chen 0001, Jiajun Bu |
ICCSA (1) | 1 |
| 2003 | Robust Real-Time Face Tracking and Modeling from Video
Ronghua Liang, Chun Chen 0001, Jiajun Bu |
ICCSA (1) | 1 |
| 2003 | Intelligent Crowd Simulation
Feng Liu 0015, Ronghua Liang |
ICCSA (1) | 2 |
| 2003 | 3D facial animation from Chinese textabstractThe reality and controllability of the facial animation are enhanced in our work by modeling mouth in a parametric approach described by 7 parameters: superior and anterior bend for upper and lower lip, width of the mouth, weight of radium of lips. These parameters were calculated by tracking only 4 points around the mouth from video, and the cost of computation is low. The coarticulation model of Cohen and Massaro was adopted to generate the natural key frames and the transitional frames. Considering the difference between Chinese and English, an algorithm was introduced to calculate the coefficients of dominance functions. Based on the muscle-based facial model and the parametric mouth model, the 3D head can talk vividly. Jiajun Bu, Chun Chen 0001, Ronghua Liang |
SMC | 4 |
| 2003 | Individual face expressions and expression cloningabstractWe develop a system to generate individual facial expression with minimal interactions, and expression cloning can be created for different 3D mesh model. In our system, expressions can be created with the muscle model, and muscular vectors adhere to the vertex in 3D mesh face model. Expressions can be generated after deforming the muscular vectors and rotation angle of jaw parameter and rotation angle of eyes. New approach for expression cloning for different 3D models is presented, radical based function (RBF) is employed to match the vertices of 3D source model and destination model, the cloning expression is obtained by deforming the proportioned muscular vector model. Cloned expression animations preserve the relative motions, dynamics, and character of the original facial animations. If manual tuning or computational costs are high in creating animations for one model, creating similar animations for new models will take similar efforts. Experimental results show the high vividness of our system. Ronghua Liang, Jiajun Bu, Chun Chen 0001 |
SMC | 1 |
| 2003 | Real-time facial features tracker with motion estimation and feedbackabstractReal-time automatic face tracking is a great challenge in computer vision and computer graphics. We develop a system to automatically track the face by integrating auto-generation of features of first frame, feature correspondence and Kalman filter based on face attribution and motion estimation. First, facial features of the first frame are acquired automatically by face detachment and Plessey corner detector according to FDP in MPEG-4 based on anatomical knowledge and general 3D model. Feature correspondence is obtained by maximizing the cross-correlation and Kalman filter. Automatic face detection makes full use of the knowledge of 3D general model and face attribute, and face tracking is more efficient by using motion estimation. Experimental results show the high prospect of this algorithm. Ronghua Liang, Chun Chen 0001, Jiajun Bu |
SMC | 1 |
| 2003 | 3D realistic talking face co-driven by text and speechabstractTo create 3D realistic talking face has been a challenge for a long time. Previous works emphasize text or speech driven talking face respectively while the animation result is not very realistic or natural-looking. In the proposed approach, text and speech are considered to drive the 3D talkingface coordinately. The text is translated into a sequence of visemes' transcription. And time vector of the sequence is extracted from the speech corresponding to the text after it is segmented into phonetic sequence. A muscle based viseme vector is defined for static viseme. And then, with the time vector and the static visemes's sequence, dynamic visemes are generated through time-related dominance function. Finally, according to the frame rate to be rendered, intermediate frames are interpolated between key frames to make the animation result looks more natural and realistic than those obtained based on the text or speech-driven only. Mingli Song, Chun Chen 0001, Jiajun Bu, Ronghua Liang |
SMC | 4 |