VLDB 2026 Research / reviewers in the wild / expert
Ruomei Wang 0001
dblp:23/6165-1
· DBLP profile ↗
71ranked-venue papers
5as first author
44since 2021 · last 2026
0000-0002-2712-4412ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 4 first-author · 33 since 2021Artificial intelligence and machine learning · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | V-HOI: Velocity-Aware Human-Object Interaction Generation
Honghui Chen, Fan Zhou 0001, Ruomei Wang 0001, Baoquan Zhao |
MMM (2) | 3 |
| 2026 | An intelligent distributed CAD system for scalable thermal performance simulation in smart clothing engineering
Yi Teng, Ruomei Wang 0001, Yueping Guo |
Expert Syst. Appl. | 3 |
| 2026 | Expressive Human Volumetric Video Generation With Rich TextabstractPlain text has become the dominant interactive interface for text-driven human volumetric video generation. However, its limited customization options hinder users from expressing motion effects with accuracy. For example, plain text struggles to specify continuous variables such as motion amplitude, speed, and joint trajectories with precision, and it fails to convey stylized motion characteristics. Additionally, crafting detailed textual prompts for complex motion sequences is cumbersome, while excessively long prompts strain text encoders. To address these limitations, we propose a rich text-based framework that supports font styles, sizes, and trajectory sketching. By extracting motion-related attributes from rich text, our method enables fine-grained control over motion styles, precise speed regulation, and accurate joint trajectory manipulation. These capabilities are realized through gradient-guided noise editing and ControlNet-based motion optimization, which operate within the latent motion diffusion process. Specifically, we design a unified gradient-guided adaptation mechanism to ensure that the generated motion video adheres strictly to the specified constraints. Furthermore, we introduce realism-oriented optimization for stylistic and joint-level control, refining motion synthesis at a granular level to produce smoother, more natural movements. We present multiple comparative evaluations showcasing volumetric video generation from both rich text and plain text. Through quantitative analysis, we demonstrate that our method surpasses strong plain-text baselines, producing expressive, customizable human volumetric motion videos. Guanghui Yue 0001, Wei Zhou 0021, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | PSAM: Parameter-Free Spatiotemporal Attention Mechanism for Video Question AnsweringabstractSpatiotemporal attention learning has always been a challenging research task in video question answering (VideoQA). It needs to consider not only the modelling of local neighbourhood dependencies between the adjacent frames in a video but also the modelling of long-term dependencies between nonadjacent frames. Although the existing methods are usually good at modelling temporal dependencies in one aspect, they cannot simultaneously and effectively model the temporal dependencies between adjacent and nonadjacent frames. To address this issue, we first derive a novel statistic-driven difference-aware generation function, which can efficiently calculate the difference between a sequence feature value and the whole mean value to identify the significance of the feature. Subsequently, we design a novel parameter-free spatiotemporal attention mechanism (PSAM), which captures the most relevant cues scattered in the context of a spatiotemporal video by generating functions and utilizes a gating mechanism to adaptively integrate and filter relevant and irrelevant information. Finally, we use the PSAM and hierarchical modelling to construct a lightweight multiscale context fusion- and reasoning-based VideoQA model. Extensive experimental research results obtained on five benchmark datasets for the VideoQA task show that our VideoQA model has high Q&A performance and lightweight characteristics. Simultaneously, comprehensive ablation experimental results show that the PSAM can not only improve the performance of the model but also significantly reduce the number of model parameters. In addition, extensive experimental findings obtained on the benchmark dataset of joint tasks (video moment retrieval and video highlight detection) further demonstrate that the PSAM is a general and effective spatiotemporal attention mechanism. Ruomei Wang 0001, Fan Zhou 0001, Yuanmao Luo |
IEEE Trans. Multim. | 2 |
| 2026 | Visual-Guided Long Temporal Context Learning Network for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection (WVAD) aims to locate events or behaviors that deviate from normal patterns in untrimmed videos using video-level labels. Recent studies typically utilize supplementary modalities to assist anomaly detection. However, these methods suffer from two main issues: (1) The limitations of long-duration anomaly event temporal modeling. The model struggles to consistently maintain key information, resulting in the forgetting phenomenon, which affects the tracking of the event’s overall dynamic evolution and complicates anomaly event analysis and understanding. (2) The multi-modal fusion strategy is insufficient, particularly when there is temporal inconsistency between visual and audio information, causing the model to overlook key information, directly affecting the accurate detection and recognition of anomalous events. To address these issues, we propose a visual-guided long-term temporal context learning network (LTCLNet). The network consists of three key components: a cross-modal interaction module, a multi-modal fusion module, and a visual-guided parameter optimization strategy. First, to address the forgetting issue in long-duration anomaly detection, we designed a cross-modal interaction module. The key part of this module is the establishment of a cross-matrix mechanism. This mechanism achieves bidirectional temporal guidance across modalities. It allows the temporal modeling of each modality to dynamically integrate information from the other modality. This enables the model to continuously track the dynamic evolution of the event. The tracking is facilitated through shared temporal information between the visual and audio modalities. Secondly, to fully exploit the complementary characteristics between different modalities, we introduced a novel temporal reversal integration method in the multi-modal fusion module. This method reverses the feature sequences of each modality to enhance the model’s perception of temporal dynamic changes. By fusing the modality features before and after reversal, the shared temporal structure between modalities is strengthened, improving the model’s ability to capture anomalous information. Additionally, our proposed visual-guided parameter optimization strategy trains a parallel visual modality network as a semantic anchor, ensuring that the model stays aligned with a semantically stable and structurally clear visual flow during the learning process, thus ensuring stability and semantic coherence in the training. Extensive experiments on datasets such as XD-Violence demonstrate that our method significantly outperforms existing approaches, particularly achieving notable improvements in the accuracy and stability of long-term anomaly detection. Our code is publicly available at https://github.com/ibliever/LTCLNet . Ruomei Wang 0001, Linxuan Han, Baoquan Zhao, Fan Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | GameMLD: A Game-Sourced Motion-Language Dataset for Stylized Motion GenerationabstractText-guided character animation generation has emerged as a significant research area with broad applications in gaming, film, interactive media, and beyond. However, existing motion-language datasets face limitations in motion quality, stylistic diversity, and annotation depth, particularly for professional applications. In contrast to existing datasets based on motion capture or video reconstruction techniques, our dataset leverages professionally crafted game animations and employs a structured annotation framework that incorporates standardized game design terminology. The dataset contains 8,700 high-fidelity motion sequences paired with 26,100 multi-level textual descriptions, generated through our proposed annotation pipeline that combines domain expertise with large language models. Through comprehensive experiments and user studies, we demonstrate GameMLD’s advantages in motion quality, style expressiveness, and annotation quality. Additionally, we showcase its practical value by developing a text-driven character animation generation system that effectively supports game production pipelines. Our experiments with state-of-the-art motion synthesis models demonstrate significant improvements in both animation quality and style control. The GameMLD dataset and source code can be reached via this link. Yiyu Fu, Ziming Cheng, Yihao Liao, Jiangfeiyang Wang, Ruomei Wang 0001, Guanghui Yue 0001, Chenlei Lv, Baoquan Zhao |
ICME | 5 |
| 2025 | MSPoint-Gait: Multi-Scale Point Cloud Analysis for 3D Gait Recognition via Cross-Modal LearningabstractRecent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, we present MSPoint-Gait, a novel framework that addresses these challenges through multi-scale analysis and cross-modal learning. At the core of our framework lies a Depth-Aware Attention Module (DAAM) that leverages rich 3D geometric information to generate attention-weighted depth representations, enabling fine-grained feature extraction from point cloud sequences. We further introduce a Multi-Scale Spatio-Temporal (MSST) network that hierarchically captures both local and global gait patterns through adaptive convolution kernels across multiple spatial and temporal scales. These components are unified through a novel cross-modal learning strategy that effectively bridges the semantic gap between raw point clouds and structured depth representations. The proposed frame-work achieves state-of-the-art performance on the challenging SUSTech1K dataset, with 91.9% Rank-1 and 98.0% Rank-5 accuracy, demonstrating significant improvements over existing methods across various walking conditions and viewpoints. Xinzhu Li, Yikun Chen, Guanghui Yue 0001, Wei Zhou 0021, Ruomei Wang 0001, Xudong Mao, Juepeng Zheng, Fan Zhou 0001, Ziqi Qiu, Baoquan Zhao |
ICME | 6 |
| 2025 | Dynamic Feature-Focusing with Cross-Modal Semantic Alignment for Video Moment Retrieval and Highlight DetectionabstractVideo moment retrieval and highlight detection (MR&HD) is a challenging multimodal understanding task that requires precise temporal localization and saliency estimation. While existing approaches have achieved promising performance, they face two critical limitations, including static modeling strategies that fail to adapt to diverse video contents and durations, as well as significant semantic gaps between video and textual query representations that hinder effective cross-modal alignment. To address these challenges, this paper presents a novel adaptive semantic-guided framework for improved video MR&HD. First, we leverage video captions as semantic prior knowledge to enhance video representation and reduce the initial semantic gap. Second, we introduce an Adaptive Feature Focusing Module (AFF) that employs dynamic convolution to flexibly capture salient information across varying temporal scales. Third, we design a Multi-perspective Semantic Sensing Module (MSS) that combines attention mechanisms with text reconstruction to achieve robust cross-modal semantic alignment. Extensive experiments on four public benchmarks demonstrate that our approach consistently outperforms state-of-the-art methods. Ablation studies and visualizations further validate the effectiveness of each component and demonstrate our method’s capability in narrowing the semantic gap between heterogeneous features. Xuehui Liang, Ruomei Wang 0001, Baoquan Zhao |
ICME | 2 |
| 2025 | Multi-granularity Frequency Difference-Aware Attention for Video Question AnsweringabstractVideo Question Answering (VideoQA) demands complex reasoning about multi-granular information, requiring both fine-grained visual details and global event understanding from videos. While existing methods employ stacked cross-modal attention modules for multi-granular feature representation, they struggle to effectively separate different granularities due to the intertwined nature of visual information in videos. To address this challenge, we introduce a novel Multi-granularity Frequency Difference-Aware Attention (MFDA) mechanism that enhances VideoQA by modeling unified multi-granular relation-ships between multimodal features in the frequency domain. MFDA comprises three key components: a heterogeneous multi-granularity dynamic-aware module, a frequency distance-aware function, and a multi-granular complementary module. These components enable the VideoQA model to effectively parse and filter multi-granularity feature information, allocate attention weights, and provide precise visual semantic cues for answer prediction. Extensive experiments demonstrate that MFDA serves as a plug-and-play cross-modal attention mechanism that significantly improves existing VideoQA models’ performance, achieving state-of-the-art results across diverse question types while reducing computational complexity. Our code is available at https://github.com/haha94322/MFDA. Fan Zhou 0001, Ruomei Wang 0001, Baoquan Zhao |
ICME | 3 |
| 2025 | MCSMoG: Multi-Conditional Diffusion for Stylized Motion Generation with Parametric ControlabstractStylized human motion synthesis remains a fundamental challenge in computer animation and graphics, with a wide spectrum of applications spanning gaming, film production, virtual reality, and beyond. While recent advances in text-driven motion generation have shown promise, existing approaches face critical limitations including the inability to maintain consistent trajectory control, the lack of fine-grained stylization intensity adjustment, and inadequate generalization across diverse motion styles. To address these challenges, We introduce MCSMoG, a novel framework for controllable stylized motion synthesis through multi-conditional guidance. First, a new Multi-Conditional Motion Latent Diffusion (MC-MLD) model is proposed to introduce additional trajectory guidance and achieve trajectory decoupling. Second, we develop a Style and Non-Style Feature Fusion Module that dynamically blends motion features through an adjustable parameter, providing control over stylization intensity. Third, we integrate MotionCLIP as our style encoder, enhancing the model’s generalization capability across diverse and unseen motion styles. Extensive experiments conducted on the combined HumanML3D and 100STYLE datasets demonstrate that our approach outperforms state-of-the-art methods, achieving a 4.6% reduction in FID scores and a 4.1% increase in motion diversity. User studies further confirm the superiority of our method in style fidelity, semantic consistency, and motion naturalness. Xinzhu Li, Guanghui Yue 0001, Wei Zhou 0021, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
ICME | 7 |
| 2025 | DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait RecognitionabstractRobust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture finer and meaningful details of gait. In this paper, we introduce a novel framework, termed DepthGait, that incorporates RGB-derived depth maps and silhouettes for enhanced gait recognition. Specifically, apart from the 2D silhouette representation of the human body, the proposed pipeline explicitly estimates depth maps from a given RGB image sequence and uses them as a new modality to capture discriminative features inherent in human locomotion. In addition, a novel multi-scale and cross-level fusion scheme has also been developed to bridge the modality gap between depth maps and silhouettes. Extensive experiments on standard benchmarks demonstrate that the proposed DepthGait achieves state-of-the-art performance compared to peer methods and attains an impressive mean rank-1 accuracy on the challenging datasets. Xinzhu Li, Juepeng Zheng, Yikun Chen, Xudong Mao, Guanghui Yue 0001, Wei Zhou 0021, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
ACM Multimedia | 8 |
| 2025 | VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question AnsweringabstractCross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video streams and managing the complexity of multi-source information retrieval. We introduce VideoForest, a novel framework that addresses these challenges through person-anchored hierarchical reasoning, enabling effective cross-video understanding without requiring end-to-end training. VideoForest integrates three key innovations: 1) a human-anchored feature extraction mechanism that employs ReID and tracking algorithms to establish robust spatiotemporal relationships across multiple video sources; 2) a multi-granularity spanning tree structure that hierarchically organizes visual content around person-level trajectories; and 3) a multi-agent reasoning framework that efficiently traverses this hierarchical structure to answer complex queries. To evaluate our method, we develop CrossVideoQA, a comprehensive benchmark specifically designed for person-centric cross-video analysis. Experimental results demonstrate VideoForest's superior performance in cross-video reasoning tasks, achieving 71.93% accuracy in person recognition, 83.75% in behavior analysis, and 51.67% in summarization and reasoning. Yiran Meng, Junhong Ye, Wei Zhou 0021, Guanghui Yue 0001, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao |
ACM Multimedia | 6 |
| 2025 | VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual AugmentationsabstractThe widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in speech-rich video content. In contrast to other video types, such as surveillance footage, which typically contain abundant visual cues, speech-rich videos convey most of their meaningful information through the audio channel. This poses challenges for improving content consumption using existing visual-based video summarization, navigation, and exploration systems. In this paper, we present VisAug, a novel interactive system designed to enhance speech-rich video navigation and engagement by automatically generating informative and expressive visual augmentations based on the speech content of videos. Our findings suggest that this system has the potential to significantly enhance the consumption and engagement of information in an increasingly video-driven digital landscape. Baoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang 0001, Fan Zhou 0001, Shujin Lin |
ACM Multimedia | 4 |
| 2024 | Improved Text-Driven Human Motion Generation via Out-of-Distribution Detection and Rectification
Yiyu Fu, Baoquan Zhao, Chenlei Lv, Guanghui Yue 0001, Ruomei Wang 0001, Fan Zhou 0001 |
CVM (1) | 5 |
| 2024 | Clip-Medfake: Synthetic Data Augmentation With AI-Generated Content for Improved Medical Image ClassificationabstractData augmentation is serving as a critical and fundamental technology to improve model generalization and performance in a wide spectrum of machine learning tasks. Despite the increasing interest in developing various pathways to artificially generate new data to reduce the overfitting issue during model training, enriching the diversity of training data in the field of medicine remains facing enormous challenges. By virtue of recent advancements in generative artificial intelligence, we present a novel data augmentation framework, CLIP-MedFake, to address the shortage of training data used in medical image classification. The proposed method first employs the Stable Diffusion model to generate new fake data based on a small amount of training data, and then adopts the paradigm of few-shot learning and uses the CLIP architecture as the backbone to pre-train the model with synthetic data and then fine-tune it with real medical images. Extensive experiment results on two publicly available datasets demonstrate the effectiveness of the proposed method in promoting medical image classification. Honghui Chen, Baoquan Zhao, Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001 |
ICIP | 6 |
| 2024 | Frequency-Domain Enhanced Cross-modal Interaction Mechanism for Joint Video Moment Retrieval and Highlight DetectionabstractThe joint video moment retrieval and highlight detection aims to locate moments and highlight clips relevant to textual queries. The main challenges lies in the fusion of multi-modal data and understanding semantic information. Current research for addressing this issue primarily focus on capturing information from different modalities in temporal order. Due to the heterogeneity of different modalities, it is difficult to capture all complementary information in temporal domain. Therefore, we propose a frequency-domain enhanced cross-modal interaction mechanism (FCIM) to capture complementary information from different modalities across all frequency components, addressing the limitations in the temporal domain. Additionally, we design a lightweight inference model for the joint (LIMJ) task. Extensive experiments on three datasets demonstrate the effectiveness and generality of our method. Ruomei Wang 0001, Yuanmao Luo |
ICME | 2 |
| 2024 | Text-Based Vector Sketch Editing with Image Editing Diffusion PriorabstractWe present a framework for text-based vector sketch editing to improve the efficiency of graphic design. The key idea behind the approach is to transfer the prior information from raster-level diffusion models, especially those from image editing methods, into the vector sketch-oriented task. The framework presents three editing modes and allows iterative editing. To meet the editing requirement of modifying the intended parts only while avoiding changing the other strokes, we introduce a stroke-level local editing scheme that automatically produces an editing mask reflecting locally editable regions and modifies strokes within the regions only. Comparisons with existing methods demonstrate the superiority of our approach. Haoran Mo, Xusheng Lin, Chengying Gao, Ruomei Wang 0001 |
ICME | 4 |
| 2024 | Hierarchical Attention Feature Fusion and Refinement Network for Point Cloud UpsamplingabstractThis paper presents a novel hierarchical attention feature fusion and refinement network designed to address challenges in existing deep learning based point cloud upsampling methods. The network combines self-attention layers with a multi-level feature extraction architecture, effectively integrating local and global features, thereby enhancing the robustness and uniformity of the point cloud. Furthermore, a spatial refinement module is employed to predict the offset between the generated coarse dense point clouds and real point clouds, thereby enhancing consistency with the ground truth. Concurrently, a filter function is applied in the loss function to handle outliers of generated point clouds. Extensive experimental results across multiple datasets indicate that our method outperforms existing approaches. Yaori Zhang, Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001 |
ICME | 4 |
| 2024 | Utilizing Text-Video Relationships: A Text-Driven Multi-modal Fusion Framework for Moment Retrieval and Highlight Detection
Ruomei Wang 0001, Zhuo Su 0001 |
PRCV (10) | 3 |
| 2024 | Predicting Plain Text Imageability for Faithful Prompt-Conditional Image Generation
Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
PRICAI (3) | 5 |
| 2024 | Dual-branch Complementary Multimodal Interaction Mechanism for Joint Video Moment Retrieval and Highlight DetectionabstractJoint video moment retrieval and highlight detection is a video understanding task that requires the model to construct multimodal interaction between heterogeneous features. Recent Transformer-based models mainly focus on promoting global interaction between features. However, local interaction and temporal asynchronism modeling are not deeply considered. To solve this problem, this paper proposes a dual-branch complementary multimodal interaction mechanism (DCMI), which consists of a global difference feature activation module (GDFA) and a local information dynamic aggregation module (LIDA). GDFA measures the difference between the target element and the global features, thus activating important information. LIDA designs a multimodal heterogeneous graph and constructs asynchronous interaction between heterogeneous features to dynamically aggregate local information. DCMI adaptively fuses the complementary dual branches to improve the model's cognitive and decision-making abilities of global and local information. Comprehensive comparisons with existing methods on public datasets verify the superiority of the proposed model. Extensive ablation experiments and qualitative analysis show the effectiveness and rationality of DCMI, which can promote the interaction between multimodal features. Xuehui Liang, Ruomei Wang 0001, Ge Lin 0002, Yuanmao Luo |
SMC | 2 |
| 2024 | Video Q &A based on two-stage deep exploration of temporally-evolving features with enhanced cross-modal attention mechanism
Yuanmao Luo, Ruomei Wang 0001, Fan Zhou 0001 |
Neural Comput. Appl. | 2 |
| 2024 | Modality-Aware Heterogeneous Graph for Joint Video Moment Retrieval and Highlight DetectionabstractThe joint task of video moment retrieval and video highlight detection is a challenging study, which requires building a model that not only captures contextual information between sequences in time but also has the ability to understand and judge significance. This paper solves these problems from three aspects. Firstly, we design a parameter-free cross-modal statistical correlation interaction method. A novel saliency enhancement function is defined to quantify the saliency differences between the important features associated with the query and other features to achieve parameter-free cross-modal fusion. Secondly, we propose a novel modality-aware heterogeneous graph reasoning mechanism (MHGR). MHGR can effectively capture the global context information between sequences, enhance the local association relationship between sequences, and deal with the complexity of multi-modal data better through the organic combination of two key modules: parameter-free cross-modal statistical correlation interaction, and heterogeneous graph reasoning mechanism. Thirdly, a lightweight solution for the joint task of video moment retrieval and highlight detection is designed based on the above two novel algorithm modules. Comprehensive experiments are conducted on publicly available benchmark data to validate the advantages of the new solution in comparison with a series of state-of-the-art peer methods. Quantitative results consistently demonstrate that the new solution is lightweight and has high inference performance so the remarkable improvement in accuracy achieved by the new solution with respect to peer methods. An extended ablation study is further conducted to show the usefulness of each module of the solution in acquiring its computational capabilities. Ruomei Wang 0001, Yuanmao Luo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Subtask Prior-Driven Optimized Mechanism on Joint Video Moment Retrieval and Highlight DetectionabstractJoint video moment retrieval and highlight detection is an emerging and challenging research task. It requires the generation of robust joint task features to satisfy the demands of video moment retrieval and video highlight detection. Moreover, it involves the interaction of multiple modalities. Presently, methods typically focus on the design of distinct enhancement modules and the addition of supplementary input data to improve the solution for joint video moment retrieval and highlight detection. However, they overlook subtask interference during joint training. Joint task learning leverages the correlations and complementarities between tasks, yet it also introduces task interference arising from the differences between tasks. In order to address task interference, we proposes a subtask prior-driven optimized mechanism. The mechanism consists of two stages. In the free stage, we train subtask model to get subtask prior features. In the constrained stage, the joint task model is constrained by the subtask. Besides, we propose a cross adaptive-gated mechanism. It addresses the issue of information loss in cross-modal fusion and filters out redundant information by conducting cross-modal interaction during feature compression and an adaptive gating process. Extensive experimental results exhibit the effectiveness of the subtask prior-driven optimized mechanism and the cross adaptive-gated transformer in joint video moment retrieval and highlight detection. Ruomei Wang 0001, Fan Zhou 0001, Zhuo Su 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | DMAP: Decoupling-Driven Multi-Level Attribute Parsing for Interpretable Outfit CollocationabstractOutfit collocation requires considering the interrelationship and adaptability among the attributes of component items. However, with the numerous and diverse attributes of fashion items, accurately capturing attribute features and modeling the complex relationships between attributes become the key challenges. To address these challenges, we propose a novel scheme Decoupling-driven Multi-level Attribute Parsing for interpretable outfit collocation. First, we decouple a series of attribute features from the item's visual feature by fully supervised, which can improve the robustness of the model in processing both relevant and irrelevant attributes of items. Furthermore, employing a deep deconvolution neural network with attention mechanisms to reconstruct the decoupled attribute features into a visual image that is close to the original item image. It ensures all attribute features can be combined to contain complete item information. Next, graph attention networks are constructed to parse multi-level attribute compatibility relationships from three perspectives: intra-attribute, inter-attribute, and item integration relationships. Finally, we use multi-layer perceptrons to fuse the score distributions of the three and output the outfit compatibility score. Experiments conducted on the IQON3000 dataset demonstrate that our model outperforms existing state-of-the-art methods and exhibits good interpretability. Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Ge Lin 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | HSSHG: Heuristic Semantics-Constrained Spatio-Temporal Heterogeneous Graph for VideoQAabstractVideo question answering is a challenging task that requires models to recognize visual information in videos and perform spatio-temporal reasoning. Current models increasingly focus on enabling objects spatio-temporal reasoning via graph neural networks. However, the existing graph network-based models still have deficiencies when constructing the spatio-temporal relationship between objects: (1) The lack of consideration of the spatio-temporal constraints between objects when defining the adjacency relationship; (2) The semantic correlation between objects is not fully considered when generating edge weights. These make the model lack representation of spatio-temporal interaction between objects, which directly affects the ability of object relation reasoning. To solve the above problems, this paper designs a heuristic semantics-constrained spatio-temporal heterogeneous graph, employing a semantic consistency-aware strategy to construct the spatio-temporal interaction between objects. The spatio-temporal relationship between objects is constrained by the object co-occurrence relationship and the object consistency. The plot summaries and object locations are used as heuristic semantic priors to constrain the weights of spatial and temporal edges. The spatio-temporal heterogeneity graph more accurately restores the spatio-temporal relationship between objects and strengthens the model's object spatio-temporal reasoning ability. Based on the spatio-temporal heterogeneous graph, this paper proposes Heuristic Semantics-constrained Spatio-temporal Heterogeneous Graph for VideoQA (HSSHG), which achieves state-of-the-art performance on benchmark MSVD-QA and FrameQA datasets, and demonstrates competitive results on benchmark MSRVTT-QA and ActivityNet-QA dataset. Extensive ablation experiments verify the effectiveness of each component in the network and the rationality of hyperparameter settings, and qualitative analysis verifies the object-level spatio-temporal reasoning ability of HSSHG. Ruomei Wang 0001, Yuanmao Luo |
IEEE Trans. Multim. | 1 |
| 2024 | Joint Stroke Tracing and Correspondence for 2D AnimationabstractTo alleviate human labor in redrawing keyframes with ordered vector strokes for automatic inbetweening, we for the first time propose a joint stroke tracing and correspondence approach. Given consecutive raster keyframes along with a single vector image of the starting frame as a guidance, the approach generates vector drawings for the remaining keyframes while ensuring one-to-one stroke correspondence. Our framework trained on clean line drawings generalizes to rough sketches, and the generated results can be imported into inbetweening systems to produce inbetween sequences. Hence, the method is compatible with standard 2D animation workflow. An adaptive spatial transformation module (ASTM) is introduced to handle non-rigid motions and stroke distortion. We collect a dataset for training with 10k+ pairs of raster frames and their vector drawings with stroke correspondence. Comprehensive validations on real clean and rough animated frames manifest the effectiveness of our method and superiority to existing methods. Haoran Mo, Chengying Gao, Ruomei Wang 0001 |
ACM Trans. Graph. | 3 |
| 2023 | MIM: Lightweight Multi-Modal Interaction Model for Joint Video Moment Retrieval and Highlight DetectionabstractJoint video moment retrieval and highlight detection aims to find the relevant moments and highlight clips in a video with natural language. It is an emerging task though its individual problems have been studied for a while. The current methods utilize transformer to interact between modals, which leads to a huge cost of parameters and computation in spite of great performance. To address this problem, we present a cross-modal attention mechanism to capture related features from different modalities in a few-parameter way. Furthermore, a lightweight multi-modal interaction model (MIM) is proposed to solve video moment retrieval and highlight detection jointly. In the case of greatly reducing the number of parameters, we achieve competitive performance and faster convergence speed compared to previous method. Extensive experiments on four datasets demonstrate the effectiveness of our method. Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001 |
ICME | 5 |
| 2023 | SyntaxLineDP: a Line-level Software Defect Prediction Model based on Extended Syntax InformationabstractThe existence of software defects greatly restricts the application of software and may bring great economic losses. Defect prediction techniques based on file level or module level can help code reviewers quickly locate defects and fix them. In recent years, researchers have used natural language techniques to explore the defectiveness of code lines by analyzing the semantic information of the code lines. However, the appearance of defects is often associated with the syntactic information of the code. To fully use the syntax information associated with code lines, we propose a code line-level representation considering the coverage of the syntax node. Specifically, we add the syntax node to its corresponding coverage lines. Combining with the BiLSTM model, we propose a new line-level defect prediction model, SyntaxLineDP, which shows good performance achieving 0.72 and 0.62 for AUC and Balanced Accuracy respectively. SyntaxLineDP outperforms the state-of-the-art model and the popular static tools (e.g., ErrorProne, PMD). Specifically, In the within-project evaluation, the SyntaxLineDP model outperforms the start-of-the-art model by 15% and 27% on the Recall@Top20% LOC metric and the Effort@Top20% Recall metric respectively. In the cross-project evaluation, the SyntaxLineDP outperforms the start-of-the-art model by 7% and 42% on the Recall@Top20% LOC metric and the Effort@Top20% Recall metric respectively. Through ablation experiments, we studied how the selection of extended syntax information contributes to the detection and found out some key features that are helpful for the defect prediction task: number of infix expressions of code line, number of blocks of code line, and number of method invocation of the code line. Jianzhong Zhu, Yuan Huang 0002, Xiangping Chen, Ruomei Wang 0001, Zibin Zheng |
QRS | 4 |
| 2023 | ERM: Energy-Based Refined-Attention Mechanism for Video Question AnsweringabstractSpatiotemporal attention learning remains a challenging video question answering (VideoQA) task as it requires a sufficient understanding of cross-modal spatiotemporal information. Existing methods usually leverage different cross-modal attention mechanisms to reveal potential associations between video and question. While these methods effectively remove irrelevant information from the spatiotemporal attention, they ignore the pseudo-related information within the cross-modal interaction attention. To address this problem, we proposed a novel energy-based refined-attention mechanism (ERM). ERM leverages the significant difference distribution as a discriminative criterion derived from question-guided cross-modal interaction information to determine question-related and question-irrelated cross-modal interaction information. The specific method is to measure the linear separability between the target neuron and other neurons in the neural network to confirm the importance of neurons. In addition, to solve the statistical bias caused by the differences between different modes in video tasks, the ERM proposed in this paper has learnable parameters. The correlation between different modes can be learned adaptively through learnable parameters. The advantages of the proposed ERM are that it is more flexible and modular while remaining lightweight. With the help of the ERM, we construct a lightweight VideoQA model that efficiently integrates the cross-modal feature representations in an energy-based manner. To evaluate the effectiveness of our method, we carried out extensive experiments on five publicly available datasets and compared them with state-of-the-art VideoQA methods. The experiment results demonstrate that our method brings a noticeable performance improvement compared to state-of-the-art VideoQA methods. ERM can be flexibly integrated into different VideoQA methods to improve their Q&A performance. Ruomei Wang 0001, Fan Zhou 0001, Yuanmao Luo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Arbitrary-Shape Scene Text Detection via Visual-Relational Rectification and Contour ApproximationabstractOne trend in the latest bottom-up approaches for arbitrary-shape scene text detection is to determine the links between text segments using Graph Convolutional Networks (GCNs). However, the performance of these bottom-up methods is still inferior to that of state-of-the-art top-down methods even with the help of GCNs. We argue that a cause of this is that bottom-up methods fail to make proper use of visual-relational features, which results in accumulated false detection, as well as the error-prone route-finding used for grouping text segments. In this paper, we improve classic bottom-up text detection frameworks by fusing the visual-relational features of text with two effective false positive/negative suppression (FPNS) mechanisms and developing a new shape-approximation strategy. First, dense overlapping text segments depicting the “characterness” and “streamline” properties of text are constructed and used in weakly supervised node classification to filter the falsely detected text segments. Then, relational features and visual features of text segments are fused with a novel Location-Aware Transfer (LAT) module and Fuse Decoding (FD) module to jointly rectify the detected text segments. Finally, a novel multiple-text-map-aware contour-approximation strategy is developed based on the rectified text segments, instead of the error-prone route-finding process, to generate the final contour of the detected text. Experiments conducted on five benchmark datasets demonstrate that our method outperforms the state-of-the-art performance when embedded in a classic text detection framework, which revitalizes the strengths of bottom-up methods. Chengpei Xu, Wenjing Jia, Tingcheng Cui, Ruomei Wang 0001, Yuan-fang Zhang, Xiangjian He |
IEEE Trans. Multim. | 4 |
| 2023 | MorphText: Deep Morphology Regularized Accurate Arbitrary-Shape Scene Text DetectionabstractBottom-up text detection methods play an important role in arbitrary-shape scene text detection but there are two restrictions preventing them from achieving their great potential, i.e., 1) the accumulation of false text segment detections, which affects subsequent processing, and 2) the difficulty of building reliable connections between text segments. Targeting these two problems, we propose a novel approach, named ``MorphText", to capture the regularity of texts by embedding deep morphology for arbitrary-shape text detection. Towards this end, two deep morphological modules are designed to regularize text segments and determine the linkage between them. First, a Deep Morphological Opening (DMOP) module is constructed to remove false text segment detections generated in the feature extraction process. Then, a Deep Morphological Closing (DMCL) module is proposed to allow text instances of various shapes to stretch their morphology along their most significant orientation while deriving their connections.Extensive experiments conducted on four challenging benchmark datasets (CTW1500, Total-Text, MSRA-TD500 and ICDAR2017) demonstrate that our proposed MorphText outperforms both top-down and bottom-up state-of-the-art arbitrary-shape scene text detection approaches. Chengpei Xu, Wenjing Jia, Ruomei Wang 0001, Xiangjian He |
IEEE Trans. Multim. | 3 |
| 2022 | NewsThumbnail: Automatic Generation of News Video ThumbnailabstractReading news is an important way for people to obtain information. People can quickly sort out the context of events through a short news video. However, there are numerous news generated around the world every day. It’s challenging to locate the interesting video. Thumbnails are often used as video covers and play an important role in displaying video content and driving views. Video owners can choose from individual images or elaborate thumbnails to upload to the site. But manually selecting from a large number of frames is time-consuming, and customizing thumbnails requires a high degree of expertise. Therefore, this paper proposes an automatic generation method of news video thumbnail, which can screen out semantically similar contents according to user query and combine them into a thumbnail. In order to facilitate the screening of graphic materials, we also propose a video content structuring method based on multiple cues, which can accurately segment the video into theme units. At the same time, we designed a visual system to display thumbnails and designed a user survey to investigate the performance of this method in news retrieval and understanding. Compared with peer methods, the thumbnails generated by our method can help users better understand the video content and locate the videos they are interested in. Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001 |
SMC | 4 |
| 2022 | Temporal-aware Mechanism with Bidirectional Complementarity for Video Q&AabstractVideo question answering (Video Q&A) is a challenging task as it requires a sufficient understanding of the video and question information. Video is composed of frame sequence, which contains multi-scale temporal relationships and corresponding contextual information. A model competently tackle Video Q&A task that needs to be able to: 1) construct long-term and neighborhood dependencies in frame sequences to extract global and local contextual features that can reflect multi-scale temporal dependencies, and deduce the temporal-aware refined features, and 2) identify static and dynamic features from pertinent moments of a video, while filtering away question-irrelated dependencies of feature sequences, to yield the most precise and reasonable temporal-aware overall contextual features. In response to the above requirements, we propose a novel Video Q&A mechanism which consists of Bidirectional Complementary Attention(BCA) module and Adaptive Temporal-aware(ATA) module. Bidirectional complementary attention module stacks multi-head self-attention layer and convolutional layer in different orders to designed two kinds of attention units, which is able to make bidirectional multi-step reasoning based on complete global information and accurate local information to obtain temporal-aware refined features. Adaptive temporal-aware module is used to filter away question-irrelated dependencies in the feature sequence to yield the most precise and reasonable temporal-aware overall contextual features. Comprehensive comparative experiments are conducted on publicly available benchmark datasets. An extended ablation study is further conducted to show the usefulness of each module of the solution in acquiring its computational Q&A capabilities. Yuanmao Luo, Ruomei Wang 0001, Fan Zhou 0001, Shujin Lin |
SMC | 2 |
| 2022 | Learning compatibility knowledge for outfit recommendation with complementary clothing matching
Ruomei Wang 0001, Zhuo Su 0001 |
Comput. Commun. | 1 |
| 2022 | Attribute-aware heterogeneous graph network for fashion compatibility prediction
Zhouyi Zhou, Zhuo Su 0001, Ruomei Wang 0001 |
Neurocomputing | 3 |
| 2022 | Multistage Spatio-Temporal Networks for Robust Sketch RecognitionabstractSketch recognition relies on two types of information, namely, spatial contexts like the local structures in images and temporal contexts like the orders of strokes. Existing methods usually adopt convolutional neural networks (CNNs) to model spatial contexts, and recurrent neural networks (RNNs) for temporal contexts. However, most of them combine spatial and temporal features with late fusion or single-stage transformation, which is prone to losing the informative details in sketches. To tackle this problem, we propose a novel framework that aims at the multi-stage interactions and refinements of spatial and temporal features. Specifically, given a sketch represented by a stroke array, we first generate a temporal-enriched image (TEI), which is a pseudo-color image retaining the temporal order of strokes, to overcome the difficulty of CNNs in leveraging temporal information. We then construct a dual-branch network, in which a CNN branch and a RNN branch are adopted to process the stroke array and the TEI respectively. In the early stages of our network, considering the limited ability of RNNs in capturing spatial structures, we utilize multiple enhancement modules to enhance the stroke features with the TEI features. While in the last stage of our network, we propose a spatio-temporal enhancement module that refines stroke features and TEI features in a joint feature space. Furthermore, a bidirectional temporal-compatible unit that adaptively merges features in opposite temporal orders, is proposed to help RNNs tackle abrupt strokes. Comprehensive experimental results on QuickDraw and TU-Berlin demonstrate that the proposed method is a robust and efficient solution for sketch recognition. Xudong Jiang 0001, Boliang Guan, Ruomei Wang 0001, Nadia Magnenat-Thalmann |
IEEE Trans. Image Process. | 4 |
| 2022 | Popularity-Guided Cost Optimization for Live Streaming in Mobile Edge ComputingabstractLive streaming service usually delivers the content in mobile edge computing (MEC) to reduce the network latency and save the backhaul capacity. Considering the limited resources, it is necessary that MEC servers collaborate with each other and form an overlay to realize more efficient delivery. The critical challenge is how to optimize the topology among the servers and allocate the link capacity so that the cost will be lower with delay constraints. Previous approaches rarely consider server collaborations for live streaming service, and the scheduling delay is usually ignored in MEC, leading to suboptimal performances. In this paper, we propose a popularity‐guided overlay model which takes the scheduling delay into consideration and utilizes MEC collaboration to achieve efficient live streaming service. The links and servers are shared among all channel streams and each stream is pushed from cloud servers to MEC servers via the trees. Considering the optimization problem is NP‐hard, we propose an effective optimization framework called cost optimization for live streaming (COLS) to predict the channel popularity by a LSTM model with multiscale input data. Finally, we compute topology graph by greedy scheme and allocate the capacity with convex programming. Experimental results show that the proposed approach achieves higher prediction accuracy, reducing the capacity cost by more than 40% with an acceptable delay compared with state‐of‐the‐art schemes. Tao He 0012, Kunxin Zhu, Ruomei Wang 0001, Fan Zhou 0001 |
Wirel. Commun. Mob. Comput. | 4 |
| 2021 | Learning Outfit Compatibility with Graph Attention Network and Visual-Semantic EmbeddingabstractFashion recommendation is an essential component of user shopping that it is capable of selecting and presenting fascinating items to customers. The fact that humans exhibit inconsistencies for fashion items in their choice is known to all due to the visual aesthetic features and fine-grained differences of fashion items. Previous research on fashion recommendations mainly focuses on sequential models, most of them only consider complex similarity relationships in fashion compatibility while neglecting the real-world compatible information often desired in practical applications. To learn the fashion compatibility and generate for the outfit, we propose an approach that jointly learns latent fashion concepts in visual-semantic space to measure compatibility between items. The fashion concepts are shaped by design elements such as color, material, and silhouette. Accordingly, we model a unified representation to learn different notions of similarity by mapping text descriptors and images into latent space to learn high-level representations. Experimental results reveal that our method effectively reaches the aimed results on the fill-in-the-blank and outfit compatibility tasks. Xiaochun Cheng, Ruomei Wang 0001, Shaohui Liu |
ICME | 3 |
| 2021 | Fusing Temporally Distributed Multi-Modal Semantic Clues for Video Question AnsweringabstractVideo Question Answering (VideoQA) is an intriguing topic, attracting increasing interest among the broad AI community. Yet videoQA is a difficult task. An algorithm competently tackle this task that needs to be able to: 1) extract rich semantics supplied in each modality of a video and incorporate them across modalities, and 2) identify and integrate such multimodal semantics from pertinent moments of a video, which may or may not be temporally adjacent or nearby, while filtering away irrelevant or even detractive portions of the video, to yield the most precise and sensible semantic context for executing the QA task. In response to the above requirements, a novel deep VideoQA solution is proposed in this paper, which comprises a multi-modal semantic clue extraction module, driven by a series of deep networks, each dedicated to digesting signals of a distinct modality type, to develop the first algorithmic QA capability, and a multi-modal temporal QA module empowered by a deep graph attention network to build the second algorithmic QA capability. Comprehensive experiments are conducted on publicly available benchmark data to validate advantages of the new solution in the end. Ruomei Wang 0001, Songhua Xu, Fan Zhou 0001 |
ICME | 2 |
| 2021 | News2Mapping: A news events correlation model for news videosabstractNews video is an important way of news communication, and people can easily get news from all over the world through the Internet. However, it lacks in the organization of news video content about temporality, presentation and relevance and fails to express the correlation between news events. In this paper, we present an event correlation model of news videos to organize news content. News content is clustered through topics, and the relationship between events is illustrated through relationship mappings. To achieve this purpose, an XLNet-based language model is presented to extract news keywords and their relationships. The clustering algorithm is designed to obtain news event topic clustering and named entity clustering. At the same time, we build the relationship mappings in news events to visualize the correlation between news events better. The user study is also designed to investigate the performance of our method in news reading and understanding. Compared with peer methods, the news information organized by our method achieves a higher user satisfaction level. Mingjie Zhou, Ruomei Wang 0001, Shujin Lin, Fan Zhou 0001, Shirou Ou |
SMC | 2 |
| 2021 | LGCPNet : Local-global combined point-based network for shape segmentation
Boliang Guan, Fan Zhou 0001, Shujin Lin, Ruomei Wang 0001 |
Comput. Graph. | 5 |
| 2021 | Joint Feature Optimization and Fusion for Compressed Action RecognitionabstractRecent methods including CoViAR and DMC-Net provide a new paradigm for action recognition since they are directly targeted at compressed videos (e.g., MPEG4 files). It avoids the cumbersome decoding procedure of traditional methods, and leverages the pre-encoded motion vectors and residuals in compressed videos to complete recognition efficiently. However, motion vectors and residuals are noisy, sparse and highly correlated information, which cannot be effectively exploited by plain and separated networks. To tackle these issues, we propose a joint feature optimization and fusion framework that better utilizes motion vectors and residuals in the following three aspects. (i) We model the feature optimization problem as a reconstruction process that represents features by a set of bases, and propose a joint feature optimization module that extracts bases in the both modalities. (ii) A low-rank non-local attention module, which combines the non-local operation with the low-rank constraint, is proposed to tackle the noise and sparsity problem during the feature reconstruction process. (iii) A lightweight feature fusion module and a self-adaptive knowledge distillation method are introduced, which use motion vectors and residuals to generate predictions similar to those from networks with optical flows. With these proposed components embedded in a baseline network, the proposed network not only achieves the state-of-the-art performance on HMDB-51 and UCF-101, but also maintains its advantage in computational complexity. Xudong Jiang 0001, Boliang Guan, Raymond Rui Ming Tan, Ruomei Wang 0001, Nadia Magnenat-Thalmann |
IEEE Trans. Image Process. | 5 |
| 2021 | General virtual sketching framework for vector line artabstractVector line art plays an important role in graphic design, however, it is tedious to manually create. We introduce a general framework to produce line drawings from a wide variety of images, by learning a mapping from raster image space to vector image space. Our approach is based on a recurrent neural network that draws the lines one by one. A differentiable rasterization module allows for training with only supervised raster data. We use a dynamic window around a virtual pen while drawing lines, implemented with a proposed aligned cropping and differentiable pasting modules. Furthermore, we develop a stroke regularization loss that encourages the model to use fewer and longer strokes to simplify the resulting vector image. Ablation studies and comparisons with existing methods corroborate the efficiency of our approach which is able to generate visually better results in less computation time, while generalizing better to a diversity of images and applications. Haoran Mo, Edgar Simo-Serra, Chengying Gao, Changqing Zou, Ruomei Wang 0001 |
ACM Trans. Graph. | 5 |
| 2020 | Multi-column point-CNN for sketch segmentation
Fei Wang 0056, Shujin Lin, Hefeng Wu, Tie Cai, Ruomei Wang 0001 |
Neurocomputing | 7 |
| 2020 | Learning rebalanced human parsing model from imbalanced datasets
Enbo Huang, Zhuo Su 0001, Fan Zhou 0001, Ruomei Wang 0001 |
Image Vis. Comput. | 4 |
| 2020 | Voxel-based quadrilateral mesh generation from point cloud
Boliang Guan, Shujin Lin, Ruomei Wang 0001, Fan Zhou 0001, Yongchuan Zheng |
Multim. Tools Appl. | 3 |
| 2019 | Learning Transmission Filtering Network for Image-Based Pm2.5 EstimationabstractPM2.5 is an important indicator of the severity of air pollution and its level can be predicted through hazy photographs caused by its degradation. Image-based PM2.5 estimation is thus extensively employed in various multimedia applications but is challenging because of its ill-posed property. In this paper, we convert it to the problem of estimating the PM2.5-relevant haze transmission and propose a learning model called the transmission filtering network. Different from most methods that generate a transmission map directly from a hazy image, our model takes the coarse transmission map derived from the dark channel prior as the input. To obtain a transmission map that satisfies the local smoothness constraint without regional boundary degradation, our model performs the edge-preserving smoothing filtering as the refinement on the map. Moreover, we introduce the attention mechanism to the network architecture for more efficient feature extraction and smoothing effects in the transmission estimation. Experimental results prove that our model performs favorably against the state-of-the-art dehazing methods in a variety of hazy scenes. Yinghong Liao, Bin Qiu, Zhuo Su 0001, Ruomei Wang 0001, Xiangjian He |
ICME | 4 |
| 2019 | SPFusionNet: Sketch Segmentation Using Multi-modal Data FusionabstractThe sketch segmentation problem remains largely unsolved because conventional methods are greatly challenged by the highly abstract appearances of freehand sketches and their numerous shape variations. In this work, we tackle such challenges by exploiting different modes of sketch data in a unified framework. Specifically, we propose a deep neural network SPFusionNet to capture the characteristic of sketch by fusing from its image and point set modes. The image modal component SketchNet learns hierarchically abstract ro-bust features and utilizes multi-level representations to produce pixel-wise feature maps, while the point set-modal component SPointNet captures local and global contexts of the sampled point set to produce point-wise feature maps. Then our framework aggregates these feature maps by a fusion network component to generate the sketch segmentation result. The extensive experimental evaluation and comparison with peer methods on our large SketchSeg dataset verify the effectiveness of the proposed framework. Fei Wang 0056, Shujin Lin, Hefeng Wu, Ruomei Wang 0001, Xiangjian He |
ICME | 5 |
| 2019 | Lecture2Note: Automatic Generation of Lecture Notes from Slide-Based Educational VideosabstractGiven rapid development witnessed by open educational resources (OER) in the past few decades, a considerable number of online educational videos emerge on various MOOC platforms such as Coursera and YouTube. Nevertheless, most educational videos on the internet are lengthy and lack of elaborate annotations, which poses a challenge for learners to explore and locate content of interest efficiently. To address this, we present an automatic note-generating method to establish correspondences between visual entities in the slide-based lecture video and their descriptive speech texts by evaluating the semantic relationship. Firstly, the visual entities are extracted and recognised from the presentation slides. Then, each of visual entities is associated with its corresponding descriptive speech text. Finally, a placement optimisation scheme is put forward to pack the visual entities and speech texts into a note-like layout in a compact fashion, which can help learners to improve their learning efficiency. The experimental results show that the efficient performances about visual entity extraction and correspondence matching are efficient. The user study is also designed to investigate the performance of Lecture2Note in facilitating learning. Compared with peer methods, the auto-generated note created by our method achieves a higher user satisfaction level regarding a properly structured layout as well as efficient content navigation and exploration. Chengpei Xu, Ruomei Wang 0001, Shujin Lin, Baoquan Zhao, Lijie Shao, Mengqiu Hu |
ICME | 2 |
| 2019 | A New Visual Interface for Searching and Navigating Slide-Based Lecture VideosabstractThe rapid development of distance education technologies, e.g. MOOCs, provide learners unprecedented access to high-quality online lecture videos at scale, anytime and anywhere. Unfortunately, these valuable resources are often underutilized by online learners. One prevailing reason is the lack of support for and the resulting difficulty of exploring and locating content of interest among lengthy recordings of course lectures. To address this deficiency, we introduce a novel visual interface that supports efficient search and navigation of video content at fine granularities and with rich semantic clues. The interface is particularly designed for slide-based lecture videos (SBLV), which represent a significant portion of online lecture videos. The interface comprehensively derives versatile semantic clues for video content indexing and visual aid generation according to visual elements, text, and mathematical expressions included on lecture slides, speeches recorded, as well as mouse and cursor pointing actions captured during a lecture. Empowered by such semantically revealing indices and visual assistance, the interface is able to noticeably enhance online learners' capabilities in searching and browsing of content in need from SBLVs. The advantages of the new interface are demonstrated through benchmarked experimental results in comparison with peer methods. Baoquan Zhao, Songhua Xu, Shujin Lin, Ruomei Wang 0001 |
ICME | 4 |
| 2019 | SFSegNet: Parse Freehand Sketches using Deep Fully Convolutional NetworksabstractParsing sketches via semantic segmentation is attractive but challenging, because (i) free-hand drawings are abstract with large variances in depicting objects due to different drawing styles and skills; (ii) distorting lines drawn on the touchpad make sketches more difficult to be recognized; (iii) the high-performance image segmentation via deep learning technologies needs enormous annotated sketch datasets during the training stage. In this paper, we propose a Sketch-target deep FCN Segmentation Network(SFSegNet) for automatic free-hand sketch segmentation, labeling each sketch in a single object with multiple parts. SFSegNet has an end-to-end network process between the input sketches and the segmentation results, composed of 2 parts: (i) a modified deep Fully Convolutional Network(FCN) using a reweighting strategy to ignore background pixels and classify which part each pixel belongs to; (ii) affine transform encoders that attempt to canonicalize the shaking strokes. We train our network with the dataset that consists of 10,000 annotated sketches, to find an extensively applicable model to segment stokes semantically in one ground truth. Extensive experiments are carried out and segmentation results show that our method outperforms other state-of-the-art networks. Junkun Jiang, Ruomei Wang 0001, Shujin Lin, Fei Wang 0056 |
IJCNN | 2 |
| 2019 | Rain Wiper: An Incremental Randomly Wired Network for Single Image DerainingabstractAbstract Single image rain removal is a challenging ill‐posed problem due to various shapes and densities of rain streaks. We present a novel incremental randomly wired network (IRWN) for single image deraining. Different from previous methods, most structures of modules in IRWN are generated by a stochastic network generator based on the random graph theory, which ease the burden of manual design and further help to characterize more complex rain streaks. To decrease network parameters and extract more details efficiently, the image pyramid is fused via the multi‐scale network structure. An incremental rectified loss is proposed to better remove rain streaks in different rain conditions and recover the texture information of target objects. Extensive experiments on synthetic and real‐world datasets demonstrate that the proposed method outperforms the state‐of‐the‐art methods significantly. In addition, an ablation study is conducted to illustrate the improvements obtained by different modules and loss items in IRWN. Xiangguo Liang, Bin Qiu, Zhuo Su 0001, Chengying Gao, X. Shi, Ruomei Wang 0001 |
Comput. Graph. Forum | 6 |
| 2019 | Learning mean progressive scattering using binomial truncated loss for image dehazingabstractIn this study, the authors propose a novel progressive dehazing network to address the single image haze removal problem based on a new mean progressive scattering model. Different from methods that learn atmosphere light and transmission maps with different networks, these two variables are optimised in a unified network. Following the methodology of traditional prior‐based methods that estimate a coarse transmission map first, a progressive refinement branch in the decoder has been designed to restore the fine‐scale transmission map. To improve the prediction accuracy of the transmission map, a novel binomial truncated loss that assigns weights to error values according to the probabilities of error occurrences has been proposed. An ablation study is conducted to verify the effectiveness of the components in the proposed method. Experiments in the synthetic datasets and real images demonstrate that the proposed method outperforms other state‐of‐the‐art methods. Bin Qiu, Xiwen Liang, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001 |
IET Image Process. | 4 |
| 2019 | Conditional progressive network for clothing parsingabstractClothing parsing is significant to many clothing applications. Recently, a lot of clothing parsing methods have been presented, which explore the innovation of the parsing pipeline or try to find more specific prior information. Although these methods perform well in some benchmarks, a few challenging problems have not been solved yet, such as the complicated mutual interference among labels. In this study, the authors propose a Conditional Progressive Network to parse clothing in different scales and prevent the mutual interference among labels. The authors’ solution consists of three sub‐networks, including Conditional Parsing Network (CPN), Pose Estimation Network (PEN) and Label Transform Network (LTN). Specifically, the CPN module generates the intermediate parsing result in the form of the multiple progressive stages, which combines with the previous outputs in each stage and the specific prior conditions. The PEN module provides a series of heat maps about the human pose information. The LTN module suppresses the redundant labels to avoid the mutual interference among labels. They demonstrate their solution in parsing the fashion clothing cases on the ATR and the Fashion dataset. In their experiments, their method obtains a better performance than the state‐of‐the‐art methods. Zhuo Su 0001, Jiaming Guo, Gengwei Zhang, Xianghui Luo, Ruomei Wang 0001, Fan Zhou 0001 |
IET Image Process. | 5 |
| 2019 | Parallel simulation model for heat and moisture transfer of clothed human body
Yuan Huang 0002, Jiapei Li, Haigang An, Xiaomin Jia, Ruomei Wang 0001 |
J. Supercomput. | 6 |
| 2018 | Learning deep similarity models with focus ranking for fabric image retrieval
Daiguo Deng, Ruomei Wang 0001, Hefeng Wu, Huayong He, Qi Li 0001 |
Image Vis. Comput. | 2 |
| 2017 | Multi-view pairwise relationship learning for sketch based 3D shape retrievalabstractRecent progress in sketch-based 3D shape retrieval creates a novel and user-friendly way to explore massive 3D shapes on the Internet. However, current methods on this topic rely on designing invariant features for both sketches and 3D shapes, or complex matching strategies. Therefore, they suffer from problems like arbitrary drawings and inconsistent viewpoints. To tackle this problem, we propose a probabilistic framework based on Multi-View Pairwise Relationship (MVPR) learning. Our framework includes multiple views of 3D shapes as the intermediate layer between sketches and 3D shapes, and transforms the original retrieval problem into the form of inferring pairwise relationship between sketches and views. We accomplish pairwise relationship inference by a novel MVPR net, which can automatically predict and merge the pairwise relationships between a sketch and multiple views, thus freeing us from exhaustively selecting the best view of 3D shapes. We also propose to learn robust features for sketches and views via fine-tuning pre-trained networks. Extensive experiments on a large dataset demonstrate that the proposed method can outperform state-of-the-art methods significantly. Hefeng Wu, Xiangjian He, Shujin Lin, Ruomei Wang 0001 |
ICME | 5 |
| 2017 | A Novel System for Visual Navigation of Educational Videos Using Multimodal CuesabstractWith recent developments and advances in distance learning and MOOCs, the amount of open educational videos on the Internet has grown dramatically in the past decade. However, most of these videos are lengthy and lack of high-quality indexing and annotations, which triggers an urgent demand for efficient and effective tools that facilitate video content navigation and exploration. In this paper, we propose a novel visual navigation system for exploring open educational videos. The system tightly integrates multimodal cues obtained from the visual, audio and textual channels of the video and presents them with a series of interactive visualization components. With the help of this system, users can explore the video content using multiple levels of details to identify content of interest with ease. Extensive experiments and comparisons against previous studies demonstrate the effectiveness of the proposed system. Baoquan Zhao, Shujin Lin, Songhua Xu, Ruomei Wang 0001 |
ACM Multimedia | 5 |
| 2017 | A Data-Driven Approach for Sketch-Based 3D Shape Retrieval via Similar Drawing-Style RecommendationabstractAbstract Sketching is a simple and natural way of expression and communication for humans. For this reason, it gains increasing popularity in human computer interaction, with the emergence of multitouch tablets and styluses. In recent years, sketch‐based interactive methods are widely used in many retrieval systems. In particular, a variety of sketch‐based 3D model retrieval works have been presented. However, almost all of these works focus on directly matching sketches with the projection views of 3D models, and they suffer from the large differences between the sketch drawing and the views of 3D models, leading to unsatisfying retrieval results. Therefore, in this paper, during the matching procedure in the retrieval, we propose to match the sketch with each 3D model from historical users instead of projection views. Yet since the sketches between the current user and the historical users can have big difference, we also aim to handle users' personalized deviations and differences. To this end, we leverage recommendation algorithms to estimate the drawing style characteristic similarity between the current user and historical users. Experimental results on the Large Scale Sketch Track Benchmark(SHREC14LSSTB) demonstrate that our method outperforms several state‐of‐the‐art methods. Fei Wang 0056, Shujin Lin, Hefeng Wu, Ruomei Wang 0001, Fan Zhou 0001 |
Comput. Graph. Forum | 5 |
| 2017 | A data-driven editing framework for automatic 3D garment modeling
Li Liu 0032, Zhuo Su 0001, Xiaodong Fu, Ruomei Wang 0001 |
Multim. Tools Appl. | 5 |
| 2017 | Distortion-Aware Correlation TrackingabstractRecently, correlation filter (CF)-based tracking methods have attracted considerable attention because of their high-speed performance. However, distortion, which refers to the phenomenon that the correlation outputs of CF-based trackers are distorted, remains a major obstacle for these methods. In this paper, we propose a distortion-aware correlation filter framework, which can detect distortions and recover from tracking failures. Our framework employs a simple yet effective feature termed normed correlation response to detect distortions. Meanwhile, we introduce a competition mechanism to handle distortions, in which we build a specialized graph to formulate and handle tracking under distortion as a maximum multi clique problem. Furthermore, a global-local context model is exploited to alleviate underlying distortions during the tracking process. Extensive experiments on the Online Tracking Benchmark show that our tracker can find the optimal target trajectory during the distortion period and retrieve the possibly missing target, consequently outperforms the state-of-the-art methods and improves the performance of CF-based trackers favorably. Hefeng Wu, Huifang Zhang, Shujin Lin, Ruomei Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2016 | A 3D model perceptual feature metric based on global height field
Yihui Guo, Shujin Lin, Zhuo Su 0001, Ruomei Wang 0001, Yang Kang |
Vis. Comput. | 5 |
| 2016 | Retiling scheme: a novel approach of direct anisotropic quad-dominant remeshing
Ruomei Wang 0001, Fan Zhou 0001 |
Vis. Comput. | 1 |
| 2014 | Mesh-based anisotropic cloth deformation for virtual fitting
Li Liu 0032, Ruomei Wang 0001, Zhuo Su 0001, Chengying Gao |
Multim. Tools Appl. | 2 |
| 2014 | An adaptive neural fuzzy network clothing comfort evaluation model and application in digital home
Ruomei Wang 0001, Fan Zhou 0001, Daiguo Deng |
Multim. Tools Appl. | 1 |
| 2013 | Material-aware cloth simulation via constrained geometric deformation
Li Liu 0032, Zhuo Su 0001, Ruomei Wang 0001 |
Comput. Graph. | 3 |
| 2012 | Image Resizing Based on Geometry Preservation with Seam CarvingabstractWhen an image or a video is transformed to an aspect ratio deferent from its original size, information lost is inevitable no matter what method is used, thus, how to keep the most attractive contents and minimize the visual distortion during the resizing process is the key issue. To address this problem, this paper proposes an object geometry preservation method based on the seam carving method. We first define a framework that measures the importance of geometry feature in the source material, then a new energy function is presented with object geometry constraint, according to the new energy function, an optimized seam carving method is used to minimize distortion while resizing the source material. The experiment results show that our method is better to transform a variety of source images to a different display size than conventional resizing methods. Fan Zhou 0001, Ruomei Wang 0001, Yun Liang 0003 |
TrustCom | 2 |
| 2011 | A multi-disciplinary strategy for computer-aided clothing thermal engineering design
Aihua Mao, Jie Luo 0009, Yi Li 0001, Ruomei Wang 0001 |
Comput. Aided Des. | 5 |
| 2008 | A CAD system for multi-style thermal functional design of clothing
Aihua Mao, Yi Li 0001, Ruomei Wang 0001, Shuxiao Wang |
Comput. Aided Des. | 4 |
| 2006 | P-smart - a virtual system for clothing thermal functional design
Yi Li 0001, Aihua Mao, Ruomei Wang 0001, Wenbang Hou, Liya Zhou, Yubei Lin |
Comput. Aided Des. | 3 |