EDBT 2026 Demo / reviewers in the wild / expert
Changbo Wang
dblp:87/7045
· DBLP profile ↗
154ranked-venue papers
16as first author
100since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 112 · 14 first-author · 72 since 2021Artificial intelligence and machine learning · 27 · 25 since 2021Human-computer interaction and ubiquitous computing · 19 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MPJudge: Towards Perceptual Assessment of Music-Induced PaintingsabstractMusic-induced painting is a unique artistic practice, where visual artworks are created under the influence of music. Evaluating whether a painting faithfully reflects the music that inspired it poses a challenging perceptual assessment task. Existing methods primarily rely on emotion recognition models to assess the similarity between music and painting, but such models introduce considerable noise and overlook broader perceptual cues beyond emotion. To address these limitations, we propose a novel framework for music-induced painting assessment that directly models perceptual coherence between music and visual art. We introduce MPD, the first large-scale dataset of music–painting pairs annotated by domain experts based on perceptual coherence. To better handle ambiguous cases, we further collect pairwise preference annotations. Building on this dataset, we present MPJudge, a model that integrates music features into a visual encoder via a modulation-based fusion mechanism. To effectively learn from ambiguous cases, we adopt Direct Preference Optimization for training. Extensive experiments demonstrate that our method outperforms existing approaches. Qualitative results further show that our model more accurately identifies music-relevant regions in paintings. Shiqi Jiang 0001, Tianyi Liang 0002, Huayuan Ye, Changbo Wang, Chenhui Li 0001 |
AAAI | 4 |
| 2026 | GT2-GS: Geometry-aware Texture Transfer for Gaussian SplattingabstractTransferring 2D textures onto complex 3D scenes plays a vital role in enhancing the efficiency and controllability of 3D multimedia content creation. However, existing 3D style transfer methods primarily focus on transferring abstract artistic styles to 3D scenes. These methods often overlook the geometric information of the scene, which makes it challenging to achieve high-quality 3D texture transfer results. In this paper, we present GT2-GS, a geometry-aware texture transfer framework for gaussian splatting. First, we propose a geometry-aware texture transfer loss that enables view-consistent texture transfer by leveraging prior view-dependent feature information and texture features augmented with additional geometric parameters. Moreover, an adaptive fine-grained control module is proposed to address the degradation of scene information caused by low-granularity texture features. Finally, a geometry preservation branch is introduced. This branch refines the geometric parameters using additionally bound Gaussian color priors, thereby decoupling the optimization objectives of appearance and geometry. Extensive experiments demonstrate the effectiveness and controllability of our method. Through geometric awareness, our approach achieves texture transfer results that better align with human visual perception. Zhongliang Liu, Junwei Shu, Changbo Wang |
AAAI | 4 |
| 2026 | KG-CPEN: Knowledge-Guided Compositional Prototype Evolution for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) serves as a pivotal bridge between Computer Vision and Natural Language Processing, aiming to parse unstructured imagery into structured semantic summaries. When dealing with long-tailed distributions, existing discriminative methods often suffer from severe data-dependency bias and ignore the compositional semantics of predicates, resulting in predictions inevitably collapsing into high-frequency head classes. To address this, we propose the Knowledge-Guided Compositional Prototype Evolution Network (KG-CPEN). To tackle the paucity of feature space caused by tail sample scarcity, we introduce a Knowledge-Guided Prototype Evolution mechanism, which utilizes external knowledge as a semantic engine to synthesize missing tail prototypes in the visual manifold, effectively replenishing feature representations in long-tailed and zero-shot scenarios. Furthermore, since traditional local greedy matching exacerbates head bias, we design an Optimal Transport Alignment module. This achieves unbiased matching by minimizing the distance between visual and prototype distributions on a global scale, thereby reducing decision layer difficulty and eliminating head class dominance. Extensive experiments on Visual Genome, GQA, and Open Images V6 consistently demonstrate that KG-CPEN establishes a new State-Of-The-Art (SOTA) for unbiased scene graph generation. Yujun Hu, Changbo Wang, Gaoqi He |
ICMR | 2 |
| 2026 | Enhancing trust through a human-center evaluation framework from an accessibility perspective: The case of graph anomaly detection
Yiding Shen, Juntong Chen, Feng Liu 0039, Chenhui Li 0001, Changbo Wang |
Int. J. Hum. Comput. Stud. | 6 |
| 2026 | MEDP: Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation
Ziheng Huang, Weiliang Meng, Changbo Wang, Gaoqi He |
Knowl. Based Syst. | 4 |
| 2026 | ChatTracker: Enhancing Visual Tracking via LLM-Driven Iterative Description RefinementabstractVisual object tracking focuses on locating a target object within a video sequence based on an initial bounding box. Recently, Vision-Language (VL) trackers have been proposed to utilize additional natural language descriptions to enhance versatility in various applications. Despite this potential, VL trackers still underperform the State-of-the-Art (SoTA) visual trackers in terms of tracking accuracy. We find that this inferiority is primarily due to their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we identify, for the first time, that over 10% of textual annotations in existing VL tracking datasets suffer from inaccuracies through manual evaluation. To address this problem, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel Reflection-based Language Description Refinement Module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed, which can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that ChatTracker achieves comparable performance to existing SoTA tracking methods. In addition, language descriptions generated by ChatTracker enhance the performance of various VL trackers and exhibit better text-to-image alignment than annotations in the original dataset. Moreover, our proposed framework can improve the performance of various visual tasks, including Referring Expression Comprehension (REC), Referring Expression Segmentation (RES), and Referring Video Object Segmentation (R-VOS) tasks by providing more accurate language descriptions, which demonstrates the universality of ChatTracker. We release the manual evaluation results and the generated textual descriptions, aiming to drive advancements in VL tracking. Yiming Sun 0006, Mi Zhang 0001, Shaoxiang Chen 0001, Yang Li 0041, Changbo Wang, Jianke Zhu, Steven C. H. Hoi |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | NewsVis: GenAI-Based Visual Storytelling for Corporate Financial NewsabstractCorporate financial news is pivotal for market decisions, but often overwhelms general audiences. While data videos effectively bridge this comprehension gap, their production remains a bottleneck for journalists. We present NewsVis, an authoring tool powered by Generative Artificial Intelligence (GenAI) that automates the transformation of unstructured narratives and raw financial datasets into professional data videos. Unlike generic models, our pipeline ensures factual accuracy through a domain-specific taxonomy of financial attributes and optimizes visual information presentation via a multimodal layout algorithm. Additionally, a human-in-the-loop interface empowers journalists to audit and calibrate generative outputs. Comprehensive quantitative and qualitative evaluations demonstrate that NewsVis significantly reduces production barriers while enhancing information accessibility for viewers. Jia Bu, Mingwei Jiang, Tong Lyu, Lumeng Wu, Shiqi Jiang 0001, Boyuan Huangfu, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2026 | RelMap: Reliable Spatiotemporal Sensor Data Visualization via Imputative Spatial InterpolationabstractAccurate and reliable visualization of spatiotemporal sensor data such as environmental parameters and meteorological conditions is crucial for informed decision-making. Traditional spatial interpolation methods, however, often fall short of producing reliable interpolation results due to the limited and irregular sensor coverage. This paper introduces a novel spatial interpolation pipeline that achieves reliable interpolation results and produces a novel heatmap representation with uncertainty information encoded. We leverage imputation reference data from Graph Neural Networks (GNNs) to enhance visualization reliability and temporal resolution. By integrating Principal Neighborhood Aggregation (PNA) and Geographical Positional Encoding (GPE), our model effectively learns the spatiotemporal dependencies. Furthermore, we propose an extrinsic, static visualization technique for interpolation-based heatmaps that effectively communicates the uncertainties arising from various sources in the interpolated map. Through a set of use cases, extensive evaluations on real-world datasets, and user studies, we demonstrate our model's superior performance for data imputation, the improvements to the interpolant with reference data, and the effectiveness of our visualization design in communicating uncertainties. Juntong Chen, Huayuan Ye, Siwei Fu, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | VizDefender: Unmasking Visualization Tampering Through Proactive Localization and Intent InferenceabstractThe integrity of data visualizations is increasingly threatened by image editing techniques that enable subtle yet deceptive tampering. Through a formative study, we define this challenge and categorize tampering techniques into two primary types: data manipulation and visual encoding manipulation. To address this, we present VizDefender, a framework for tampering detection and analysis. The framework integrates two core components: 1) a semi-fragile watermark module that protects the visualization by embedding a location map to images, which allows for the precise localization of tampered regions while preserving visual quality, and 2) an intent analysis module that leverages Multimodal Large Language Models (MLLMs) to interpret manipulation, inferring the attacker's intent and misleading effects. Extensive evaluations and user studies demonstrate the effectiveness of our methods. Sicheng Song, Zixin Chen, Huamin Qu, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D GenerationabstractText-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and a user study demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers. Tianyi Liang 0002, Haiwen Huang, Shiqi Jiang 0001, Yifei Huang 0006, Liangyu Chen 0001, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2026 | VisGuard: Securing Visualization Dissemination through Tamper-Resistant Data RetrievalabstractThe dissemination of visualizations is primarily in the form of raster images, which often results in the loss of critical information such as source code, interactive features, and metadata. While previous methods have proposed embedding metadata into images to facilitate Visualization Image Data Retrieval (VIDR), most existing methods lack practicability since they are fragile to common image tampering during online distribution such as cropping and editing. To address this issue, we propose VisGuard, a tamper-resistant VIDR framework that reliably embeds metadata link into visualization images. The embedded data link remains recoverable even after substantial tampering upon images. We propose several techniques to enhance robustness, including repetitive data tiling, invertible information broadcasting, and an anchor-based scheme for crop localization. VisGuard enables various applications, including interactive chart reconstruction, tampering detection, and copyright protection. We conduct comprehensive experiments on VisGuard's superior performance in data retrieval accuracy, embedding capacity, and security against tampering and steganalysis, demonstrating VisGuard's competence in facilitating and safeguarding visualization dissemination and information conveyance. Huayuan Ye, Juntong Chen, Shenzhuo Zhang, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationabstractRecent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions. However, exerting control over the motion of objects in videos generated by any video diffusion model remains a challenging problem. In this paper, we propose a novel zero-shot moving object trajectory control framework, Motion-Zero, to enable arbitrary single-object-trajectory control for the text-to-video diffusion model. To this end, an initial noise prior module is designed to provide a position-based prior to improve the stability of the appearance of the moving object and the accuracy of position. In addition, based on the attention map of the U-Net, spatial constraints are directly applied to the denoising process of diffusion models, which further ensures the positional consistency of moving objects during the inference. Furthermore, temporal consistency is guaranteed with a proposed shift temporal attention mechanism. Our method can be flexibly applied to various state-of-the-art video diffusion models without any training process. Extensive experiments demonstrate our proposed method can control the motion trajectories of arbitrary objects while preserving the original ability to generate high-quality videos. Changgu Chen, Junwei Shu, Gaoqi He, Changbo Wang, Yang Li 0041 |
AAAI | 4 |
| 2025 | Multi-granularity Feature Extraction Based on Long-Short Chains for Motion Retargeting
Weiliang Meng, Changbo Wang, Gaoqi He |
CGI (3) | 4 |
| 2025 | SandTouch: Empowering Virtual Sand Art in VR with AI Guidance and Emotional Relief
Junbin Ren, Zeyuan Fan, Chenhui Li 0001, Gaoqi He, Changbo Wang, Yang Gao 0025, Chen Li 0035 |
CHI | 6 |
| 2025 | Robust Message Embedding via Attention Flow-Based SteganographyabstractImage steganography can hide information in a host image and obtain a stego image that is perceptually indistinguishable from the original one. This technique has tremendous potential in scenarios like copyright protection and information retrospection. Some previous studies have proposed to enhance the robustness of the methods against image disturbances to increase their applicability. However, they generally cannot achieve a satisfying balance between the steganography quality and robustness. Instead of image-in-image steganography, we focus on the issue of message-in-image embedding that is robust to various real- world image distortions. This task aims to embed information into a natural image and the decoding result is required to be completely accurate, which increases the difficulty of data concealing and revealing. Inspired by the recent developments in transformer-based vision models, we discover that the tokenized representation of image is naturally suitable for steganography task. In this paper, we propose a novel message embedding framework, called Robust Message Steganography (RMSteg), which is competent to hide message via QR Code in a host image based on an normalizing flow-based model. The stego image derived by our method has imperceptible changes and the encoded message can be accurately restored even if the image is printed out and photographed. To our best knowledge, this is the first work that integrates the advantages of transformer models into normalizing flow. The code is available at https://github.com/huayuan4396/RMSteg. Huayuan Ye, Shenzhuo Zhang, Shiqi Jiang 0001, Jing Liao 0001, Shuhang Gu, Dejun Zheng, Changbo Wang, Chenhui Li 0001 |
CVPR | 7 |
| 2025 | Dynamic Stereotype Theory Induced Micro-expression Recognition with Oriented DeformationabstractMicro-expression recognition (MER) aims to uncover genuine emotions and underlying psychological states. However, existing MER methods struggle with three main challenges. 1) Scarcity of micro-expression samples. 2) Difficulty in modeling nearly imperceptible facial movements. 3) Reliance on apex frame annotations. To address these issues, we propose a Self-supervised Oriented Deformation model for Apex-free Micro-expression Recognition (SODA4MER). Our approach enhances local deformation perception using muscle-group priors and amplifies subtle features through Dynamic Stereotype Theory (DST) based enhancement, while contrastive learning eliminates the need for manual apex annotations. Specifically, the Oriented deformation estimator of SODA4MER is first pretrained in a self-supervised manner. Secondly, a Gated Temporal Variance Gaussian model (GTVG) is introduced to adaptively integrate facial muscle-group priors, enhancing local deformation perception and mitigating noise from head movements. Then, contrastive learning is employed to achieve apex detection by identifying the frame with the most significant local deformation. Finally, guided by DST, we introduced a feature enhancement strategy that models the temporal dynamics of local deformation in the activation and decay phases, leading to richer deformation features. Our rigorous experiments confirm the competitive performance and practical applicability of SODA4MER. Bohao Zhang, Changbo Wang, Gaoqi He |
CVPR | 3 |
| 2025 | TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image GenerationabstractText-to-image (T2I) generation has made remarkable progress in producing high-quality images, but a fundamental challenge remains: creating backgrounds that naturally accommodate text placement without compromising image quality.
This capability is non-trivial for real-world applications like graphic design, where clear visual hierarchy between content and text is essential.
Prior work has primarily focused on arranging layouts within existing static images, leaving unexplored the potential of T2I models for generating text-friendly backgrounds.
We present TextCenGen, a training-free approach that actively relocates objects before optimizing text regions, rather than directly reducing cross-attention which degrades image quality. Our method introduces: (1) a force-directed graph approach that detects conflicting objects and guides them relocation using cross-attention maps, and (2) a spatial attention constraint that ensures smooth background generation in text regions. Our method is plug-and-play, requiring no additional training while well balancing both semantic fidelity and visual quality.
Evaluated on our proposed text-friendly T2I benchmark of 27,000 images across three seed datasets, TextCenGen outperforms existing methods by achieving 23\% lower saliency overlap in text regions while maintaining 98\% of the original semantic fidelity measured by CLIP score and our proposed Visual-Textual Concordance Metric (VTCM). Tianyi Liang 0002, Jiangqi Liu, Yifei Huang 0006, Shiqi Jiang 0001, Jianshen Shi, Changbo Wang, Chenhui Li 0001 |
ICML | 6 |
| 2025 | DPSN: Dual Prior Knowledge Induced Tactile paving and Obstacle Joint Segmentation NetworkabstractAccurate semantic segmentation of both tactile paving and the obstacle is crucial for the safe mobility of visually impaired individuals. However, existing methods face two major challenges: (i) discontinuous segmentation fragments; (ii) Inaccurate obstacle recognition. To address challenge (i), we propose incorporating appearance priors of complete tactile pavings to prevent the model from directly learning irregular ground truth masks. To tackle challenge (ii), we propose introducing cross-modal semantic priors to complement the semantic information of obstacles. We implemented these strategies in proposed Dual Prior knowledge induced tactile paving and obstacle joint Segmentation Network (DPSN). Based on bilateral network architecture, DPSN merges obstacle category masks into tactile paving categories, constructing a complete tactile paving mask. Utilizing the complete mask, DPSN transfer appearance prior knowledge to detail features from boundary and structural perspectives. Concurrently, DPSN leverages the CLIP Text Encoder to guide visual feature decoding by attention mechanisms, transferring rich cross-modal semantic prior knowledge to the visual feature maps. Furthermore, we propose the TPO-Dataset, the first dataset for joint tactile paving and obstacle segmentation acquired from actual scenes. Experiments demonstrate that DPSN achieves state-of-the-art results on the TPO-Dataset, with relative gains of 27.16% in obstacle IoU and 30.53% in accuracy metrics compared to baseline methods. Notably, DPSN achieves real-time performance at 88.25 FPS on the maximum scale of 2048×512 resolution. Youqi Song, Zilong Jin, Changbo Wang, Gaoqi He |
IROS | 6 |
| 2025 | 3D Scene Graph Generation with Cross-Modal Alignment and Adversarial Learning
Yujun Hu, Changbo Wang, Weiliang Meng, Gaoqi He |
ICMR | 3 |
| 2025 | Mitigating Long-tail Distribution in Oracle Bone Inscriptions: Dataset, Model, and BenchmarkabstractThe oracle bone inscription (OBI) recognition plays a significant role in understanding the history and culture of ancient China. However, the existing OBI datasets suffer from a long-tail distribution problem, leading to biased performance of OBI recognition models across majority and minority classes. With recent advancements in generative models, OBI synthesis-based data augmentation has become a promising avenue to expand the sample size of minority classes. Unfortunately, current OBI datasets lack large-scale structure-aligned image pairs for generative model training. To address these problems, we first present the Oracle-P15K, a structure-aligned OBI dataset for OBI generation and denoising, consisting of 14,542 images infused with domain knowledge from OBI experts. Second, we propose a diffusion model-based pseudo OBI generator, called OBIDiff, to achieve realistic and controllable OBI generation. Given a clean glyph image and a target rubbing-style image, it can effectively transfer the noise style of the original rubbing to the glyph image. Extensive experiments on OBI downstream tasks and user preference studies show the effectiveness of the proposed Oracle-P15K dataset and demonstrate that OBIDiff can accurately preserve inherent glyph structures while transferring authentic rubbing styles effectively. The dataset, code, and pre-trained models are available at https://github.com/LJHolyGround/Oracle-P15K. Jinhao Li 0001, Zijian Chen 0001, Runze Jiang, Tingzhu Chen, Changbo Wang, Guangtao Zhai |
ACM Multimedia | 5 |
| 2025 | PPJudge: Towards Human-Aligned Assessment of Artistic Painting ProcessabstractArtistic image assessment has become a prominent research area in computer vision. In recent years, the field has witnessed a proliferation of datasets and methods designed to evaluate the aesthetic quality of paintings. However, most existing approaches focus solely on static final images, overlooking the dynamic and multi-stage nature of the artistic painting process. To address this gap, we propose a novel framework for human-aligned assessment of painting processes. Specifically, we introduce the Painting Process Assessment Dataset (PPAD)-the first large-scale dataset comprising real and synthetic painting process images, annotated by domain experts across eight detailed attributes. Furthermore, we present PPJudge (Painting Process Judge), a Transformer-based model enhanced with temporally-aware positional encoding and a heterogeneous mixture-of-experts architecture, enabling effective assessment of the painting process. Experimental results demonstrate that our method outperforms existing baselines in accuracy, robustness, and alignment with human judgment, offering new insights into computational creativity and art education. Shiqi Jiang 0001, Xinpeng Li 0002, Xi Mao, Changbo Wang, Chenhui Li 0001 |
ACM Multimedia | 4 |
| 2025 | Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation LearningabstractEmotion alignment between music and palettes is crucial for effective multimedia content, yet misalignment creates confusion that weakens the intended message. However, existing methods often generate only a single dominant color, missing emotion variation. Others rely on indirect mappings through text or images, resulting in the loss of crucial emotion details. To address these challenges, we present Music2Palette, a novel method for emotion-aligned color palette generation via cross-modal representation learning. We first construct MuCED, a dataset of 2,634 expert-validated music-palette pairs aligned through Russell-based emotion vectors. To directly translate music into palettes, we propose a cross-modal representation learning framework with a music encoder and color decoder. We further propose a multi-objective optimization approach that jointly enhances emotion alignment, color diversity, and palette coherence. Extensive experiments demonstrate that our method outperforms current methods in interpreting music emotion and generating attractive and diverse color palettes. Our approach enables applications like music-driven image recoloring, video generating, and data visualization, bridging the gap between auditory and visual emotion experiences. Jiayun Hu, Yueyi He, Tianyi Liang 0002, Changbo Wang, Chenhui Li 0001 |
ACM Multimedia | 4 |
| 2025 | FluidGS: Physics Informed Gaussian Splatting for Dynamic Fluid Reconstruction from Sparse Views
Youchen Xie, Chen Li 0035, Sheng Qiu, Zhi-Jun Wang, Chenhui Li 0001, Yibo Zhao 0001, Zan Gao 0001, Changbo Wang |
ACM Multimedia | 8 |
| 2025 | TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary GenerationabstractSoccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) show promising capabilities in temporal grounding and video understanding. However, generating soccer commentary requires both precise temporal localization and semantically rich descriptions over long-form videos. Existing soccer MLLMs often rely on temporal priors for caption generation, which limits their ability to process the entire video in an end-to-end manner. Traditional approaches, on the other hand, follow a complex two-step paradigm that fails to capture the global context, leading to suboptimal performance. To solve the above issues, we present TimeSoccer, the first end-to-end soccer MLLM for Single-anchor Dense Video Captioning (SDVC) in full-match soccer videos. TimeSoccer jointly predicts timestamps and generates captions in a single pass, enabling global context modeling across 45-minute matches. To support long video understanding of soccer matches, we introduce MoFA-Select, a training-free, motion-aware frame compression module that adaptively selects representative frames via a coarse-to-fine strategy, and incorporates complementary training paradigms to strengthen the model's ability to handle long temporal sequences. Extensive experiments demonstrate that our TimeSoccer achieves State-of-The-Art (SoTA) performance on the SDVC task in an end-to-end form, generating high-quality commentary with accurate temporal alignment and strong semantic relevance. For more information, please visit: https://vpx-ecnu.github.io/TimeSoccer-Website/. Ling You, Wenxuan Huang 0001, Xinni Xie, Xiangyi Wei, Bangyan Li, Shaohui Lin, Yang Li 0041, Changbo Wang |
ACM Multimedia | 8 |
| 2025 | Regulatory Focus Theory Induced Micro-Expression Analysis with Structured Representation Learning
Bohao Zhang, Haoxin Xu, Jingzhong Lin, Changbo Wang, Gaoqi He |
ACM Multimedia | 4 |
| 2025 | TactPav: A Vision-Language Annotated Multi-modal Dataset for Tactile Paving Navigation
Youqi Song, Zilong Jin, Yunjie Xie, Changbo Wang, Gaoqi He |
PRCV (12) | 7 |
| 2025 | Detail-preserving shape completion of point cloud models with articulated structure
Yi Quan, Chen Li 0035, Yang Li 0041, Changbo Wang, Hong Qin 0001 |
Comput. Aided Geom. Des. | 4 |
| 2025 | Prompt2Color: A prompt-based framework for image-derived color generation and visualization optimization
Jiayun Hu, Shiqi Jiang 0001, Haiwen Huang, Changbo Wang, Chenhui Li 0001 |
Comput. Graph. | 6 |
| 2025 | CeRF: Convolutional neural radiance derivative fields for new view synthesis
Ling You, Dingbo Lu, Yang Li 0041, Changbo Wang |
Comput. Graph. | 6 |
| 2025 | SUPQA: LLM-based Geo-Visualization for Subjective Urban Performance Question-AnsweringabstractAbstract As urbanization accelerates, urban performance has become a growing concern, impacting every aspect of residents' lives. However, urban performance exploration is a tedious and highly subjective process for users. Users need to manually collect and integrate various information, or spend a large amount of time and effort due to the steep learning curves of existing specialized tools. To address these challenges, we introduce SUPQA, a novel approach for urban performance exploration using natural language as input and interactive geographic visualizations as output. Our approach leverages Large Language Models (LLMs) to effectively interpret user intents and quantify various urban performance measures. We integrate progressive navigation and multi‐geographic scale analysis in our visualization system, explaining the reasoning process and streamlining users' decision‐making workflow. Two usage scenarios and evaluations demonstrate the effectiveness of SUPQA in helping residents and planners acquire desired information more efficiently and enhancing the quality of decision‐making. Haiwen Huang, Juntong Chen, Changbo Wang, Chenhui Li 0001 |
Comput. Graph. Forum | 3 |
| 2025 | EmoDiffGes: Emotion-Aware Co-Speech Holistic Gesture Generation with Progressive Synergistic DiffusionabstractAbstract Co‐speech gesture generation, driven by emotional expression and synergistic bodily movements, is essential for applications such as virtual avatars and human‐robot interaction. Existing co‐speech gesture generation methods face two fundamental limitations: (1) producing inexpressive gestures due to ignoring the temporal evolution of emotion; (2) generating incoherent and unnatural motions as a result of either holistic body oversimplification or independent part modeling. To address the above limitations, we propose EmoDiffGes, a diffusion‐based framework grounded in embodied emotion theory, unifying dynamic emotion conditioning and part‐aware synergistic modeling. Specifically, a Dynamic Emotion‐Alignment Module (DEAM) is first applied to extract dynamic emotional cues and inject emotion guidance into the generation process. Then, a Progressive Synergistic Gesture Generator (PSGG) iteratively refines region‐specific latent codes while maintaining full‐body coordination, leveraging a Body Region Prior for part‐specific encoding and Progressive Inter‐Region Synergistic Flow for global motion coherence. Extensive experiments validate the effectiveness of our methods, showcasing the potential for generating expressive, coordinated, and emotionally grounded human gestures. Jingzhong Lin, Bohao Zhang, Changbo Wang, Gaoqi He |
Comput. Graph. Forum | 5 |
| 2025 | DC-APIC: A decomposed compatible affine particle in cell transfer scheme for non-sticky solid-fluid interactions in MPMabstractDespite the material point method (MPM) provides a unified particle simulation framework for coupling of different materials, MPM suffers from sticky numerical artifacts, which is inherently restricted to sticky and no-slip interactions. In this paper, we propose a novel transfer scheme called Decomposed Compatible Affine Particle in Cell (DC-APIC) within the MPM framework for simulating the two-way coupled interaction between elastic solids and incompressible fluids under free-slip boundary conditions on a unified background grid. Firstly, we adopt particle-grid compatibility to describe the relationship between grid nodes and particles at the fluid–solid interface, which serves as the guideline for subsequent particle–grid–particle transfers. Then we develop a phase-field gradient method to track the compatibility and normal directions at the interface. Secondly, to facilitate automatic MPM collision resolution during solid–fluid coupling, in the proposed DC-APIC integrator, the tangential component will not be transferred between incompatible grid nodes to prevent velocity smoothing in another phase, while the normal component is transferred without limitations. Finally, our comprehensive results confirm that our approach effectively reduces diffusion and unphysical viscosity compared to traditional MPM. • Developed a decomposed compatible APIC transfer scheme to reduce numerical viscosity on a unified grid. • Modified traditional MPM with DC-APIC integrator to enforce free-slip and separation boundary conditions. • Created an efficient parallel framework utilizing hierarchical GPU architecture for phase-field gradient method. Jianyang Zhang, Chen Li 0035, Changbo Wang |
Graph. Model. | 4 |
| 2025 | DCS-RISR: Dynamic channel splitting for efficient real-world image super-resolution
Junbo Qiao, Shaohui Lin, Yulun Zhang 0001, Wei Li 0002, Jie Hu 0021, Gaoqi He, Changbo Wang, Lizhuang Ma |
Neural Networks | 7 |
| 2025 | GVVST: Image-Driven Style Extraction From Graph Visualizations for Visual Style TransferabstractIncorporating automatic style extraction and transfer from existing well-designed graph visualizations can significantly alleviate the designer's workload. There are many types of graph visualizations. In this paper, our work focuses on node-link diagrams. We present a novel approach to streamline the design process of graph visualizations by automatically extracting visual styles from well-designed examples and applying them to other graphs. Our formative study identifies the key styles that designers consider when crafting visualizations, categorizing them into global and local styles. Leveraging deep learning techniques such as saliency detection models and multi-label classification models, we develop end-to-end pipelines for extracting both global and local styles. Global styles focus on aspects such as color scheme and layout, while local styles are concerned with the finer details of node and edge representations. Through a user study and evaluation experiment, we demonstrate the efficacy and time-saving benefits of our method, highlighting its potential to enhance the graph visualization design process. Sicheng Song, Yanna Lin, Huamin Qu, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Data augmentation with attention framework for robust deepfake detection
Sardor Mamarasulov, Lianggangxu Chen, Changgu Chen, Changbo Wang |
Vis. Comput. | 5 |
| 2024 | Multi-Prototype Space Learning for Commonsense-Based Scene Graph GenerationabstractIn the domain of scene graph generation, modeling commonsense as a single-prototype representation has been typically employed to facilitate the recognition of infrequent predicates. However, a fundamental challenge lies in the large intra-class variations of the visual appearance of predicates, resulting in subclasses within a predicate class. Such a challenge typically leads to the problem of misclassifying diverse predicates due to the rough predicate space clustering. In this paper, inspired by cognitive science, we maintain multi-prototype representations for each predicate class, which can accurately find the multiple class centers of the predicate space. Technically, we propose a novel multi-prototype learning framework consisting of three main steps: prototype-predicate matching, prototype updating, and prototype space optimization. We first design a triple-level optimal transport to match each predicate feature within the same class to a specific prototype. In addition, the prototypes are updated using momentum updating to find the class centers according to the matching results. Finally, we enhance the inter-class separability of the prototype space through iterations of the inter-class separability loss and intra-class compactness loss. Extensive evaluations demonstrate that our approach significantly outperforms state-of-the-art methods on the Visual Genome dataset. Lianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu, Yang Li 0041, Changbo Wang, Gaoqi He |
AAAI | 7 |
| 2024 | Kumaraswamy Wavelet for Heterophilic Scene Graph GenerationabstractGraph neural networks (GNNs) has demonstrated its capabilities in the field of scene graph generation (SGG) by updating node representations from neighboring nodes. Actually it can be viewed as a form of low-pass filter in the spatial domain, which smooths node feature representation and retains commonalities among nodes. However, spatial GNNs does not work well in the case of heterophilic SGG in which fine-grained predicates are always connected to a large number of coarse-grained predicates. Blind smoothing undermines the discriminative information of the fine-grained predicates, resulting in failure to predict them accurately. To address the heterophily, our key idea is to design tailored filters by wavelet transform from the spectral domain. First, we prove rigorously that when the heterophily on the scene graph increases, the spectral energy gradually shifts towards the high-frequency part. Inspired by this observation, we subsequently propose the Kumaraswamy Wavelet Graph Neural Network (KWGNN). KWGNN leverages complementary multi-group Kumaraswamy wavelets to cover all frequency bands. Finally, KWGNN adaptively generates band-pass filters and then integrates the filtering results to better accommodate varying levels of smoothness on the graph. Comprehensive experiments on the Visual Genome and Open Images datasets show that our method achieves state-of-the-art performance. Lianggangxu Chen, Youqi Song, Shaohui Lin, Changbo Wang, Gaoqi He |
AAAI | 4 |
| 2024 | AACP: Aesthetics Assessment of Children's Paintings Based on Self-Supervised LearningabstractThe Aesthetics Assessment of Children's Paintings (AACP) is an important branch of the image aesthetics assessment (IAA), playing a significant role in children's education. This task presents unique challenges, such as limited available data and the requirement for evaluation metrics from multiple perspectives. However, previous approaches have relied on training large datasets and subsequently providing an aesthetics score to the image, which is not applicable to AACP. To solve this problem, we construct an aesthetics assessment dataset of children's paintings and a model based on self-supervised learning. 1) We build a novel dataset composed of two parts: the first part contains more than 20k unlabeled images of children's paintings; the second part contains 1.2k images of children's paintings, and each image contains eight attributes labeled by multiple design experts. 2) We design a pipeline that includes a feature extraction module, perception modules and a disentangled evaluation module. 3) We conduct both qualitative and quantitative experiments to compare our model's performance with five other methods using the AACP dataset. Our experiments reveal that our method can accurately capture aesthetic features and achieve state-of-the-art performance. Shiqi Jiang 0001, Changbo Wang, Chenhui Li 0001 |
AAAI | 5 |
| 2024 | SalienTime: User-driven Selection of Salient Time Steps for Large-Scale Geospatial Data VisualizationabstractThe voluminous nature of geospatial temporal data from physical monitors and simulation models poses challenges to efficient data access, often resulting in cumbersome temporal selection experiences in web-based data portals. Thus, selecting a subset of time steps for prioritized visualization and pre-loading is highly desirable. Addressing this issue, this paper establishes a multifaceted definition of salient time steps via extensive need-finding studies with domain experts to understand their workflows. Building on this, we propose a novel approach that leverages autoencoders and dynamic programming to facilitate user-driven temporal selections. Structural features, statistical variations, and distance penalties are incorporated to make more flexible selections. User-specified priorities, spatial regions, and aggregations are used to combine different perspectives. We design and implement a web-based interface to enable efficient and context-aware selection of time steps and evaluate its efficacy and usability through case studies, quantitative evaluations, and expert interviews. Juntong Chen, Haiwen Huang, Huayuan Ye, Zhong Peng, Chenhui Li 0001, Changbo Wang |
CHI | 6 |
| 2024 | DoodleTunes: Interactive Visual Analysis of Music-Inspired Children Doodles with Automated Feature AnnotationabstractMusic and visual arts are essential in children’s arts education, and their integration has garnered significant attention. Existing data analysis methods for exploring audio-visual correlations are limited. Yet, relevant research is necessary for innovating and promoting arts integration courses. In our work, we collected substantial volumes of music-inspired doodles created by children and interviewed education experts to comprehend the challenges they encountered in the relevant analysis. Based on the insights we obtained, we designed and constructed an interactive visualization system DoodleTunes. DoodleTunes integrates deep learning-driven methods for automatically annotating several types of data features. The visual designs of the system are based on a four-level analysis structure to construct a progressive workflow, facilitating data exploration and insight discovery between doodle images and corresponding music pieces. We evaluated the accuracy of our feature prediction results and collected usage feedback on DoodleTunes from five domain experts. Jia Bu, Huayuan Ye, Juntong Chen, Shiqi Jiang 0001, Mingtian Tao, Changbo Wang, Chenhui Li 0001 |
CHI | 8 |
| 2024 | CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive Learningabstract3D Scene Graph Generation (3DSGG) aims to classify objects and their predicates within 3D point cloud scenes. However, current 3DSGG methods struggle with two main challenges. 1) The dependency on labor-intensive ground-truth annotations. 2) Closed-set classes training hampers the recognition of novel objects and predicates. Addressing these issues, our idea is to extract cross-modality features by CLIP from text and image data naturally related to 3D point clouds. Cross-modality features are used to train a robust 3D scene graph (3DSG)feature extractor. Specifically, we propose a novel Cross-Modality Contrastive Learning 3DSGG (CCL-3DSGG) method. Firstly, to align the text with 3DSG, the text is parsed into word level that are consistent with the 3DSG annotation. To enhance robustness during the alignment, adjectives are exchanged for different objects as negative samples. Then, to align the image with 3DSG, the camera view is treated as a positive sample and other views as negatives. Lastly, the recognition of novel object and predicate classes is achieved by calculating the cosine similarity between prompts and 3DSG features. Our rigorous experiments confirm the superior open-vocabulary capability and applicability of CCL-3DSGG in real-world contexts. Lianggangxu Chen, Jiale Lu, Shaohui Lin, Changbo Wang, Gaoqi He |
CVPR | 5 |
| 2024 | Shadow Constrained DEM Refinement Based on Differentiable RenderingabstractDigital elevation models (DEMs) are the fundamental for modeling and analyzing spatial topographic information in geographic information system, 3D video games, and many other fields. However, due to various terrain factors in data acquisition, open access datasets often contain inaccurate data or miss data, leading to undesirable models. This paper proposes a terrain refinement method based on shadow constraints by taking full advantages of differentiable rendering enabled efficient optimization. To be specific, we introduce an iterative approach to optimize shadow masks from satellite images based on differentiable rendering, which provides extra geometric clues for further terrain refinement. Thereafter, we propose to synthesize high-quality data in a randomization manner via differentiable renderer to expose the latent correlation between shadow distribution and terrain geometry, and generalize to real-world DEMs. Moreover, structure lines extracted from forward rendering results are also utilized to provide comprehensive geometric constraints for terrains. Extensive experiments demonstrate the effectiveness of our proposed methods. Fan Tian, Peichi Zhou, Chen Li 0035, Changbo Wang |
ICME | 4 |
| 2024 | Summarizing Charts of Financial Document via Context-Aware Multi-ModelingabstractIn the field of financial analysis, investment research analysts depend on a detailed understanding of complex financial documents to guide their decision-making process. Charts, while providing visual insights into data, present challenges in summarization. To address this issue, we present a novel approach that leverages contextual awareness, both in terms of textual semantics and visual perception. Our method begins with object detection technology to accurately locate and identify charts. Subsequently, a pre-trained language model is employed for vectorizing text and chart captions, enabling effective correlation between charts and their textual descriptions. Utilizing a large language model and strategic prompt engineering, we generate concise yet informative chart summaries, and incorporate visual saliency to assign scores, quantifying the importance of each chart for more effective data interpretation. Our study, supported by dedicated datasets, validates efficiency and accuracy improvements in financial analysis, expediting well-informed investment decisions. Xiaoyue Huang, Yaxuan Zheng, Xiping Wang, Yanpeng Hu, Changbo Wang, Chenhui Li 0001 |
IJCNN | 5 |
| 2024 | HeRF: A Hierarchical Framework for Efficient and Extendable New View SynthesisabstractRecently, neural radiance fields have made significant advancements in rendering new views. However, limited research has focused on dynamically loading implicit radiance fields with efficient memory utilization and extended scene representation. This paper introduces HeRF, a novel framework with a hierarchical scene representation based on layered sparse voxels. With such an adaptive design, our method is able to partition scenes into different levels for faster modeling and reduced memory cost. Furthermore, these partitioned scenes can be dynamically loaded and joined for a better immersive experience. Quantitative and qualitative analysis using objectlevel, indoor, and outdoor datasets demonstrates the effectiveness of HeRF. Remarkably, our proposed method requires only about 38% of the training rays and 45% of the GPU memory cost, yet achieves a 9% improvement in PSNR compared to NeRFusion on the ScanNet dataset. The code is available at https://github.com/Minisal/HeRF. Dingbo Lu, Ling You, Yang Li 0041, Changbo Wang |
IJCNN | 6 |
| 2024 | Saliency-Aware Projection Usability Enhancement for Dimensionality Reduction through Generative ModelsabstractDimensionality reduction (DR), also known as projection, is one of the most commonly used methods for visualizing high-dimensional data. Despite its effectiveness in handling large datasets with high dimensions, users often face the challenge of tuning the parameters for optimal performance. Additionally, due to the lack of intuitive standards, users often struggle to quickly identify satisfactory results from the vast number of possible outcomes. Therefore, enhancing the usability of DR algorithms is an urgent problem that needs to be addressed. In this paper, we present a method based on generative models aimed at circumventing the parameter tuning process for DR. Furthermore, to provide users with valid recommendations, we introduce mixed quality metrics based on visual saliency for visualizing DR results. These quality metrics are mapped to a continuous latent space constructed by the generative model using interpolation. We demonstrate the validity and effectiveness of our method through a series of quantitative experiments. Subsequently, we develop a visual interface that combines the proposed method and metrics. The evaluation results demonstrate that our method can quickly recommend good DR results, leading to a more user-friendly and efficient visualization analysis experience. Yaxuan Zheng, Wenli Xiong, Changbo Wang, Chenhui Li 0001 |
IJCNN | 3 |
| 2024 | Channel Robust Strategies with Data Augmentation for Audio Anti-spoofing
Sardor Mamarasulov, Yang Li 0041, Changbo Wang |
ISC (2) | 3 |
| 2024 | Generative Data Augmentation with Liveness Information Preserving for Face Anti-SpoofingabstractFace anti-spoofing is a critical aspect of ensuring security in the context of human-robot interaction and collaboration. Recently, disentangled-based data augmentation methods have achieved great success in face anti-spoofing tasks. The underlying assumption of those methods is that the liveness information could be completely disentangled and the labeling of the augmented data could totally depend on the liveness-related feature branch. However, we observe that it is almost impossible to extract the liveness-related information completely, which makes the current labeling strategy inaccurate. In this paper, we rethink the disentangling process and propose a novel generative-based data augmentation framework without forcing liveness information encoded into any specific feature space. Specifically, the original images are decomposed into statistic feature space and spatial feature space with liveness information preserving. With these two feature spaces, synthesized liveness-preserving images are generated with the Cartesian product to further approach the distribution of real face anti-spoofing data. Along with the original samplings, the augmented data are fed to a ResNet-based classifier with our proposed pseudo-label strategy for liveness information augmentation. Both qualitative and quantitative experiments demonstrate promising results to show the effectiveness of our proposed method. Changgu Chen, Yang Li 0041, Jian Zhang 0079, Changbo Wang |
ICMR | 5 |
| 2024 | FIND: Fine-tuning Initial Noise Distribution with Policy Optimization for Diffusion ModelsabstractIn recent years, large-scale pre-trained diffusion models have demonstrated their outstanding capabilities in image and video generation tasks. However, existing models tend to produce visual objects commonly found in the training dataset, which diverges from user input prompts. The underlying reason behind the inaccurate generated results lies in the model's difficulty in sampling from specific intervals of the initial noise distribution corresponding to the prompt. Moreover, it is challenging to directly optimize the initial distribution, given that the diffusion process involves multiple denoising steps. In this paper, we introduce a Fine-tuning Initial Noise Distribution (FIND) framework with policy optimization, which unleashes the powerful potential of pre-trained diffusion networks by directly optimizing the initial distribution to align the generated contents with user-input prompts. To this end, we first reformulate the diffusion denoising procedure as a one-step Markov decision process and employ policy optimization to directly optimize the initial distribution. In addition, a dynamic reward calibration module is proposed to ensure training stability during optimization. Furthermore, we introduce a ratio clipping algorithm to utilize historical data for network training and prevent the optimized distribution from deviating too far from the original policy to restrain excessive optimization magnitudes. Extensive experiments demonstrate the effectiveness of our method in both text-to-image and text-to-video tasks, surpassing SOTA methods in achieving consistency between prompts and the generated content. Our method achieves 10 times faster than the SOTA approach. Changgu Chen, Libing Yang, Lianggangxu Chen, Gaoqi He, Changbo Wang, Yang Li 0041 |
ACM Multimedia | 6 |
| 2024 | ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelabstractVisual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility in various applications. However, VL trackers are still inferior to State-of-The-Art (SoTA) visual trackers in terms of tracking performance. We found that this inferiority primarily results from their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel reflection-based prompt optimization module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed and can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that our proposed ChatTracker achieves a performance comparable to existing methods. Yiming Sun 0006, Shaoxiang Chen 0001, Junwei Huang, Yang Li 0041, Chenhui Li 0001, Changbo Wang |
NeurIPS | 8 |
| 2024 | Prototype-based contrastive substructure identification for molecular property predictionabstractSubstructure-based representation learning has emerged as a powerful approach to featurize complex attributed graphs, with promising results in molecular property prediction (MPP). However, existing MPP methods mainly rely on manually defined rules to extract substructures. It remains an open challenge to adaptively identify meaningful substructures from numerous molecular graphs to accommodate MPP tasks. To this end, this paper proposes Prototype-based cOntrastive Substructure IdentificaTion (POSIT), a self-supervised framework to autonomously discover substructural prototypes across graphs so as to guide end-to-end molecular fragmentation. During pre-training, POSIT emphasizes two key aspects of substructure identification: firstly, it imposes a soft connectivity constraint to encourage the generation of topologically meaningful substructures; secondly, it aligns resultant substructures with derived prototypes through a prototype-substructure contrastive clustering objective, ensuring attribute-based similarity within clusters. In the fine-tuning stage, a cross-scale attention mechanism is designed to integrate substructure-level information to enhance molecular representations. The effectiveness of the POSIT framework is demonstrated by experimental results from diverse real-world datasets, covering both classification and regression tasks. Moreover, visualization analysis validates the consistency of chemical priors with identified substructures. The source code is publicly available at https://github.com/VRPharmer/POSIT. Gaoqi He, Changbo Wang, Kai Zhang 0001, Honglin Li 0003 |
Briefings Bioinform. | 4 |
| 2024 | Improving rare relation inferring for scene graph generation using bipartite graph network
Jiale Lu, Lianggangxu Chen, Haoyue Guan, Shaohui Lin, Chunhua Gu, Changbo Wang, Gaoqi He |
Comput. Vis. Image Underst. | 6 |
| 2024 | A novel transformer-based graph generation model for vectorized road designabstractAbstract Road network design, as an important part of landscape modeling, shows a great significance in automatic driving, video game development, and disaster simulation. To date, this task remains labor‐intensive, tedious and time‐consuming. Many improved techniques have been proposed during the last two decades. Nevertheless, most of the state‐of‐the‐art methods still encounter problems of intuitiveness, usefulness and/or interactivity. As a rapid deviation from the conventional road design, this paper advocates an improved road modeling framework for automatic and interactive road production driven by geographical maps (including elevation, water, vegetation maps). Our method integrates the capability of flexible image generation models with powerful transformer architecture to afford a vectorized road network. We firstly construct a dataset that includes road graphs, density map and their corresponding geographical maps. Secondly, we develop a density map generation network based on image translation model with an attention mechanism to predict a road density map. The usage of density map facilitates faster convergence and better performance, which also serves as the input for road graph generation. Thirdly, we employ the transformer architecture to evolve density maps to road graphs. Our comprehensive experimental results have verified the efficiency, robustness and applicability of our newly‐proposed framework for road design. Peichi Zhou, Chen Li 0035, Jian Zhang 0070, Changbo Wang, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2024 | Toward Efficient Hyperspectral Anomaly Detection With Subspace Transformation LearningabstractCurrent research in hyperspectral anomaly detection often incorporates low-rank (LR) or total variation (TV) priors to encode the background matrix. However, applying such regularizers to the detection model increases the computational burden. In this letter, we propose a subspace transformation learning-based anomaly detector (termed STLAD). In STLAD, we employ an orthogonal transformation to represent the background in its subspace, where both the background and the transformation share spatial smoothness prior and approximate sparsity properties based on carefully selected basis vectors. By leveraging this background characterization, the anomaly component can be effectively described using the ℓ2,1 mixed norm. To solve the STLAD model, we design an alternating direction method of multipliers (ADMM) with guaranteed convergence. Experiments conducted on benchmark hyperspectral datasets demonstrate that STLAD outperforms several state-of-the-art anomaly detection methods. The demo of STLAD will be publicly available at: https://github.com/XiangfeiShen/STLAD. Changbo Wang, Laihang Yu, Jian Zhang 0118, Xiangfei Shen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | A Unified MPM Framework Supporting Phase-field Models and Elastic-viscoplastic Phase TransitionabstractRecent years have witnessed the rapid deployment of numerous physics-based modeling and simulation algorithms and techniques for fluids, solids, and their delicate coupling in computer animation. However, it still remains a challenging problem to model the complex elastic-viscoplastic behaviors during fluid–solid phase transitions and facilitate their seamless interactions inside the same framework. In this article, we propose a practical method capable of simulating granular flows, viscoplastic liquids, elastic-plastic solids, rigid bodies, and interacting with each other, to support novel phenomena all heavily involving realistic phase transitions, including dissolution, melting, cooling, expansion, shrinking, and so on. At the physics level, we propose to combine and morph von Mises with Drucker–Prager and Cam–Clay yield models to establish a unified phase-field-driven EVP model, capable of describing the behaviors of granular, elastic, plastic, viscous materials, liquid, non-Newtonian fluids, and their smooth evolution. At the numerical level, we derive the discretization form of Cahn–Hilliard and Allen–Cahn equations with the material point method to effectively track the phase-field evolution, so as to avoid explicit handling of the boundary conditions at the interface. At the application level, we design a novel heuristic strategy to control specialized behaviors via user-defined schemes, including chemical potential, density curve, and so on. We exhibit a set of numerous experimental results consisting of challenging scenarios to validate the effectiveness and versatility of the new unified approach. This flexible and highly stable framework, founded upon the unified treatment and seamless coupling among various phases, and effective numerical discretization, has its unique advantage in animation creation toward novel phenomena heavily involving phase transitions with artistic creativity and guidance. Zaili Tu, Chen Li 0035, Zipeng Zhao, Changbo Wang, Hong Qin 0001 |
ACM Trans. Graph. | 6 |
| 2024 | SenseMap: Urban Performance Visualization and Analytics Via Semantic Textual SimilarityabstractAs urban populations grow, effectively accessing urban performance measures such as livability and comfort becomes increasingly important due to their significant socioeconomic impacts. While Point of Interest (POI) data has been utilized for various applications in location-based services, its potential for urban performance analytics remains unexplored. In this article, we present SenseMap, a novel approach for analyzing urban performance by leveraging POI data as a semantic representation of urban functions. We quantify the contribution of POIs to different urban performance measures by calculating semantic textual similarities on our constructed corpus. We propose Semantic-adaptive Kernel Density Estimation which takes into account POIs' influential areas across different Traffic Analysis Zones and semantic contributions to generate semantic density maps for measures. We design and implement a feature-rich, real-time visual analytics system for users to explore the urban performance of their surroundings. Evaluations with human judgment and reference data demonstrate the feasibility and validity of our method. Usage scenarios and user studies demonstrate the capability, usability and explainability of our system. Juntong Chen, Qiaoyun Huang, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Image-Driven Harmonious Color Palette Generation for Diverse Information VisualizationabstractColor has been widely used to encode data in all types of visualizations. Effective color palettes contain discriminable and harmonious colors, which allow information from visualizations to be accurately and aesthetically conveyed. However, predefined color palettes not only lack the flexibility of custom color palette generation but also ignore the context in which the visualizations are used. Designing an effective color palette is a time-consuming and challenging process for users, even experts. In this work, we propose the generation of an image-based visualization color palette to exploit the human perception of visually appealing images while considering visualization cognition. By analyzing color palette constraints, including harmony, discrimination, and context, we propose an image-driven color generation method. We design a color clustering method in the saliency-hue plane based on visual importance detection and then select the palette based on the visualization color constraints. In addition, we design two color optimization and assignment strategies for visualizations of different data types. Evaluations through numeric indicators and user experiments demonstrate that the palettes predicted by our method are visually related to the original images and are aesthetically pleasing, supporting diverse visualization contexts and data types in practical applications. Mingtian Tao, Yifei Huang 0006, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | GraphDecoder: Recovering Diverse Network Graphs From Visualization Images via Attention-Aware LearningabstractDNGs are diverse network graphs with texts and different styles of nodes and edges, including mind maps, modeling graphs, and flowcharts. They are high-level visualizations that are easy for humans to understand but difficult for machines. Inspired by the process of human perception of graphs, we propose a method called GraphDecoder to extract data from raster images. Given a raster image, we extract the content based on a neural network. We built a semantic segmentation network based on U-Net. We increase the attention mechanism module, simplify the network model, and design a specific loss function to improve the model's ability to extract graph data. After this semantic segmentation network, we can extract the data of all nodes and edges. We then combine these data to obtain the topological relationship of the entire DNG. We also provide an interactive interface for users to redesign the DNGs. We verify the effectiveness of our method by evaluations and user studies on datasets collected on the internet and generated datasets. Sicheng Song, Chenhui Li 0001, Juntong Chen, Changbo Wang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | InvVis: Large-Scale Data Embedding for Invertible VisualizationabstractWe present InvVis, a new approach for invertible visualization, which is reconstructing or further modifying a visualization from an image. InvVis allows the embedding of a significant amount of data, such as chart data, chart information, source code, etc., into visualization images. The encoded image is perceptually indistinguishable from the original one. We propose a new method to efficiently express chart data in the form of images, enabling large-capacity data embedding. We also outline a model based on the invertible neural network to achieve high-quality data concealing and revealing. We explore and implement a variety of application scenarios of InvVis. Additionally, we conduct a series of evaluation experiments to assess our method from multiple perspectives, including data embedding quality, data restoration accuracy, data encoding capacity, etc. The result of our experiments demonstrates the great potential of InvVis in invertible visualization. Huayuan Ye, Chenhui Li 0001, Yang Li 0041, Changbo Wang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Explicit Invariant Feature Induced Cross-Domain Crowd CountingabstractCross-domain crowd counting has shown progressively improved performance. However, most methods fail to explicitly consider the transferability of different features between source and target domains. In this paper, we propose an innovative explicit Invariant Feature induced Cross-domain Knowledge Transformation framework to address the inconsistent domain-invariant features of different domains. The main idea is to explicitly extract domain-invariant features from both source and target domains, which builds a bridge to transfer more rich knowledge between two domains. The framework consists of three parts, global feature decoupling (GFD), relation exploration and alignment (REA), and graph-guided knowledge enhancement (GKE). In the GFD module, domain-invariant features are efficiently decoupled from domain-specific ones in two domains, which allows the model to distinguish crowds features from backgrounds in the complex scenes. In the REA module both inter-domain relation graph (Inter-RG) and intra-domain relation graph (Intra-RG) are built. Specifically, Inter-RG aggregates multi-scale domain-invariant features between two domains and further aligns local-level invariant features. Intra-RG preserves taskrelated specific information to assist the domain alignment. Furthermore, GKE strategy models the confidence of pseudolabels to further enhance the adaptability of the target domain. Various experiments show our method achieves state-of-theart performance on the standard benchmarks. Code is available at https://github.com/caiyiqing/IF-CKT. Yiqing Cai, Lianggangxu Chen, Haoyue Guan, Shaohui Lin, Changhong Lu, Changbo Wang, Gaoqi He |
AAAI | 6 |
| 2023 | Learning Local Features of Motion Chain for Human Motion Prediction
Lianggangxu Chen, Chen Li 0035, Changbo Wang, Gaoqi He |
CGI (3) | 4 |
| 2023 | KDEM: A Knowledge-Driven Exploration Model for Indoor Crowd Evacuation Simulation
Yuji Shen, Bohao Zhang, Chen Li 0035, Changbo Wang, Gaoqi He |
CGI (3) | 4 |
| 2023 | GVQA: Learning to Answer Questions about Graphs with Visualizations via Knowledge BaseabstractGraphs are common charts used to represent the topological relationship between nodes. It is a powerful tool for data analysis and information retrieval tasks involve asking questions about graphs. In formative study, we found that questions for graphs are not only about the relationship of nodes but also about the properties of graph elements. We propose a pipeline to answer natural language questions about graph visualizations and generate visual answers. We first extract the data from graphs and convert them into GML format. We design data structures to encode graph information and convert them into an knowledge base. We then extract topic entities from questions. We feed questions, entities and knowledge bases into our question-answer model to obtain the SPARQL queries for textual answers. Finally, we design a module to present the answers visually. A user study demonstrates that these visual and textual answers are useful, credible and and transparent. Sicheng Song, Juntong Chen, Chenhui Li 0001, Changbo Wang |
CHI | 4 |
| 2023 | Multi-Source Templates Learning for Real-Time Aerial TrackingabstractAerial tracking aims at tracking an arbitrary visual object in a video captured by Unmanned Aerial Vehicles (UAV). Due to the scarce computation resources, the deployment of high-consuming state-of-the-art trackers on UAV becomes impractical. On the other hand, lightweight trackers suffer from inferior performance caused by the low sampling frequency and resolution of UAV videos. In this paper, we propose a novel multi-source templates learning method to alleviate the paradox of efficiency and effectiveness for aerial tracking. Besides conventional static and dynamic templates, our work introduces an additional general-object template to learn common feature properties of a general object during training time. To exploit all templates information, a multi-source templates fusion scheme is proposed to capture characteristics of object in low quality UAV video streams. Furthermore, a joint optimization process is employed to enforce the lightness of model while achieving comparable tracking performance. Our experimental results demonstrate an appealing performance trade-off between accuracy and speed. The proposed tracker achieves 200 FPS on GPU, 100 FPS on CPU, and 12 FPS on Nvidia Jetson Xavier NX, respectively. Our code will be released at https://github.com/vpx-ecnu/MSTL. Yiming Sun 0006, Yang Li 0041, Changbo Wang |
ICASSP | 3 |
| 2023 | MTT-DynGL: Towards Multidimensional Topology-oriented Time-series Dynamic Graphs Learning ModelabstractDynamic graph learning has received increasing attention in recent years. However, real-world graph data sets are characterized by significant structural complexity, attribute diversity, and temporal variability. Importantly, there are complex and significant influence mechanisms between them. All them pose great challenges to dynamic graph learning (DGL). To address them, we propose a novel dynamic graph learning framework, MTT-DynGL. First, graph attention networks (GAT) is used to efficiently aggregate the topology and multidimensional attribute features on each snapshot. Then, a temporal variation matrix with strength factors is designed to further measure the interaction mechanism between structures and attributes over time. Further, to effectively integrate the above results, a MTT-based dynamic graph learning network is designed. It consists of an MTT integration mechanism and a bidirectional dilated causal convolution network. The former is used to learn temporal variation features in an integrated manner, and the latter is used to improve learning quality and training efficiency. Finally, the effectiveness of our method is verified by multiple experiments. Yujie Mao, Yiding Shen, Wenli Xiong, Feng Liu 0039, Chenhui Li 0001, Changbo Wang |
ICDM | 7 |
| 2023 | Scene Graph Generation using Depth-based Multimodal NetworkabstractScene graph generation (SGG) provides an efficient way for scene understanding. However, it has been plagued by the inaccurate classification of relative spatial relationship and incorrect feature information aggregation from distant objects. In this paper, we innovatively introduce the depth information of objects into SGG and propose a multimodal edge-featured graph attention network (MEGA-Net). MEGA-Net primarily comprises three modules. First, the edge-aware message passing (EMP) module extracts multimodal features and fuses them as edge features in the graph network via a quadrilinear model. Multimodal features consist of depth features, visual features, spatial features, and linguistic features. The depth feature in EMP provides the relative spatial relationship among objects which prevents the tail spatial predicates from being recognized as the head predicates. Second, we propose a depth-based self-supervised graph attention (DSGAT) module to predict the correlation probability between object pairs. By encoding the depth ranking of different object pairs in 2D images, DSGAT learns more accurate directional attention to avoid unrelated neighbors. Third, we introduce a predicate aware loss (PA-Loss) to alleviate the feature redundancy problem caused by extra depth information. This is achieved by introducing semantic frequency information that reflects the priority between different types of relationships. Systematic experiments show that our method achieves state-of-the-art performance on two popular datasets, VG and VRD. Lianggangxu Chen, Jiale Lu, Changbo Wang, Gaoqi He |
ICME | 3 |
| 2023 | Estimating Market Value of Companies Based on Finance Statement through Data FusionabstractThe evaluation of a company's value can serve as a guide for investors to assess the company and make informed investment decisions. However, conventional valuation techniques are not applicable to Initial Public Offering (IPO) companies in China, mainly due to the absence of historical market performance. In contrast, a company's finance statement provides a periodic overview of the company's operational and production activities, which is linked to its market performance. Traditional methods often rely on the selection of a limited number of financial indicators from the finance statement and the application of regression analysis. These approaches fail to fully exploit the comprehensive data available in the finance statement. This study proposes a comprehensive method that leverages all relevant information contained in the finance statement, including industry interconnections, financial indices, and additional insights obtained from the report. The structured data is analyzed through tree models, while the interrelationships between different companies are modeled through graph neural networks. Our approach offers a multi-perspective evaluation of IPO companies. The results of our experiments demonstrate that our method can effectively utilize the valuable information in finance statements and improve outcomes. Shiqi Jiang 0001, Yaxuan Zheng, Wenli Xiong, Yanpeng Hu, Changbo Wang, Chenhui Li 0001 |
IJCNN | 7 |
| 2023 | Beware of Overcorrection: Scene-induced Commonsense Graph for Scene Graph GenerationabstractA scene graph generation task is largely restricted under a class imbalance. Previous methods have alleviated the class imbalance problem by incorporating commonsense information into the classification, enabling the prediction model to rectify the incorrect head class into the correct tail class. However, the results of commonsense-based models are typically overcorrected, e.g., the visually correct head class is forcibly modified into the wrong tail class. We argue that there are two principal reasons for this phenomenon. First, existing models ignore the semantic gap between commonsense knowledge and real scenes. Second, current commonsense fusion strategies propagate the neighbors in the visual-linguistic contexts without long-range correlation. To alleviate overcorrection, we formulate the commonsense-based scene graph generation task as two sub-problems: scene-induced commonsense graph generation (SI-CGG) and commonsense-inspired scene graph generation (CI-SGG). In SI-CGG module, unlike conventional methods using fixed commonsense graph, we adaptively adjust the node embeddings in a commonsense graph according to their visual appearance and configure the new reasoning edge under a specific visual context. The CI-SGG module is proposed to propagate the information from scene-induced commonsense graph back to the scene graph. It updates the representations of each node in scene graph by the aggregation of neighbourhood information at different scales. Through maximum likelihood optimisation of the logarithmic Gaussian process, the scene graph automatically adapt to the different neighbors in the visual-linguistic contexts. Systematic experiments on the Visual Genome dataset show that our full method achieves state-of-the-art performance. Lianggangxu Chen, Jiale Lu, Youqi Song, Changbo Wang, Gaoqi He |
ACM Multimedia | 4 |
| 2023 | Prior Knowledge-driven Dynamic Scene Graph Generation with Causal InferenceabstractThe task of dynamic scene graph generation (DSGG) aims at constructing a set of frame-level scene graphs for the given video. It suffers from two kinds of spurious correlation problems. First, the spurious correlation between input object pair and predicate label is caused by the biased predicate sample distribution in dataset. Second, the spurious correlation between contextual information and predicate label arises from interference caused by background content in both the current frame and adjacent frames of the video sequence. To alleviate spurious correlations, our work is formulated into two sub-tasks: video-specific commonsense graph generation (VsCG) and causal inference (CI). VsCG module aims to alleviate the first correlation by integrating prior knowledge into prediction. Information of all the frames in current video is used to enhance the commonsense graph constructed from co-occurrence patterns of all training samples. Thus, the commonsense graph has been augmented with video-specific temporal dependencies. Then, a CI strategy with both intervention and counterfactual is used. The intervention component further eliminates the first correlation by forcing the model to consider all possible predicate categories fairly, while the counterfactual component resolves the second correlation by removing the bad effect from context. Comprehensive experiments on the Action Genome dataset show that the proposed method achieves state-of-the-art performance. Jiale Lu, Lianggangxu Chen, Youqi Song, Shaohui Lin, Changbo Wang, Gaoqi He |
ACM Multimedia | 5 |
| 2023 | Enhancing Visual Understanding by Removing Dithering with Global and Self-Conditioned TransformationabstractPNG-8 images are commonly used on the web due to their small size, but their limited color palette often leads to dithering artifacts. Unfortunately, restoring these images using a conventional convolutional neural network (CNN) often results in suboptimal performance since the spatial distribution of dithering is not uniform across the image. This is because the convolutional operator is spatially consistent, meaning it applies the same kernel to all pixels, which we refer to as a global transformation. To address this issue, we propose PNG8IRNet, one approach that combines global and self-conditioned transformations to remove dithering artifacts. Our method incorporates a multilayer perceptron (MLP) to generate diverse kernels for each pixel, taking into account the spatial non-uniformity of dithering, which we define as a self-conditioned transformation. PNG8IRNet demonstrates its performance on multiple datasets, substantially enhancing visual comprehension through a comprehensive set of experiments. Yifei Huang 0006, Chenhui Li 0001, Risheng Liu, Tianyi Liang 0002, Changbo Wang |
VINCI | 5 |
| 2023 | Metaphor Design of Dockless Bike-sharing Based on Spatio-temporal Geographic DataabstractDockless Bike-sharing systems can be treated as a typical paradigm of the sharing economy. Due to the indiscriminate use and placement of huge bikes, their scheduling rules are complicated, which poses a challenge to the macro-control of the companies To enhance the information dimension of bike Origin-Destination (OD) data, an algorithm based on grid features is proposed to reconstruct travel trajectories of OD data. This paper starts by analyzing and designing metaphors to describe the scheduling rules and trip characteristics of Dockless Bike-Sharing. The evaluation shows that the proposed metaphors are user-friendly to novice users. We argue that metaphors provide an effective way for users to understand abstract ideas in a visual design with a large dataset. Kelin Li, Kang Zhang 0001, Changbo Wang |
VINCI | 5 |
| 2023 | iARVis: Mobile AR Based Declarative Information Visualization Authoring, Exploring and SharingabstractWe present iARVis, a proof-of-concept toolkit for creating, experiencing, and sharing mobile AR-based information visualization environments. Over the past years, AR has emerged as a promising medium for information and data visualization beyond the physical media and the desktop, enabling interactivity and eliminating spatial limits. However, the creation of such environments remains difficult and frequently necessitates low-level programming expertise and lengthy hand encodings. We present a declarative approach for defining the augmented reality (AR) environment, including how information is automatically positioned, laid out, and interacted with, to improve the efficiency and flexibility of constructing AR-based information visualization environments. We provide fundamental layout and visual components such as the grid, rich text, images, and charts for the development of complex visualization widgets, as well as automatic targeting methods based on image and object tracking for the development of the AR environment. To increase design efficiency, we also provide features such as hot-reload and several creation levels for both novice and advanced users. We also investigate how the augmented reality-based visualization environment could persist and be shared through the internet and provide ways for storing, sharing, and restoring the environment to give a continuous and seamless experience. To demonstrate the viability and extensibility, we evaluate iARVis using a variety of use cases along with performance evaluation and expert reviews. Chenhui Li 0001, Sicheng Song, Changbo Wang |
VR | 4 |
| 2023 | Contact-conditioned hand-held object reconstruction from single-view images
Yang Li 0041, Adnane Boukhayma, Changbo Wang, Marc Christie |
Comput. Graph. | 4 |
| 2023 | MPM-driven dynamic desiccation cracking and curling in unsaturated soilsabstractAbstract Desiccation cracking of soil‐like materials is a common phenomenon in natural dry environment, however, it remains a challenge to model and simulate complicated multi‐physical processes inside the porous structure. With the goal of tracking such physical evolution accurately, we propose an MPM based method to simulate volumetric shrinkage and crack during moisture diffusion. At the physical level, we introduce Richards equations to evolve the dynamic moisture field to model evaporation and diffusion in unsaturated soils, with which a elastoplastic model is established to simulate strength changes and volumetric shrinkage via a novel saturation‐based hardening strategy during plastic treatment. At the algorithmic level, we develop an MPM‐fashion numerical solver for the proposed physical model and achieve stable yet efficient simulation towards delicate deformation and fracture. At the geometric level, we propose a correlating stretching criteria and a saturation‐aware extrapolation scheme to extend existing surface reconstruction for MPM, producing visual compelling soil appearance. Finally, we manifest realistic simulation results based on the proposed method with several challenging scenarios, which demonstrates usability and efficiency of our method. Zaili Tu, Chen Li 0035, Changbo Wang, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2023 | Dynamic leader role modeling for self-organizing crowd evacuation simulationabstractAbstract The role of the leader in the self‐organizing crowd has a significant influence on the crowd evacuation when an emergency occurs. However, current crowd evacuation models either ignore the leadership characteristics of pedestrians or predefine an invariable leader identity. This article clearly distinguishes leaders from followers in a crowd and evaluates the impact of the leader's role on crowd evacuation under dynamic situations. First, the leadership model of the pedestrian is built based on the classic OCEAN personality parameters. Social dominance tendency is measured for more credibility. Then, the dynamic leader role model (DLRM) is proposed to describe the relationship between leaders and followers, considering the interaction among pedestrians during the evacuation. Finally, a novel crowd simulation algorithm based on the above DLRM is presented using the extended social force model. Various experiments in several typical scenarios verified that our proposed simulation model has better realism. Changbo Wang, Gaoqi He |
Comput. Animat. Virtual Worlds | 2 |
| 2023 | ORCANet: Differentiable multi-parameter learning for crowd simulationabstractAbstract Realistic crowd simulation has always been an important research field in computer graphics. While both agent‐based motion models and data‐driven behavior models have made some progress, they are still suffering from either huge effort of multi‐parameter tuning or limited realistic motion. In this article, we propose a novel and differentiable multi‐parameter learning method for crowd simulation, which is called ORCANet. The main idea is to learn from real data and inverse evaluating the multi‐parameter for subsequent simulation. ORCANet uses classic optimal reciprocal collision avoidance (ORCA) as a basic motion model which is integrated into the deep learning framework. Addressing the feature of linear programming and non‐differentiable operation, a Gaussian kernel is added to approximate the role of neighbor distance in collision avoidance, which turns the original discrete operation into a fully differentiable forward simulation. Furthermore, we leverage ORCANet to optimize the multi‐parameter combination in synthetic and real‐world datasets. ORCANet is proved to rapidly converge to correct parameter values and regenerate the input synthetic sequence. Moreover, experiments on real‐world datasets by the metric of pedestrian trajectories verified that a more realistic crowd simulation has been generated through ORCANet. Chen Li 0035, Changbo Wang, Gaoqi He |
Comput. Animat. Virtual Worlds | 3 |
| 2023 | Video-based spatio-temporal scene graph generation with efficient self-supervision tasks
Lianggangxu Chen, Yiqing Cai, Changhong Lu, Changbo Wang, Gaoqi He |
Multim. Tools Appl. | 4 |
| 2023 | Global Representation Guided Adaptive Fusion Network for Stable Video Crowd CountingabstractModern crowd counting methods in natural scenes, even when video datasets are available, are mostly based on images. Because of background interference or occlusion in the scene, these methods can easily lead to mutations and instability in density prediction. There has been minimal research on how to exploit the inherent consistency among adjacent frames to achieve high estimation accuracy of video sequences. In this study, we explore the long-term global temporal consistency in the video sequence and propose a novel Global Representation Guided Adaptive Fusion Network (GRGAF) for video crowd counting. The primary aim is to establish a long-term temporal representation among consecutive frames to guide the density estimation of local frames, which can alleviate the prediction instability caused by background noise and occlusions in crowd scenes. Moreover, in order to further enforce the temporal consistency, we apply the generative adversarial learning scheme and design a global-local joint loss, which can make the estimated density maps more temporally coherent. Extensive experiments on four challenging video-based crowd counting datasets (FDST, DroneCrowd, MALL and UCSD) demonstrate that our method makes effective use of spatio-temporal information of video and outperforms the other state-of-the-art approach. Yiqing Cai, Zhenwei Ma, Changhong Lu, Changbo Wang, Gaoqi He |
IEEE Trans. Multim. | 4 |
| 2023 | Sampling-Based Planning for Retrieving Near-Cylindrical Objects in Cluttered Scenes Using Hierarchical GraphsabstractWe present an incremental sampling-based task and motion planner for retrieving near-cylindrical objects, like bottle, in cluttered scenes, which computes a plan for removing obstacles to generate a collision-free motion of a robot to retrieve the target object. Our proposed planner uses a two-level hierarchy, including the first-level roadmap for the target object motion and the second-level retrieval graph for the entire robot motion, to aid in deciding the order and trajectory of object removal. We use an incremental expansion strategy to update the roadmap and retrieval graph from the collisions between the target object, the robot, and the obstacles, in order to optimize the object removal sequence. The performance of our method is highlighted in several benchmark scenes, including a fixed robotic arm in a cluttered scene with known obstacle locations and a scene, where locations of some objects or even the target object are unknown due to occlusions. Our method can also efficiently solve the high-dimensional planning problem of object retrieval using a mobile manipulator and be combined with the symbolic planner to plan complex multistep tasks. We deploy our method to a physical robot and integrate it with nonprehensile actions to improve operational efficiency. Compared to the state-of-the-art approaches, our method reduces task and motion planning time up to 24.6$\%$with a higher success rate, and still provides a near-optimal plan. Hao Tian 0003, Chaoyang Song 0001, Changbo Wang, Xinyu Zhang 0002, Jia Pan 0001 |
IEEE Trans. Robotics | 3 |
| 2023 | VividGraph: Learning to Extract and Redesign Network Graphs From Visualization ImagesabstractNetwork graphs are common visualization charts. They often appear in the form of bitmaps in articles, web pages, magazine prints, and designer sketches. People often want to modify graphs because of their poor design, but it is difficult to obtain their underlying data. In this article, we present VividGraph, a pipeline for automatically extracting and redesigning graphs from static images. We propose using convolutional neural networks to solve the problem of graph data extraction. Our method is robust to hand-drawn graphs, blurred graph images, and large graph images. We also present a graph classification module to make it effective for directed graphs. We propose two evaluation methods to demonstrate the effectiveness of our approach. It can be used to quickly transform designer sketches, extract underlying data from existing graphs, and interactively redesign poorly designed graphs. Sicheng Song, Chenhui Li 0001, Yujing Sun 0003, Changbo Wang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Fashion sub-categories and attributes prediction model using deep learning
Amin Muhammad Shoib, Changbo Wang, Summaira Jabeen |
Vis. Comput. | 2 |
| 2022 | DH-GCN: Saliency-Aware Complex Scene Graph Generation Using Dual-Hierarchy Graph Convolutional NetworkabstractIn reality, complex scene plagues numerous scene graph generation models because realistic scene contains myriad of objects and complicated relationships. Most current methods suffer poor performance when encountering complex scenes. We find that there are two principal reasons for this phenomenon. First, the construction of graph loses sight of the hierarchy of objects. Second, there exists redundant information in feature optimization. To facilitate this issue, this paper proposes an innovative dual-hierarchy graph convolutional network (DH-GCN), which is a conceptually elegant and efficient top-down approach. In specific, DH-GCN leverages salient object detector to hierarchize objects and give gist nodes more accurate representation. Moreover, the dual-hierarchy message propagation is designed to refine the representation hierarchically and eliminate redundant information. Systematic experiments on Visual Genome dataset show the superiority of our method over strong baseline methods. Jiale Lu, Lianggangxu Chen, Yiqing Cai, Haoyue Guan, Changhong Lu, Changbo Wang, Gaoqi He |
ICME | 6 |
| 2022 | Skill-Oriented Hierarchical Structure for Deep Knowledge TracingabstractKnowledge tracing (KT) which aims to trace stu-dents' knowledge state is an effective technique in intelligent tutoring systems. Although most KT models have exploited the question side information, plentiful hierarchical information between skills hasn't been well extracted for making more accurate predictions. In this paper, a novel model called Skill-oriented Hierarchical structure for Deep Knowledge Tracing (SHDKT) is proposed to discover the relations between questions, which are implicit in the hierarchical skill structure. SHDKT comprises three modules. First, The skill concurrency graph (SCG) is constructed by incorporating students' response infor-mation into the question-skill bipartite graph, which contains both sequence and co-occurrence relations between skills. Second, a hierarchical skill representation module (HSRM) is proposed to exploit the hierarchical information of skills based on the SCG. Finally, a question representation module (QRM) is presented by learning explicit and implicit interactions of question side infor-mation. Hence we can predict the student response accurately through question representation. Extensive experiments on the KT datasets validate the effectiveness of our model. Zhenyuan Yang, Shimeng Xu, Changbo Wang, Gaoqi He |
ICTAI | 3 |
| 2022 | Unsupervised Textured Terrain Generation via Differentiable RenderingabstractConstructing large-scale realistic terrains using modern modeling tools is an extremely challenging task even for professional users, undermining the effectiveness of video games, virtual reality, and other applications. In this paper, we present a step towards unsupervised and realistic modeling of textured terrains from DEM and satellite imagery, built upon two-stage illumination and texture optimization via differentiable rendering. First, a differentiable renderer for satellite imagery is established based on the Lambert diffuse model that allows inverse optimization of material and lighting parameters towards specific objective. Second, the original illumination direction of satellite imagery is recovered by reducing the difference between the shadow distribution generated by the renderer and that of the satellite image in YCrCb colour space, leveraging the abundant geometric information of DEM. Third, we propose to generate the original texture of the shadowed region by introducing visual consistency and smoothness constraints via differentiable rendering to arrive at an end-to-end unsupervised architecture. Comprehensive experiments demonstrate the effectiveness and efficiency of our proposed method as a potential tool to achieve virtual terrain modeling for widespread graphics applications. Peichi Zhou, Dingbo Lu, Chen Li 0035, Jian Zhang 0070, Changbo Wang |
ACM Multimedia | 6 |
| 2022 | Exploring Contextual Relationships in 3D Cloud Points by Semantic Knowledge MiningabstractAbstract 3D scene graph generation (SGG) aims to predict the class of objects and predicates simultaneously in one 3D point cloud scene with instance segmentation. Since the underlying semantic of 3D point clouds is spatial information, recent ideas of the 3D SGG task usually face difficulties in understanding global contextual semantic relationships and neglect the intrinsic 3D visual structures. To build the global scope of semantic relationships, we first propose two types of Semantic Clue (SC) from entity level and path level, respectively. SC can be extracted from the training set and modeled as the co‐occurrence probability between entities. Then a novel Semantic Clue aware Graph Convolution Network (SC‐GCN) is designed to explicitly model each SC of which the message is passed in their specific neighbor pattern. For constructing the interactions between the 3D visual and semantic modalities, a visual‐language transformer (VLT) module is proposed to jointly learn the correlation between 3D visual features and class label embeddings. Systematic experiments on the 3D semantic scene graph (3DSSG) dataset show that our full method achieves state‐of‐the‐art performance. Lianggangxu Chen, Jiale Lu, Yiqing Cai, Changbo Wang, Gaoqi He |
Comput. Graph. Forum | 4 |
| 2022 | Authoring multi-style terrain with global-to-local control
Jian Zhang 0070, Chen Li 0035, Peichi Zhou, Changbo Wang, Gaoqi He, Hong Qin 0001 |
Graph. Model. | 4 |
| 2022 | Learning frequency-aware convolutional neural network for spatio-temporal super-resolution water surface wavesabstractAbstract As a usual component in virtual scenes, water surface plays an important role in various graphical applications, including special effects, video games, and virtual reality. Although recent years have witnessed significant progress based on Navier–Stokes equations and simplified water models, large‐scale water surface waves with high‐frequency visual details remain computationally expensive for interactive applications. This article proposes a novel frequency‐aware neural network to synthesize consistent and detailed water surface waves from low‐resolution input. At its core, our approach leverage the wavelet transformation theory over space, frequency and direction, and incremental supervision to decompose the 4D amplitude function into multiple smaller subproblems. Specifically, we first customize four subnetworks and corresponding loss functions for super‐resolution of spatial resolution, temporal evolution, wave direction subdivision, and wave number, respectively. Then, to enforce the upsampling along each dimension orthogonal to each other, we introduce a cooperative training scheme to fine‐tune and integrate the proposed subnetworks with carefully designed training dataset. Our method can visually enhance high‐resolution spatial details, temporal coherence, interactions with complex boundaries, and various wave patterns with flexible control along multiple dimensions. Through extensive experiments, our method arrives at 13 speedup for 32 upsampling of various simulation scenarios. We also validate the effectiveness and robustness of our method to produce realistic water surface waves toward artistic innovation. Zaili Tu, Sheng Qiu, Chen Li 0035, Changbo Wang, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2022 | VFDP: Visual Analysis of Flight Delay and Propagation on a Geographical MapabstractThe propagation of flight delays is challenging to analyze because delay events depend on multiple variables. This phenomenon has become even worse with the increasing number of aircraft in China, and research into delay propagation has shown limited progress. In this paper, we design a visual analysis system for flight delay propagation. Unlike conventional flight delay research, this work focuses on the flight delay propagation trends in one region and representing the relationship of delays occurring in multiple airports. First, we construct a Bayesian network to analyze the delay parameters and select delay factors for visualization. Second, the system employs a series of visualization methods to present the propagation of flight delays, including density and flow visualizations. Third, the system combines multiple available visual representations for analyzing flight delays from different aspects. We demonstrate our methods with real data in multiple types of cases, and we evaluate our visual design through user studies. The results help identify several benefits of our system and confirm its usefulness for delay propagation analysis. Chen Chen 0168, Chenhui Li 0001, Changbo Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Harmonious Textual Layout Generation Over Natural Images via Deep Aesthetics LearningabstractAutomatic typography is important because it helps designers avoid highly repetitive tasks and amateur users achieve high-quality textual layout designs. However, there are often many parameters and complicated aesthetic rules that need to be adjusted in automatic typography work. In this paper, we propose an efficient deep aesthetics learning approach to generate harmonious textual layout over natural images, which can be decomposed into two stages, saliency-aware text region proposal and aesthetics-based textual layout selection. Our method incorporates both semantic features and visual perception principles. First, we propose a semantic visual saliency detection network combined with a text region proposal algorithm to generate candidate text anchors with various positions and sizes. Second, a discriminative deep aesthetics scoring model is developed to assess the aesthetic quality of the candidate textual layouts. We build a new Textual Layout Aesthetics dataset with dense annotations of each image and design a reasonable evaluation metric to compare our method with richer baselines. The results demonstrate that our method can generate harmonious textual layouts in various actual scenarios with better performance. Chenhui Li 0001, Peiying Zhang 0002, Changbo Wang |
IEEE Trans. Multim. | 3 |
| 2022 | DDLVis: Real-time Visual Query of Spatiotemporal Data Distribution via Density Dictionary LearningabstractVisual query of spatiotemporal data is becoming an increasingly important function in visual analytics applications. Various works have been presented for querying large spatiotemporal data in real time. However, the real-time query of spatiotemporal data distribution is still an open challenge. As spatiotemporal data become larger, methods of aggregation, storage and querying become critical. We propose a new visual query system that creates a low-memory storage component and provides real-time visual interactions of spatiotemporal data. We first present a peak-based kernel density estimation method to produce the data distribution for the spatiotemporal data. Then a novel density dictionary learning approach is proposed to compress temporal density maps and accelerate the query calculation. Moreover, various intuitive query interactions are presented to interactively gain patterns. The experimental results obtained on three datasets demonstrate that the presented system offers an effective query for visual analytics of spatiotemporal data. Chenhui Li 0001, George Baciu, Changbo Wang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | CoPaint: Guiding Sketch Painting with Consistent Color and Coherent Generative Adversarial Networks
Shiqi Jiang 0001, Chenhui Li 0001, Changbo Wang |
CGI | 3 |
| 2021 | Leveraging Intra-Domain Knowledge to Strengthen Cross-Domain Crowd CountingabstractUnsupervised cross-domain counting research using synthetic datasets becomes imminent when considering the laborious labeling for supervised methods. However, the existing methods only focus on learning domain shared knowledge to narrow the gap between the source domain and target domain (inter-domain gap). Nevertheless, these methods do not consider the enormous distribution gap among the target domain data itself (intra-domain gap). In this paper, we propose a two-step domain adaptation method with multi-level feature response branches, which further uses the intra-domain knowledge to strengthen the target domain’s adaptability. Specifically, we first use different feature response branches to learn inter-domain knowledge more robustly, reducing the prediction inconsistency of different scenarios. Subsequently, the trained model is used to generate pseudo-labels for the target domain. The entire model was retrained by using pseudo-labels. Various experiments on synthetic dataset GCC and three real public datasets validate our proposed method’s availability with higher accuracy. Yiqing Cai, Lianggangxu Chen, Zhenwei Ma, Changhong Lu, Changbo Wang, Gaoqi He |
ICME | 5 |
| 2021 | Industry Chain Graph Building Based on Text Semantic Association MiningabstractThe current volume of data in the field of securities investment is increasing dramatically. Simultaneously, the linkage of data from multiple parties makes investment reasoning decisions more challenging than ever. In response to this problem, the financial field's knowledge graph can improve the efficiency, depth, and breadth of financial practitioners' information analysis. Some existing financial knowledge graphs analyze the shareholding relationship between companies. Still, because they are limited to observing data from the company's perspective, users without professional industry background cannot quickly find the industry factors of stock market changes. This paper proposes a financial knowledge graph from the industry chain's perspective. This paper builds upstream and downstream relationships between industries through Transformer-based bidirectional encoder to mine potential industry chain associations from text data and completes the long industry chain of the stock market. This paper also builds a visualization system to display and explore the connection between listed companies and industries. Users can inspect the industry chain's composition and each company's revenue status and stock market conditions in the industry chain. The experiment shows that when the market price fluctuation is detected, the stock price fluctuation can be traced back to its origin in the knowledge graph. Jipeng Li, Yujing Sun 0003, Chenhui Li 0001, Yanpeng Hu, Changbo Wang |
IJCNN | 5 |
| 2021 | NVNet: An Enhanced Attention Network for Segmenting Neck Vascular from Ultrasound ImagesabstractUltrasound images often contain much noise, and the examination process is easily affected by many factors. Therefore, it is often necessary for ultrasound surgeons to have rich experience in accurately identifying neck vascular from ultrasound images. The NVNet proposed in this paper can accurately segment neck vessels and accurately segment carotid intima-media from ultrasound images. We use an improved full-scale skip connection to obtain richer feature information from the encoder and introduce enhanced attention mechanism, making it possible for NVNet to identify neck vascular from ultrasound images containing much noise accurately. Due to the lack of available datasets, we collate an entirely new carotid longitudinal sectional ultrasound dataset and carry out data annotation under ultrasound surgeons' guidance. The experiment is carried out on the collated dataset and another public dataset of cross-sectional ultrasound images, including carotid artery and internal jugular vein. The final experimental results prove that the segmentation accuracy of NVNet exceeds that of many well-known models in recent years. Bohao Zhang, Changbo Wang, Chenhui Li 0001 |
IJCNN | 2 |
| 2021 | OpinionManager: Visual Exploration of Online Reviews in P2P AccommodationabstractUser-generated online reviews are critical in P2P accommodations. They contain a wealth of information about the opinions and experiences of users, which help better understand consumer decisions and improve products and services. However, the huge volume of reviews makes it difficult for potential customers to gain useful insights and for managers to track customer opinions. To address these problems, we first use topic modeling techniques for customer opinion mining. Then, we build a deep learning network for sentiment analysis. Finally, we perform sentiment analysis of the reviews at the aspect level to obtain the sentiment vector representation of the accommodation. Moreover, we design a visual analytic system with a user-friendly interface to facilitate interactive analysis. Evaluation including user and case studies demonstrates the usefulness and effectiveness of this system. Changbo Wang, Sicheng Song, Kirlin Li, Chenhui Li 0001 |
VINCI | 3 |
| 2021 | A Rapid, End-to-end, Generative Model for Gaseous Phenomena from Limited ViewsabstractAbstract Despite the rapid development and proliferation of computer graphics hardware devices for scene capture in the most recent decade, the high‐resolution 3D/4D acquisition of gaseous scenes (e.g., smokes) in real time remains technically challenging in graphics research nowadays. In this paper, we explore a hybrid approach to simultaneously taking advantage of both the model‐centric method and the data‐driven method. Specifically, this paper develops a novel conditional generative model to rapidly reconstruct the temporal density and velocity fields of gaseous phenomena based on the sequence of two projection views. With the data‐driven method, we can achieve the strong coupling of density update and the estimation of flow motion, as a result, we can greatly improve the reconstruction performance for smoke scenes. First, we employ a conditional generative network to generate the initial density field from input projection views and estimate the flow motion based on the adjacent frames. Second, we utilize the differentiable advection layer and design a velocity estimation network with the long‐term mechanism to help achieve the end‐to‐end training and more stable graphics effects. Third, we can re‐simulate the input scene with flexible coupling effects based on the estimated velocity field subject to artists' guidance or user interaction. Moreover, our generative model could accommodate single projection view as input. In practice, more input projection views are enabling and facilitating the high‐fidelity reconstruction with more realistic and finer details. We have conducted extensive experiments to confirm the effectiveness, efficiency, and robustness of our new method compared with the previous state‐of‐the‐art techniques. Sheng Qiu, Chen Li 0035, Changbo Wang, Hong Qin 0001 |
Comput. Graph. Forum | 3 |
| 2021 | An end-to-end model for chinese calligraphy generation
Peichi Zhou, Zipeng Zhao, Kang Zhang 0001, Chen Li 0035, Changbo Wang |
Multim. Tools Appl. | 5 |
| 2021 | Learning Representations for High-Dynamic-Range Image Color Transfer in a Self-Supervised WayabstractReference-based color transfer between images has been a fundamental function in image editing. However, existing approaches pay less attention to high-dynamic-range (HDR) images. It is worth noting that designing an appropriate representation for HDR images to achieve satisfying color transfer is challenging. In this paper, we propose an innovative high-dynamic-range image color transfer generative adversarial network (HDRCTGAN) to encode the original image into fine representations that allow transfer of the color of the reference image to the target image. We propose to learn fine representations through a generative adversarial network (GAN) in a self-supervised way. Particularly, the proposed method is self-supervised learning that requires only unlabeled HDR images instead of supervised learning that requires lots of ground truth pairs. HDRCTGAN consists of a generator to transfer the color of the reference image to the target image over the feature domain and a discriminator to suppress the artifacts caused by the generator. We also design a loss function to ensure that HDRCTGAN possesses two required properties: (a) high fidelity and (b) self-identity. The proposed approach yields a pleasing visual result. We have carried out HDR specific evaluations including both objective quantitative experiments with HDR metrics and subjective user studies operated on HDR display devices to demonstrate the effectiveness of our method. Furthermore, we have verified the applicability of the proposed approach to several applications, such as color transfer of HDR images captured by smartphones, color transfer of fabric images, and reference-based grayscale image colorization. Yifei Huang 0006, Sheng Qiu, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Learning Physical Parameters and Detail Enhancement for Gaseous Scene Design Based on Data GuidanceabstractThis article articulates a novel learning framework for both parameter estimation and detail enhancement for Eulerian gas based on data guidance. The key motivation of this article is to devise a new hybrid, grid-based simulation that could inherit modeling and simulation advantages from both physically-correct simulation methods and powerful data-driven methods, while combating existing difficulties exhibited in both approaches. We first employ a convolutional neural network (CNN) to estimate the physical parameters of gaseous phenomena in Eulerian settings, then we can use the just-learnt parameters to re-simulate (with or without artists' guidance) for specific scenes with flexible coupling effects. Next, a second CNN is adopted to reconstruct the high-resolution velocity field to guide a fast re-simulation on the finer grid, achieving richer and more realistic details with little extra computational expense. From the perspective of physics-based simulation, our trained networks respect temporal coherence and physical constraints. From the perspective of the data-driven machine-learning approaches, our network design aims at extracting a meaningful parameters and reconstructing visually realistic details. Additionally, our implementation based on parallel acceleration could significantly enhance the computational performance of every involved module. Our comprehensive experiments confirm the controllability, effectiveness, and accuracy of our novel approach when producing various gaseous scenes with rich details for widespread graphics applications. Chen Li 0035, Sheng Qiu, Changbo Wang, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | VisCode: Embedding Information in Visualization Images using Encoder-Decoder NetworkabstractWe present an approach called VisCode for embedding information into visualization images. This technology can implicitly embed data information specified by the user into a visualization while ensuring that the encoded visualization image is not distorted. The VisCode framework is based on a deep neural network. We propose to use visualization images and QR codes data as training data and design a robust deep encoder-decoder network. The designed model considers the salient features of visualization images to reduce the explicit visual loss caused by encoding. To further support large-scale encoding and decoding, we consider the characteristics of information visualization and propose a saliency-based QR code layout algorithm. We present a variety of practical applications of VisCode in the context of information visualization and conduct a comprehensive evaluation of the perceptual quality of encoding, decoding success rate, anti-attack capability, time performance, etc. The evaluation results demonstrate the effectiveness of VisCode. Peiying Zhang 0002, Chenhui Li 0001, Changbo Wang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | VEFP: visual evaluation of flight procedure in airport terminal
Chen Chen 0168, Chenhui Li 0001, Yannan Qi, Changbo Wang |
Vis. Comput. | 4 |
| 2020 | Novel Sketch-Based 3D Model Retrieval via Cross-domain Feature Clustering and Matching
Jian Zhang 0070, Chen Li 0035, Changbo Wang, Gaoqi He, Hong Qin 0001 |
ICANN (1) | 4 |
| 2020 | Smarttext: Learning To Generate Harmonious Textual Layout Over Natural ImageabstractAutomatic typography is important because it helps designers avoid highly repetitive tasks and amateur users achieve high-quality textual layout designs. However, there are often many parameters that need to be adjusted in automatic typography work. In this paper, we propose an efficient content-aware learning-based framework to generate harmonious textual layout over natural image. Our method incorporates both semantic features and visual perception principles. First, we combine a semantic visual saliency detection network with diffusion equations and a text-region proposal algorithm to generate candidate text anchors with various positions and sizes. Second, we develop a deep scoring network to assess the aesthetic quality of the candidate results. We design multiple evaluations to compare our method with several baselines and a commercial poster design tool. The results demonstrate that our method can generate harmonious textual layout in various actual scenarios with better performance. Peiying Zhang 0002, Chenhui Li 0001, Changbo Wang |
ICME | 3 |
| 2020 | DeSmoothGAN: Recovering Details of Smoothed Images via Spatial Feature-wise Transformation and Full AttentionabstractRecently, generative adversarial networks (GAN) have been widely used to solve image-to-image translation problems such as edges to photos, labels to scenes, and colorizing grayscale images. However, how to recover details of smoothed images is still unexplored. Naively training a GAN like pix2pix causes insufficiently perfect results due to the fact that we ignore two main characteristics including spatial variability and spatial correlation as for this problem. In this work, we propose DeSmoothGAN to utilize both characteristics specifically. The spatial variability indicates that the details of different areas of smoothed images are distinct and they are supposed to be recovered differently. Therefore, we propose to perform spatial feature-wise transformation to recover individual areas differently. The spatial correlation represents that the details of different areas are related to each other. Thus, we propose to apply full attention to consider the relations between them. The proposed method generates satisfying results on several real-world datasets. We have conducted quantitative experiments including smooth consistency and image similarity to demonstrate the effectiveness of DeSmoothGAN. Furthermore, ablation studies are performed to illustrate the usefulness of our proposed feature-wise transformation and full attention. Yifei Huang 0006, Chenhui Li 0001, Xiaohu Guo, Jing Liao 0001, Changbo Wang |
ACM Multimedia | 6 |
| 2020 | Cross-domain retrieving sketch and shape using cycle CNNs
Mingjia Chen, Changbo Wang, Ligang Liu 0001 |
Comput. Graph. | 2 |
| 2020 | A Novel Plastic Phase-Field Method for Ductile Fracture with GPU OptimizationabstractAbstract In this paper, we articulate a novel plastic phase‐field (PPF) method that can tightly couple the phase‐field with plastic treatment to efficiently simulate ductile fracture with GPU optimization. At the theoretical level of physically‐based modeling and simulation, our PPF approach assumes the fracture sensitivity of the material increases with the plastic strain accumulation. As a result, we first develop a hardening‐related fracture toughness function towards phase‐field evolution. Second, we follow the associative flow rule and adopt a novel degraded von Mises yield criterion. In this way, we establish the tight coupling of the phase‐field and plastic treatment, with which our PPF method can present distinct elastoplasticity, necking, and fracture characteristics during ductile fracture simulation. At the numerical level towards GPU optimization, we further devise an advanced parallel framework, which takes the full advantages of hierarchical architecture. Our strategy dramatically enhances the computational efficiency of preprocessing and phase‐field evolution for our PPF with the material point method (MPM). Based on our extensive experiments on a variety of benchmarks, our novel method's performance gain can reach 1.56× speedup of the primary GPU MPM. Finally, our comprehensive simulation results have confirmed that this new PPF method can efficiently and realistically simulate complex ductile fracture phenomena in 3D interactive graphics and animation. Zipeng Zhao, Kemeng Huang, Chen Li 0035, Changbo Wang, Hong Qin 0001 |
Comput. Graph. Forum | 4 |
| 2020 | Novel hierarchical strategies for SPH-centric algorithms on GPGPU
Kemeng Huang, Zipeng Zhao, Chen Li 0035, Changbo Wang, Hong Qin 0001 |
Graph. Model. | 4 |
| 2020 | An advanced hybrid smoothed particle hydrodynamics-fluid implicit particle method on adaptive grid for condensation simulationabstractAbstract In this article, we propose a novel hybrid framework by combining smoothed particle hydrodynamics and adaptive narrow band fluid implicit particle method (NB‐FLIP) to faithfully model the multiphysical processes involving heat transfer and phase transition, and to precisely simulate the dynamics of condensed droplets moving along intricate objects. We first formulate a governing physical model built upon an improved phase transition model and an augmented on‐surface drop analysis method to achieve realistic condensation effects over intricate hydrophilic/hydrophobic interface. To achieve both high‐fidelity interactions and high‐resolution visual effects, we further develop an adaptive NB‐FLIP solver with octree‐dictated background grid in order to further enhance the performance of our framework. Experimental results have shown that our approach can be used to efficiently and realistically simulate the small‐scale interaction details between condensed drops and complex objects with arbitrary geometry. Jiajun Shi, Chen Li 0035, Changbo Wang, Hong Qin 0001, Gaoqi He |
Comput. Animat. Virtual Worlds | 3 |
| 2020 | GenerativeMap: Visualization and Exploration of Dynamic Density Maps via Generative Learning ModelabstractThe density map is widely used for data sampling, time-varying detection, ensemble representation, etc. The visualization of dynamic evolution is a challenging task when exploring spatiotemporal data. Many approaches have been provided to explore the variation of data patterns over time, which commonly need multiple parameters and preprocessing works. Image generation is a well-known topic in deep learning, and a variety of generating models have been promoted in recent years. In this paper, we introduce a general pipeline called GenerativeMap to extract dynamics of density maps by generating interpolation information. First, a trained generative model comprises an important part of our approach, which can generate nonlinear and natural results by implementing a few parameters. Second, a visual presentation is proposed to show the density change, which is combined with the level of detail and blue noise sampling for a better visual effect. Third, for dynamic visualization of large-scale density maps, we extend this approach to show the evolution in regions of interest, which costs less to overcome the drawback of the learning-based generative model. We demonstrate our method on different types of cases, and we evaluate and compare the approach from multiple aspects. The results help identify the effectiveness of our approach and confirm its applicability in different scenarios. Chen Chen 0168, Changbo Wang, Peiying Zhang 0002, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Transferring Grasp Configurations using Active Learning and Local ReplanningabstractWe present a new approach to transfer grasp configurations from prior example objects to novel objects. We assume the novel and example objects have the same topology and similar shapes. We perform 3D segmentation on these objects using geometric and semantic shape characteristics. We compute a grasp space for each part of the example object using active learning. We build bijective contact mapping between these model parts and compute the corresponding grasps for novel objects. Finally, we assemble the individual parts and use local replanning to adjust grasp configurations while maintaining its stability and physical constraints. Our approach is general, can handle all kind of objects represented using mesh or point cloud and a variety of robotic hands. Hao Tian 0003, Changbo Wang, Dinesh Manocha, Xinyu Zhang 0002 |
ICRA | 2 |
| 2019 | Visual Analysis of Retailing Store Location SelectionabstractAn appropriate location of the retailing store is vital for achieving business success. However, a huge amount of complex information needs to be considered in location selection, such as customer flow, the business environment, and current business performance. Unlike traditional location recommendation method on account of statistical sampling, we establish a model of business-district attractiveness based on customer flow and provide a method of data-driven visual comparisons. In addition, we build an interactive visual analysis system with a user-friendly interface for an interactive visual query about complex business and environment information. Our system can help users select retailing store location, support interactive visual queries, and display rich information to facilitate manager in decision-making. Kelin Li, Yi-Na Li, Yanpeng Hu, Changbo Wang |
VINCI | 6 |
| 2019 | EdgeNet: Deep metric learning for 3D shapes
Mingjia Chen, Qianfang Zou 0001, Changbo Wang, Ligang Liu 0001 |
Comput. Aided Geom. Des. | 3 |
| 2019 | Hybrid modeling of Lagrangian-Eulerian method for high-speed fluid simulation
Changbo Wang, Shenfan Zhang, Chen Li 0035, Hong Qin 0001 |
Comput. Graph. | 1 |
| 2019 | Deep Video-Based Performance Synthesis from Sparse Multi-View CaptureabstractAbstract We present a deep learning based technique that enables novel‐view videos of human performances to be synthesized from sparse multi‐view captures. While performance capturing from a sparse set of videos has received significant attention, there has been relatively less progress which is about non‐rigid objects (e.g., human bodies). The rich articulation modes of human body make it rather challenging to synthesize and interpolate the model well. To address this problem, we propose a novel deep learning based framework that directly predicts novel‐view videos of human performances without explicit 3D reconstruction. Our method is a composition of two steps: novel‐view prediction and detail enhancement. We first learn a novel deep generative query network for view prediction. We synthesize novel‐view performances from a sparse set of just five or less camera videos. Then, we use a new generative adversarial network to enhance fine‐scale details of the first step results. This opens up the possibility of high‐quality low‐cost video‐based performance synthesis, which is gaining popularity for VA and AR applications. We demonstrate a variety of promising results, where our method is able to synthesis more robust and accurate performances than existing state‐of‐the‐art approaches when only sparse views are available. Mingjia Chen, Changbo Wang |
Comput. Graph. Forum | 2 |
| 2019 | Data-driven retrieval of spray details with random forest-based distanceabstractAbstract Generating realistic spray details in liquid simulations remains computationally expensive. This paper proposes a data‐driven method to simulate high‐resolution sprays on low‐resolution grids by retrieving details with the most compatible details from a precomputed repository efficiently. We first employ a random forest‐based distance (RFD) to measure the similarity of liquid regions. In consideration of spatiotemporal relationships between one liquid region and its neighbors, we define a multinary label for RFD instead of the original binary one. Our improved RFD enables us to retrieve details that fit ground truth the best. To ensure temporal continuity of our result and to generate new details from existing ones, we formulate a series of forests with a training set from different time steps. Then, we synthesize results of each forest according to their distances. Finally, we put the synthesis result in correct positions to generate desired sprays motion. In our method, a state‐of‐the‐art cascade forest is employed for a higher accuracy. Several experiments with various grid resolutions validate our method both in visual effect and computational cost. Zipeng Zhao, Chen Li 0035, Changbo Wang, Hong Qin 0001, Hongyan Quan |
Comput. Animat. Virtual Worlds | 4 |
| 2019 | Extracting-mapping scheme for the dynamic details in fluid re-simulations from videos
Hongyan Quan, Ning Wang 0081, Jimeng Li, Changbo Wang |
Multim. Syst. | 4 |
| 2019 | Realtime Hand-Object Interaction Using Learned Grasp Space for Virtual EnvironmentsabstractWe present a realtime virtual grasping algorithm to model interactions with virtual objects. Our approach is designed for multi-fingered hands and makes no assumptions about the motion of the user's hand or the virtual objects. Given a model of the virtual hand, we use machine learning and particle swarm optimization to automatically pre-compute stable grasp configurations for that object. The learning pre-computation step is accelerated using GPU parallelization. At runtime, we rely on the pre-computed stable grasp configurations, and dynamics/non-penetration constraints along with motion planning techniques to compute plausible looking grasps. In practice, our realtime algorithm can perform virtual grasping operations in less than 20ms for complex virtual objects, including high genus objects with holes. We have integrated our grasping algorithm with Oculus Rift HMD and Leap Motion controller and evaluated its performance for different tasks corresponding to grabbing virtual objects and placing them at arbitrary locations. Our user evaluation suggests that our virtual grasping algorithm can increase the user's realism and participation in these tasks and offers considerable benefits over prior interaction algorithms, such as pinch grasping and raycast picking. Hao Tian 0003, Changbo Wang, Dinesh Manocha, Xinyu Zhang 0002 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Example-based rapid generation of vegetation on terrain via CNN-based distribution learning
Jian Zhang 0070, Changbo Wang, Chen Li 0035, Hong Qin 0001 |
Vis. Comput. | 2 |
| 2019 | Procedural modeling of rivers from single image toward natural scene production
Jian Zhang 0070, Changbo Wang, Hong Qin 0001, Yan Gao 0004 |
Vis. Comput. | 2 |
| 2018 | Desnet: Deep Residual Networks for Descalloping of Scansar ImagesabstractScalloping is one of the critical problems in ScanSAR images. It not only affects image visualization, but also influences the quantitative applications such as surface wind and wave retrievals in the ocean area. The existing method of descalloping needs artificial parameter setting and lacks generality in the image domain. A novel deep neural network based on residual learning for descalloping of ScanSAR images is proposed in this paper. The proposed method can eliminate scalloping patterns and has strong adaptive ability, which can handle inhomogeneous scalloping patterns and different scenarios. Experiments on GF-3 ScanSAR images verify the good performance of this method. The code for our models is available online. Shang liang Xu, Xiaolan Qiu, Changbo Wang, Li-Hua Zhong |
IGARSS | 3 |
| 2018 | BehaviorTracker: Visual Analytics of Customer Switching Behavior in O2O MarketabstractVisualization of customer behavior is urgently needed for an increasing number of customer orders on O2O (online to offline) platform. Although many works have been done on visualizing customer opinion or customer click events of one store, visualizing customer switching behavior among stores is still challenging. The challenge is to show customer order records over time and structure the inter-connection among different stores when customer switching behavior happens. In this work, we focus on Takeout O2O service to present a novel visual analysis system for retailers focusing on customer switching behavior patterns. Firstly we define five customer segments based on switching behavior. Then this system enables temporal-spatial driver exploration for different segments through several interactive views. Moreover, in order to visualize inter-connection sequences, augmented streamgraph with the bundled parallel coordinates is proposed as one alternative technique to visualize temporal event sequences. Case studies through collaboration with domain experts also demonstrate the usefulness and effectiveness of this system in helping customer relationship management. Yaru Du, Changbo Wang, Chenhui Li 0001 |
VINCI | 2 |
| 2018 | Jointly learning shape descriptors and their correspondence via deep triplet CNNs
Mingjia Chen, Changbo Wang, Hong Qin 0001 |
Comput. Aided Geom. Des. | 2 |
| 2018 | Translucent Image Recoloring through Homography EstimationabstractAbstract Image color editing techniques are of great significance for users who wish to adjust the image color. However, previous works paid less attention to the translucent images. In this paper, we propose a new method to recolor the translucent images while preserving detailed information and color relationships of the source image. We consider the recolor problem as a location transformation problem and solve it in two steps: automatic palette extraction and homography estimation. First, we propose the Hmeans method to extract the dominant colors of the source image based on histogram statistics and clustering. Then, we propose homography estimation to map the source colors to desired colors in the CIE‐LAB color space. Further, we adopt a non‐linear optimization approach to improve the result generated by the last step. The proposed method maintains high fidelity of the source image. Experiments have shown that our method generates a state‐of‐the‐art visual result, in particular in the shadow areas. The source images with ground truth generated by a ray tracer further verify the effectiveness of our method. Yifei Huang 0006, Changbo Wang, Chenhui Li 0001 |
Comput. Graph. Forum | 2 |
| 2018 | Pore-scale flow simulation in anisotropic porous material via fluid-structure coupling
Chen Li 0035, Changbo Wang, Shenfan Zhang, Sheng Qiu, Hong Qin 0001 |
Graph. Model. | 2 |
| 2018 | Augmented Flow Simulation Based on Tight Coupling Between Video Reconstruction and Eulerian Models
Feng-Yu Li, Changbo Wang, Hong Qin 0001, Hongyan Quan |
J. Comput. Sci. Technol. | 2 |
| 2018 | Thickness-aware voxelizationabstractAbstract Voxelization is a crucial process for many computer graphics applications such as collision detection, rendering of translucent objects, and global illumination. However, in some situations, although the mesh looks good, the voxelization result may be undesirable. In this paper, we describe a novel voxelization method that uses the graphics processing unit for surface voxelization. Our improvements on the voxelization algorithm can address a problem of state‐of‐the‐art voxelization, which cannot deal with thin parts of the mesh object. We improve the quality of voxelization on both normal mediation and surface correction. Furthermore, we investigate our voxelization methods on indirect illumination, showing the improvement on the quality of real‐time rendering. Zhuopeng Zhang, Shigeo Morishima, Changbo Wang |
Comput. Animat. Virtual Worlds | 3 |
| 2017 | Overall quality evaluation of graph layouts based on regression analysisabstractCombining1 subjective evaluation with aesthetic criteria, this paper proposes an objective overall quality assessment method for graph layout algorithms. Firstly, we build the subjective rating database of graph layouts. The subjective experiment is designed to rate different graph layouts. Then, for each graph layout, we use the readability metrics of the layout as independent variables, the subjective score of users as the dependent variable, to establish the regression model. Through the regression model, we can get the overall quality score of a graph layout. Jiafan Li, Changbo Wang, Yuhua Liu, Zhijie Qiao |
VINCI | 2 |
| 2017 | An application of optimization method for storyline based on cluster analysisabstractAs a new visualization technology1, storyline intuitively illustrates the dynamic relationships between entities in a story, which is useful in many applications, including the description of characters' interactions in movies, the evolution of community structure in dynamic social networks, the marital status between people, etc. Previous works optimize the storyline's layout from the perspective of aesthetic standard, significantly reducing line crossings, line wiggles and layout space. But when dealing with large-scale data, there is room for improvement with regard to three issues: insufficient memory space, large time consumption and weak data expression. Therefore, this paper introduces the idea of cluster analysis to the storyline to present the clustering information and reduce the time complexity under the condition where a large number of entities interact in the same period. Meantime a scalable, reusable visualization library of storyline is implemented including some novel interactions. Yuhua Liu, Hanfei Lin, Yitao Liang, Changbo Wang |
VINCI | 4 |
| 2017 | Hybrid modeling of multiphysical processes for particle-based volcano animationabstractAbstract Many complex natural phenomena with dramatic spatial and temporal variation are difficult to animate accurately with anticipated performance in many graphics tasks and applications, because oftentimes in prior art, a single type of physical process could not afford high fidelity and effective scene production. Volcano eruption and its subsequent interaction with earth is one such complicated phenomenon that must depend on multiphysical processes and their tight coupling. This paper documents a novel and effective particle‐based solution for volcano animation that embraces multiphysical processes and their tight unification. First, we introduce a governing physical model consisting of multiphysical processes enabling flexible state transition among solid, fluid, and gas. This computational physics model is dictated by temperature and accommodates dynamic viscosity that is changing according to the temperature. Second, we propose an augmented smoothed particle hydrodynamics as the underlying numerical model to simulate the behavior of lava and smoke with several required physical attributes. Third, multiphysical quantities are tightly coupled to support the interaction with surroundings including fluid–solid coupling, ground friction, and lava–smoke coupling. We also develop a temperature‐directed rendering technique with nearly no extra computational cost and demonstrate realistic graphics effects of volcano eruption and its interaction with earth with visual appeal. Shenfan Zhang, Fanlong Kong, Chen Li 0035, Changbo Wang, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2017 | GenealogyVis: A System for Visual Analysis of Multidimensional Genealogical DataabstractThe study of genealogy is an increasingly popular activity pursued by millions of people, ranging from hobbyists to professional researchers. Such genealogical datasets provide a great opportunity for social science analysts, historians, and the public to study a wide variety of topics in demography, family and household, kinship, stratification, and health. Nevertheless, the large scale and characteristics of the data such as hierarchical, spatiotemporal, and multidimensional also pose special challenges for effective data analysis. In this paper, we introduce GenealogyVis, a visual analytic system to analyze family history and evolution by using the China Multigenerational Panel Dataset-Liaoning, which has more than 1.5 million observations and provides socioeconomic, demographic, and other information for more than 260 000 residents, and further enable users to explore the correlation between the development of families and the social context of environments, economics, policies, and so on. This system includes five main linked views: the Scatter-plot View to provide an overview of the data and further explore the correlation analysis, the Tree View to show the family structure and details for individuals, the Migration View to present the genealogical migratory behaviors, the Matrix View to analyze the reproduction pattern between two generations, and the Stream View to show various statistical information such as demographic information and temporal information. A design study was conducted with a research group led by a domain expert of humanities and social sciences in an iterative manner over half a year. Several in-depth case studies, involving the research group, are described to demonstrate the usefulness of GenealogyVis and discuss new findings. Yuhua Liu, Sicheng Dai, Changbo Wang, Zhiguang Zhou, Huamin Qu |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2017 | Fluid re-simulation based on physically driven model from video
Hongyan Quan, Changbo Wang, Yahui Song |
Vis. Comput. | 2 |
| 2017 | Video-based fluid reconstruction and its coupling with SPH simulation
Changbo Wang, Hong Qin 0001, Tai-you Zhang |
Vis. Comput. | 2 |
| 2016 | EduVis: Visualization for Education Knowledge Graph Based on Web DataabstractHow to clearly present the internal structure of knowledge graph is particularly important, however, the current visualization researches about it are rare. We construct education knowledge graph utilizing extracted entities and entity relations, and construct a visual analysis platform, EduVis. In EduVis, we design and implement a) a layout of events network based on topological structure to explore in details, b) a layout of event network based on timeline to explore time information, c) a click tracking path to record the history of users' clicks and help users backtrack. Kai Sun 0008, Yuhua Liu, Zongchao Guo, Changbo Wang |
VINCI | 4 |
| 2016 | Efficient global penetration depth computation for articulated models
Hao Tian 0003, Xinyu Zhang 0002, Changbo Wang, Jia Pan 0001, Dinesh Manocha |
Comput. Aided Des. | 3 |
| 2015 | Fast animation of debris flow with mixed adaptive grid refinementabstractAbstract Animating debris flow is one of the most challenging tasks in computer graphics, because of its complex dynamic mechanism and the interaction between flows and solids in so large scale region. The difficulty focuses on how to resolve the contradiction between lower computational load and higher request of animating quality. A highly effective method of modeling and animating of debris flow with adaptive grid is presented. First, the debris flow is modeled as Bingham plastic fluid with view‐dependent adaptive grid that is adopted to model the flow volume, and the boundless grids can cover the large scale region of debris flow. Then the mixed grids are built for confluent flows, and the two‐way coupling interaction between flows and environment is considered. After extracting the debris flow surface, adaptive surface tension combining wave particles equation is used to enhance the details and sprays are generated by particles considering the interaction between two fluid volumes. Finally, different dynamic realistic scenes with debris flow are successfully animating at interactive rates. Copyright © 2013 John Wiley & Sons, Ltd. Changbo Wang, Fanlong Kong, Yusheng Gao |
Comput. Animat. Virtual Worlds | 1 |
| 2015 | Novel adaptive SPH with geometric subdivision for brittle fracture animation of anisotropic materials
Chen Li 0035, Changbo Wang, Hong Qin 0001 |
Vis. Comput. | 2 |
| 2014 | A Probabilistic Associative Model for Segmenting Weakly Supervised ImagesabstractWeakly-supervised image segmentation is an important yet challenging task in image processing and pattern recognition fields. It is defined as: in the training stage, semantic labels are only at the image-level, without regard to their specific object/scene location within the image. Given a test image, the goal is to predict the semantics of every pixel/superpixel. In this paper, we propose a new weakly-supervised image segmentation model, focusing on learning the semantic associations between superpixel sets (graphlets in this work). In particular, we first extract graphlets from each image, where a graphlet is a small-sized graph measures the potential of multiple spatially neighboring superpixels (i.e., the probability of these superpixels sharing a common semantic label, such as the "sky" or the "sea"). To compare dierent-sized graphlets and to incorporate image-level labels, a manifold embedding algorithm is designed to transform all graphlets into equal-length feature vectors. Finally, we present a hierarchical Bayesian network (BN) to capture the semantic associations between post-embedding graphlets, based on which the semantics of each superpixel is inferred accordingly. Experimental results demonstrate that: 1) our approach performs competitively compared with the state-of-the-art approaches on three public data sets, and 2) considerable performance enhancement is achieved when using our approach on segmentation-based photo cropping and image categorization. Yi Yang 0001, Yue Gao 0002, Yi Yu 0001, Changbo Wang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2013 | Time-space varying visual analysis of micro-blog sentimentabstractMicro-blog sentiment analysis attracts much attention by companies, governments and other organizations. It could help companies to estimate the extent of product acceptance and to determine marketing strategies, governments to monitor online public perception and to improve government-public relation, etc. Researchers mainly focused on time-varying analysis or space varying analysis. Chenghai Zhang, Yuhua Liu, Changbo Wang |
VINCI | 3 |
| 2013 | Simulation of free-surface flow using a boundless grid
Changbo Wang, Fanlong Kong |
Sci. China Inf. Sci. | 1 |
| 2013 | Flexible and rapid animation of brittle fracture using the smoothed particle hydrodynamics formulationabstractABSTRACT This paper presents a hybrid animation approach to the flexible and rapid crack simulation of brittle material. At the physical level, the local stress tensors induced by collision are analyzed by using the smoothed particle hydrodynamics (SPH) formulation. Specifically, in order to determine the internal stress when rigid bodies collide with each other or neighboring environments, we treat all of them as completely rigid body that has infinite stiffness and then evaluate virtual displacement for colliding particles. At the geometric level, in order to faithfully maintain the fracture interface during the crack simulation, we utilize an efficient shape representation of solid based on the tetrahedral decomposition of the original solid geometry. This novel hybrid approach resorts to local particle models, whose goal is to avoid heavy computational burden during crack interface updating and topological changing, and meanwhile, it facilitates the user‐initiated interactive control during the crack generation and propagation. Our animation experiments demonstrate the effectiveness of our novel particle‐based method to simulate the crack of brittle material. Copyright © 2013 John Wiley & Sons, Ltd. Feibin Chen, Changbo Wang, Buying Xie, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2013 | SentiView: Sentiment Analysis and Visualization for Internet Popular TopicsabstractThere would be value to several domains in discovering and visualizing sentiments in online posts. This paper presents SentiView, an interactive visualization system that aims to analyze public sentiments for popular topics on the Internet. SentiView combines uncertainty modeling and model-driven adjustment. By searching and correlating frequent words in text data, it mines and models the changes of the sentiment on public topics. In addition, using a time-varying helix together with an attribute astrolabe to represent sentiments, it can visualize the changes of multiple attributes and relationships among demographics of interest and the sentiments of participants on popular topics. The relationships of interest among different participants are presented in a relationship map. Using a new evolution model that is based on cellular automata, it is able to compare the time-varying features for sentiment-driven forums on both simulated and real data. Adaptable for different social networking platforms, such as Twitter, blog and forum, the methods demonstrate the effectiveness of SentiView in analyzing and visualizing public sentiments on the Web. Changbo Wang, Zhao Xiao, Yuhua Liu, Yanru Xu, Aoying Zhou, Kang Zhang 0001 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2013 | Hybrid particle-grid fluid animation with enhanced details
Changbo Wang, Fanlong Kong, Hong Qin 0001 |
Vis. Comput. | 1 |
| 2012 | ExtractVis: dynamic visualization of extracting multidimensional dataabstractDue to the excessive items and multiple dimensions of parallel data, traditional visualization methods can not show the prominent information from their characters. This paper proposes a novel method of entity extracting to perform the multi-scale and hierarchical visualization of multi-attribute data set. Firstly, the relationship between these characters can be expressed as entity-relationship and data dimension is expressed as entity attributes, which can eliminate data redundancy and reduce data dimensions. Then a scalable dynamic visualization mode is proposed to show the characters at different levels of details. The method can interactively operate to visualize different data sets, such as electronic commerce data, weather forecast data, and gene expressions data, generating effective visualization results. Zhao Xiao, Changbo Wang, Yuhua Liu, Chenming Pang |
VINCI | 2 |
| 2012 | Simulation of multiple fluids with solid-liquid phase transitionabstractABSTRACT Physically based multiphase fluid simulation has been a hot topic in computer graphics. Since there are complex changing interface topology and interactions among air, solid, and different fluids, few papers have devoted to simulate the multiple fluids phenomena with solid–liquid phase transition. In this paper, the thermal fluid model for phase transition combined with free surface tracking is used to describe the interaction between air and fluids. Then a new model based on hierarchical lattice is proposed to process the solid–liquid interaction and the phase transition in the solid–liquid interface. Further, with the use of hybrid interaction with multidistribution functions, different realistic multiple fluids phenomena are rendered with different lattice sizes. Copyright © 2012 John Wiley & Sons, Ltd. Changbo Wang, Huajun Xiao, Qiuyan Shen |
Comput. Animat. Virtual Worlds | 1 |
| 2011 | Behavior-Based Simulation of Real-Time Crowd EvacuationabstractEmergency evacuation has many applications in computer animation, virtual reality, architecture planning, safety science, etc. However, current methods most focus on the agent-based modeling and simulation. These simulation results can not consider the human behavior fully and their reliabilities are doubtable. This paper presents a new method to simulate the large-scale crowds in real-time and verify the evacuation data in complex environment. Through analyzing the characteristics of human behavior in emergent condition, a mixed geometry-based ant colony evacuation model is firstly proposed. Then, many behaviors of human are considered to calculate the best evacuation path, including autonomous avoidance, human's warning time, and preferential path selecting. The experimental results show that it is an effective method to simulate large-scale crowds in real time, because the verification makes the simulation more reliable as well as making human behavior logical and the virtual scene realistic. Changbo Wang, Chenhui Li 0001, Yuhua Liu, Tianlun Zhang |
CAD/Graphics | 1 |
| 2011 | Visualization and Analysis of Information for Checkin in LBS ModeabstractWith the rapid development of application of LBS (Location Based Service) model, "checkin" has become a hot word. Since the geographic position is combined with local business resources, users can interact with shops through mobile. This Mobile SNS conceals a great deal of valuable information, including users' consumption tendency, merchants' marketing strategies and so on. If we exploit the highly fragmental information and present it with visualization techniques after information integration, many potential values and regulations can be found. Based on such background, we create a multi-axes coordinate system combined with checkin points to visualize this information. Through several interactive techniques, we make a further analysis on different shops in the same commercial circle and the same shop in different commercial circles. From these visualization results, we can make proper suggestions to customers according to their consumption tendency. Changbo Wang, Linling Ma, Qunyan Zhang |
ICIG | 1 |
| 2011 | Adaptive lattice-based light rendering of participating mediaabstractABSTRACT The visual world around us displays a rich set of light effects because of translucent and participating media. It is hard and time consuming to render these effects with scattering, caustic, and shaft because of the complex interaction between light and different media. This paper presents a new rendering method based on adaptive lattice for lighting participating media of translucent materials such as marble, wax, and shaft light. Firstly, on the basis of the lattice‐based photon tracing model, multi‐scale hierarchical lattice was constructed by mixed lattice types sampling combined cubic Cartesian and face‐centered cubic with view‐dependent adaptive resolution. Then, an adaptive method to trace diffuse photons and marked specular photons with different phase functions was suggested. Multiple lights and heterogeneous materials were also considered here. Further, the mixed rendering method and GPU accelerate technology were introduced to render different light effects under different participating media. Copyright © 2011 John Wiley & Sons, Ltd. Changbo Wang, Chenhui Li 0001, Jinqiu Dai, Yang Li 0041 |
Comput. Animat. Virtual Worlds | 1 |
| 2009 | Physically based simulation of tidal boreabstractTidal bore is a peculiar nature phenomenon which is caused by the lunar and solar gravitation. Based on the physical characters of tidal bores, in this paper we propose a novel method to model and render this phenomenon, especially the tidal waves in Qiantang estuary. According to Boltzmann equation for tidal waves, we solve it with the novel triangle mesh of kinectic flux vector splitting (KFVS) mode. Then a method combining a curve forecasting wave and particles model is proposed to render the dynamic scenes of overturning tidal waves. Finally, with some rendering technologies, various realistic tidal waves under diversified conditions is rendered in real time. Changbo Wang, Kun Ji |
CAD/Graphics | 1 |
| 2009 | Real-time fluid simulation with adaptive SPHabstractAbstract We present a new adaptive model for real‐time fluid simulation based on Smoothed Particle Hydrodynamics (SPH) framework. Unlike traditional time‐consuming SPH methods, our model can simulate fluid at a considerably faster speed without losing realism. In our model, we first introduce the non‐uniform particle system and propose a generalized distance field function which considers not only geometrical complexity but also physical complexity of fluid body. And the new sampling rules for splitting and merging of particles are also presented. This can greatly reduce the computation time of the dynamic fluid simulation. Then, a new pressure state equation and an adaptive surface tension model are proposed to enhance the stability of the system and to make the free surface more realistic. To further accelerate the computation, a special fluid solver is designed and implemented using GPU. Various fluid phenomena like breaking wave and flood are simulated at real‐time. Experiments demonstrate that our new adaptive model can greatly enhance the computation efficiency of fluid simulation compared with previous adaptive methods. Copyright © 2009 John Wiley & Sons, Ltd. Zhangye Wang, Changbo Wang, Qunsheng Peng 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2008 | Real-time modeling and rendering of raining scenes
Changbo Wang, Zhangye Wang, Zhiliang Yang, Qunsheng Peng 0001 |
Vis. Comput. | 1 |
| 2007 | Real-Time Rendering of Daylight Sky Scene for Virtual Environment
Changbo Wang |
ICEC | 1 |
| 2007 | Real-time rendering of sky scene considering scattering and refractionabstractAbstract Realistic rendering of sky scene is important in game development and virtual reality. Traditional methods did not consider both the affect of atmospheric scattering and refraction, thus failing to realistically simulate the change of shape, color, and light ring of sky scene. In this paper, a new sky light model considering atmospheric scattering and refraction is proposed. We first calculate the refractional track of light through the atmosphere according to the refraction index. Then we adopt the scattered volume model to simplify the calculation of scattering light intensity. By adapting a path tracing algorithm considering refraction, the intensity distribution of sky light is calculated. Finally, various sky scenes in sunny day, foggy day, and that with rainbow and mirage under different conditions are realistically rendered in real time. Copyright © 2007 John Wiley & Sons, Ltd. Changbo Wang, Zhangye Wang, Qunsheng Peng 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2006 | Real-Time Simulation of Dynamic Mirage Scenes
Changbo Wang, Zhangye Wang, Zhidong Jin, Qunsheng Peng 0001 |
Computer Graphics International | 1 |
| 2006 | Real-time snowing simulation
Changbo Wang, Zhangye Wang, Qunsheng Peng 0001 |
Vis. Comput. | 1 |
| 2005 | Dynamic modeling and rendering of grass wagging in windabstractAbstract Simulation of dynamic natural scene is one of the most challenging tasks in computer graphics. In this paper, we propose a new approach to dynamic modeling and rendering of grasses wagging in wind. Through length preserving free‐form deformation of the 3D skeleton lines of each grass blade and using the alpha test to implement transparent texture mapping, we successfully model grasses of different shapes with rich details. To simulate the real time waggle of grasses, the grasses of a meadow are represented in LOD, while their skeleton lines are dynamically deformed according to some physical models. Simplification technique is also employed to accelerate the collision detection between neighboring grasses. Experiments show that our method can realistically render the animated grass scenes under wind of different speeds and types in real time. Copyright © 2005 John Wiley & Sons, Ltd. Changbo Wang, Zhangye Wang, Chengfang Song, Qunsheng Peng 0001 |
Comput. Animat. Virtual Worlds | 1 |