Rui Ma 0011

dblp:85/5058-11 · DBLP profile ↗
← Back
33ranked-venue papers
2as first author
29since 2021 · last 2026
0000-0002-3477-1466ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 17 · 17 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LPA-Aug: Learning to Place and Adjust Synthetic Objects for LiDAR Data Augmentation
abstract
LiDAR point cloud data is indispensable for autonomous driving systems, enabling critical 3D perception tasks such as 3D object detection and segmentation. However, a shortage of labeled LiDAR data hinders the development of robust deep learning algorithms for these tasks. Data augmentation, as an important and effective technique to increase labeled data, has been applied to LiDAR data in different ways such as geometric transformation, mixup, and inserting synthetic objects. In this paper, we focus on exploring how to utilize synthetic objects to augment LiDAR data in a more effective manner. Given synthetic objects which are in the form of CAD models, we consider three factors that significantly impact the realism and utility of augmented data: insertion pose, point cloud spatial distribution, and point intensity value. Unlike previous insertion methods which only addressed one or two of these factors, our novel framework, LPA Aug, learns to place synthetic objects in and adjust them to existing scenes. A pose prediction module first learns to generate object placement locations and orientations based on scene context. Then, a novel 2D image-based distribution adjustment module adjusts the spatial distribution of points sampled from the inserted synthetic object. Finally, a further intensity prediction module learns in a 2D manner to predict the intensity of each object point. We evaluate LPA-Aug by using the augmented data for 3D object detection on the KITTI dataset. The results demonstrate the superiority of LPA-Aug over prior methods of LiDAR data augmentation.
Shida Wei, Ruifeng Zhai, Rui Ma 0011
Comput. Vis. Media4
2026 Multiscale group behavior perception for intelligent social bot detection
Feng Liu 0045, Bingchuan Jiang, Rui Ma 0011
Expert Syst. Appl.3
2026 Robust social bot detection via multi-relational and noise-aware representation modeling
Feng Liu 0045, Bingchuan Jiang, Rui Ma 0011
Inf. Process. Manag.3
2025 SigStyle: Signature Style Transfer via Personalized Text-to-Image Models
abstract
Style transfer enables the seamless integration of artistic styles from a style image into a content image, resulting in visually striking and aesthetically enriched outputs. Despite numerous advances in this field, existing methods did not explicitly focus on the signature style, which represents the distinct and recognizable visual traits of the image such as geometric and structural patterns, color palettes and brush strokes etc. In this paper, we introduce SigStyle, a framework that leverages the semantic priors that embedded in a personalized text-to-image diffusion model to capture the signature style representation. This style capture process is powered by a hypernetwork that efficiently fine-tunes the diffusion model for any given single style image. Style transfer then is conceptualized as the reconstruction process of content image through learned style tokens from the personalized diffusion model. Additionally, to ensure the content consistency throughout the style transfer process, we introduce a time-aware attention swapping technique that incorporates content information from the original image into the early denoising steps of target image generation. Beyond enabling high-quality signature style transfer across a wide range of styles, SigStyle supports multiple interesting applications, such as local style transfer, texture transfer, style fusion and style-guided text-to-image generation. Quantitative and qualitative evaluations demonstrate our approach outperforms existing style transfer methods for recognizing and transferring the signature styles.
Tongyuan Bai, Xuping Xie, Zili Yi, Rui Ma 0011
AAAI6
2025 Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image Generation
abstract
Recent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity, foreground-background inconsistencies, limited diversity, and reduced control flexibility. These challenges arise from current end-to-end inpainting models, which suffer from inaccurate training masks, limited foreground semantic understanding, data distribution biases, and inherent interference between visual and textual prompts. To overcome these limitations, we present Anywhere, a multi-agent framework that departs from the traditional end-to-end approach. In this framework, each agent is specialized in a distinct aspect, such as foreground understanding, diversity enhancement, object integrity protection, and textual prompt consistency. Our framework is further enhanced with the ability to incorporate optional user textual inputs, perform automated quality assessments, and initiate re-generation as needed. Comprehensive experiments demonstrate that this modular design effectively overcomes the limitations of existing end-to-end models, resulting in higher fidelity, quality, diversity and controllability in foreground-conditioned image generation. Additionally, the Anywhere framework is extensible, allowing it to benefit from future advancements in each individual agent.
Tianyidan Xie, Rui Ma 0011, Xiaoqian Ye, Feixuan Liu, Ying Tai, Zhenyu Zhang 0005, Lanjun Wang, Zili Yi
AAAI2
2025 SingleDream: Attribute-Driven T2I Customization from a Single Reference Image
Zili Yi, Tieru Wu, Rui Ma 0011
CVM (2)5
2025 FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free Prompts
abstract
Controllability plays a crucial role in the practical applications of 3D indoor scene synthesis. Existing works either allow rough language-based control, that is convenient but lacks fine-grained scene customization, or employ graph-based control, which offers better controllability but demands considerable knowledge for the cumbersome graph design process. To address these challenges, we present FreeScene, a user-friendly framework that enables both convenient and effective control for indoor scene synthesis. Specifically, FreeScene supports free-form user inputs including text description and/or reference images, allowing users to express versatile design intentions. The user inputs are adequately analyzed and integrated into a graph representation by a VLM-based Graph Designer. We then propose MG-DiT, a Mixed Graph Diffusion Transformer, which performs graph-aware denoising to enhance scene generation. Our MG-DiT not only excels at preserving graph structure but also offers broad applicability to various tasks, including, but not limited to, text-to-scene, graph-to-scene, and rearrangement, all within a single model. Extensive experiments demonstrate that FreeScene provides an efficient and user-friendly solution that unifies text-based and graph-based scene synthesis, outperforming state-of-the-art methods in terms of both generation quality and controllability in a range of applications.
Tongyuan Bai, Wangyuanfan Bai, Dong Chen 0044, Tieru Wu, Manyi Li, Rui Ma 0011
CVPR6
2025 OmniStyle: Filtering High Quality Style Transfer Data at Scale
abstract
In this paper, we introduce OmniStyle-1M, a large-scale paired style transfer dataset comprising over one million content-style-stylized image triplets across 1,000 diverse style categories, each enhanced with textual descriptions and instruction prompts. We show that OmniStyle-1M can not only enable efficient and scalable of style transfer models through supervised training but also facilitate precise control over target stylization. Especially, to ensure the quality of the dataset, we introduce OmniFilter, a comprehensive style transfer quality assessment framework, which filters high-quality triplets based on content preservation, style consistency, and aesthetic appeal. Building upon this foundation, we propose OmniStyle, a framework based on the Diffusion Transformer (DiT) architecture designed for high-quality and efficient style transfer. This framework supports both instruction-guided and imageguided style transfer, generating high resolution outputs with exceptional detail. Extensive qualitative and quantitative evaluations demonstrate OmniStyle’s superior performance compared to existing approaches, highlighting its efficiency and versatility. OmniStyle-1M and its accompanying methodologies provide a significant contribution to advancing high-quality style transfer, offering a valuable resource for the research community.
Jiang Lin, Zili Yi, Rui Ma 0011
CVPR7
2025 DecoupledGaussian: Object-Scene Decoupling for Physics-Based Interaction
abstract
We present DecoupledGaussian, a novel system that decouples static objects from their contacted surfaces captured in-the-wild videos, a key prerequisite for realistic Newtonian-based physical simulations. Unlike prior methods focused on synthetic data or elastic jittering along the contact surface, which prevent objects from fully detaching or moving independently, DecoupledGaussian allows for significant positional changes without being constrained by the initial contacted surface. Recognizing the limitations of current 2D inpainting tools for restoring 3D locations, our approach proposes joint Poisson fields to repair and expand the Gaussians of both objects and contacted scenes after separation. This is complemented by a multi-carve strategy to refine the object’s geometry. Our system enables realistic simulations of decoupling motions, collisions, and fractures driven by user-specified impulses, supporting complex interactions within and across multiple scenes. We validate DecoupledGaussian through a comprehensive user study and quantitative benchmarks. This system enhances digital interaction with objects and scenes in real-world environments, benefiting industries such as VR, robotics, and autonomous driving. Our project page is at: https://wangmiaowei.github.io/DecoupledGaussian.github.io/.
Miaowei Wang, Weiwei Xu 0003, Rui Ma 0011, Changqing Zou, Daniel D. Morris
CVPR4
2025 One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild
abstract
Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and real-world test samples. While some researchers aim to enhance model generalizability through sophisticated training procedures, advanced architectures, or by creating more diverse datasets, we adopt the test-time fine-tuning paradigm to customize a pre-trained Text2Image (T2I) model. However, naively applying test-time tuning results in inconsistencies in facial identities and appearance attributes. To address this, we introduce a Visual Consistency Module (VCM), which enhances appearance consistency by combining the face, text, and image embedding. Our approach, named OnePoseTrans, requires only a single source image to generate high-quality pose transfer results, offering greater stability than state-of-the-art data-driven methods. For each test case, OnePoseTrans customizes a model in around 48 seconds with an NVIDIA V100 GPU.
Dongqi Fan, Rui Ma 0011, Qiang Tang 0002, Zili Yi
ICASSP4
2025 Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
abstract
Recent advances in large language models (LLMs) have spurred interests in encoding images as discrete tokens and leveraging autoregressive (AR) frameworks for visual generation. However, the quantization process in AR-based visual generation models inherently introduces information loss that degrades image fidelity. To mitigate this limitation, recent studies have explored to autoregressively predict continuous tokens. Unlike discrete tokens that reside in a structured and bounded space, continuous representations exist in an unbounded, high-dimensional space, making density estimation more challenging and increasing the risk of generating out-of-distribution artifacts. Based on the above findings, this work introduces DisCon (Discrete-Conditioned Continuous Autoregressive Model), a novel framework that reinterprets discrete tokens as conditional signals rather than generation targets. By modeling the conditional probability of continuous representations conditioned on discrete tokens, DisCon circumvents the optimization challenges of continuous token modeling while avoiding the information loss caused by quantization. DisCon achieves a gFID score of 1.38 on ImageNet 256$\times$256 generation, outperforming state-of-the-art autoregressive approaches by a clear margin. Project page: https://pengzheng0707.github.io/DisCon.
Yizhou Yu, Rui Ma 0011, Zuxuan Wu
ICCV5
2025 Diff3DS: Generating View-Consistent 3D Sketch via Differentiable Curve Rendering
abstract
3D sketches are widely used for visually representing the 3D shape and structure of objects or scenes. However, the creation of 3D sketch often requires users to possess professional artistic skills. Existing research efforts primarily focus on enhancing the ability of interactive sketch generation in 3D virtual systems. In this work, we propose Diff3DS, a novel differentiable rendering framework for generating view-consistent 3D sketch by optimizing 3D parametric curves under various supervisions. Specifically, we perform perspective projection to render the 3D rational Bézier curves into 2D curves, which are subsequently converted to a 2D raster image via our customized differentiable rasterizer. Our framework bridges the domains of 3D sketch and raster image, achieving end-to-end optimization of 3D sketch through gradients computed in the 2D image domain. Our Diff3DS can enable a series of novel 3D sketch generation tasks, including text-to-3D sketch and image-to-3D sketch, supported by the popular distillation-based supervision, such as Score Distillation Sampling (SDS). Extensive experiments have yielded promising results and demonstrated the potential of our framework. Project: https://yiboz2001.github.io/Diff3DS/
Changqing Zou, Tieru Wu, Rui Ma 0011
ICLR5
2025 TaylorGrid: Towards Fast and High-Quality Implicit Field Learning via Direct Taylor-Based Grid Optimization
Renyi Mao, Rui Ma 0011
PRCV (2)6
2025 Nav2Scene: Navigation-driven fine-tuning for robot-friendly scene generation
abstract
The integration of embodied intelligence in indoor scene synthesis holds significant potential for future interior design applications. Nevertheless, prevailing methodologies for indoor scene synthesis predominantly adhere to data-driven learning paradigms. Despite achieving photorealistic 3D renderings through such approaches, current frameworks systematically neglect to incorporate agent-centric functional metrics essential for optimizing navigational topology and task-oriented interactivity in embodied AI systems like service robotics platforms or autonomous domestic assistants. For example, poorly arranged furniture may prevent robots from effectively interacting with the environment, and this issue cannot be fully resolved by merely introducing prior constraints. To fill this gap, we propose Nav2Scene, a novel plug-and-play fine-tuning mechanism that can be deployed on existing scene generators to enhance the suitability of generated scenes for efficient robot navigation. Specifically, we first introduce path planning score (PPS), which is defined based on the results of the path planning algorithm and can be used to evaluate the robot navigation suitability of a given scene. Then, we pre-compute the PPS of 3D scenes from existing datasets and train a ScoreNet to efficiently predict the PPS of the generated scenes. Finally, the predicted PPS is used to guide the fine-tuning of existing scene generators and produce indoor scenes with higher PPS, indicating improved suitability for robot navigation. We conduct experiments on the 3D-FRONT dataset for different tasks including scene generation, completion and re-arrangement. The results demonstrate that by incorporating our Nav2Scene mechanism, the fine-tuned scene generators can produce scenes with improved navigation compatibility for home robots, while maintaining superior or comparable performance in terms of scene quality and diversity.
Bowei Jiang, Tongyuan Bai, Tieru Wu, Rui Ma 0011
Graph. Model.5
2025 DP-Adapter: Dual-pathway adapter for boosting fidelity and text consistency in customizable human image generation
abstract
With the growing popularity of personalized human content creation and sharing, there is a rising demand for advanced techniques in customized human image generation. However, current methods struggle to simultaneously maintain the fidelity of human identity and ensure the consistency of textual prompts, often resulting in suboptimal outcomes. This shortcoming is primarily due to the lack of effective constraints during the simultaneous integration of visual and textual prompts, leading to unhealthy mutual interference that compromises the full expression of both types of input. Building on prior research that suggests visual and textual conditions influence different regions of an image in distinct ways, we introduce a novel Dual-Pathway Adapter (DP-Adapter) to enhance both high-fidelity identity preservation and textual consistency in personalized human image generation. Our approach begins by decoupling the target human image into visually sensitive and text-sensitive regions. For visually sensitive regions, DP-Adapter employs an Identity-Enhancing Adapter (IEA) to preserve detailed identity features. For text-sensitive regions, we introduce a Textual-Consistency Adapter (TCA) to minimize visual interference and ensure the consistency of textual semantics. To seamlessly integrate these pathways, we develop a Fine-Grained Feature-Level Blending (FFB) module that efficiently combines hierarchical semantic features from both pathways, resulting in more natural and coherent synthesis outcomes. Additionally, DP-Adapter supports various innovative applications, including controllable headshot-to-full-body portrait generation, age editing, old-photo to reality, and expression editing. Extensive experiments demonstrate that DP-Adapter outperforms state-of-the-art methods in both visual fidelity and text consistency, highlighting its effectiveness and versatility in the field of human image generation.
Xuping Xie, Lanjun Wang, Zili Yi, Rui Ma 0011
Graph. Model.6
2025 From the past to the present: A social bot detection method based on spatio-temporal interactive perception
Feng Liu 0045, Rui Ma 0011
Knowl. Based Syst.2
2025 A Community-Aware Spatio-Temporal Hypergraph Contrastive Learning Method for Social Bot Detection
abstract
Social bot detection plays a crucial role in enabling social media platforms and governments to effectively manage and regulate social networks. Existing methods, which typically rely on multi-modal feature fusion, often concentrate on interactions between paired accounts and their immediate neighbors, overlooking the dynamic evolution of social networks over time. To address this gap, we propose a community-aware spatio-temporal hypergraph contrastive learning method for social bot detection, namely BotSTHCL. Specifically, we construct a temporal community-aware hypergraph and apply a contrastive learning framework in a semi-supervised setting, facilitating the effective extraction and representation of node features. Our model not only captures interactions among multiple accounts within a community but also accounts for the evolving nature of social networks. Extensive experiments on publicly available datasets demonstrate the effectiveness and superiority of our approach. The code is available at https://github.com/FengLiuii/BotSTHCL.
Feng Liu 0045, Zhenyu Li 0004, Rui Ma 0011
IEEE Trans. Inf. Forensics Secur.3
2025 BotCF: Improving the Social Bot Detection Performance By Focusing on the Community Features
abstract
Various malicious activities performed by social bots have brought a crisis of trust to online social networks. Existing social bot detection methods often overlook the significance of community structure features and effective fusion strategies for multimodal features. To counter these limitations, we propose BotCF, a novel social bot detection method that incorporates community features and utilizes cross-attention fusion for multimodal features. In BotCF, we extract community features using a community division algorithm based on deep autoencoder-like non-negative matrix factorization. These features capture the social interactions and relationships within the network, providing valuable insights for bot detection. Furthermore, we employ cross-attention fusion to integrate the features of the account’s semantic content, properties, and community structure. This fusion strategy allows the model to learn the interdependencies between different modalities, leading to a more comprehensive representation of each account. Extensive experiments conducted on three publicly available benchmark datasets (Twibot20, Twibot22, and Cresci-2015) demonstrate the effectiveness of BotCF. Compared to state-of-the-art social bot detection models, BotCF achieves significant improvements in accuracy, with an average increase of 1.86%, 1.67%, and 0.47% on the respective datasets. The detection accuracy is boosted to 86.53%, 81.33%, and 98.21%, respectively.
Feng Liu 0045, Zhenyu Li 0004, Chunfang Yang, Daofu Gong, Fenlin Liu, Rui Ma 0011, Adrian G. Bors
IEEE Trans. Netw. Serv. Manag.6
2024 EPSD: Early Pruning with Self-Distillation for Efficient Model Compression
abstract
Neural network compression techniques, such as knowledge distillation (KD) and network pruning, have received increasing attention. Recent work `Prune, then Distill' reveals that a pruned student-friendly teacher network can benefit the performance of KD. However, the conventional teacher-student pipeline, which entails cumbersome pre-training of the teacher and complicated compression steps, makes pruning with KD less efficient. In addition to compressing models, recent compression techniques also emphasize the aspect of efficiency. Early pruning demands significantly less computational cost in comparison to the conventional pruning methods as it does not require a large pre-trained model. Likewise, a special case of KD, known as self-distillation (SD), is more efficient since it requires no pre-training or student-teacher pair selection. This inspires us to collaborate early pruning with SD for efficient model compression. In this work, we propose the framework named Early Pruning with Self-Distillation (EPSD), which identifies and preserves distillable weights in early pruning for a given SD task. EPSD efficiently combines early pruning and self-distillation in a two-step process, maintaining the pruned network's trainability for compression. Instead of a simple combination of pruning and SD, EPSD enables the pruned network to favor SD by keeping more distillable weights before training to ensure better distillation of the pruned network. We demonstrated that EPSD improves the training of pruned networks, supported by visual and quantitative analyses. Our evaluation covered diverse benchmarks (CIFAR-10/100, Tiny-ImageNet, full ImageNet, CUB-200-2011, and Pascal VOC), with EPSD outperforming advanced pruning and SD techniques.
Dong Chen 0044, Ning Liu 0007, Yichen Zhu 0001, Zhengping Che, Rui Ma 0011, Fachao Zhang, Xiaofeng Mou, Jian Tang 0008
AAAI5
2024 3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
abstract
Text-driven 3D scene generation techniques have made rapid progress in recent years. Their success is mainly at-tributed to using existing generative models to iteratively perform image warping and inpainting to generate 3D scenes. However, these methods heavily rely on the out-puts of existing models, leading to error accumulation in geometry and appearance that prevent the models from being used in various scenarios (e.g., outdoor and unreal sce-narios). To address this limitation, we generatively refine the newly generated local views by querying and aggregating global 3D information, and then progressively generate the 3D scene. Specifically, we employ a tri-plane features-based NeRF as a unified representation of the 3D scene to constrain global 3D consistency, and propose a generative refinement network to synthesize new contents with higher quality by exploiting the natural image prior from 2D dif-fusion model as well as the global 3D information of the current scene. Our extensive experiments demonstrate that, in comparison to previous methods, our approach supports wide variety of scene generation and arbitrary camera tra-jectories with improved visual quality and 3D consistency.
Songchun Zhang, Quan Zheng 0004, Rui Ma 0011, Wei Hua 0002, Hujun Bao, Weiwei Xu 0003, Changqing Zou
CVPR4
2024 SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
Zili Yi, Rui Ma 0011
ECCV (30)4
2024 DiffPop: Plausibility-Guided Object Placement Diffusion for Image Composition
abstract
Abstract In this paper, we address the problem of plausible object placement for the challenging task of realistic image composition. We propose DiffPop, the first framework that utilizes plausibility‐guided denoising diffusion probabilistic model to learn the scale and spatial relations among multiple objects and the corresponding scene image. First, we train an unguided diffusion model to directly learn the object placement parameters in a self‐supervised manner. Then, we develop a human‐in‐the‐loop pipeline which exploits human labeling on the diffusion‐generated composite images to provide the weak supervision for training a structural plausibility classifier. The classifier is further used to guide the diffusion sampling process towards generating the plausible object placement. Experimental results verify the superiority of our method for producing plausible and diverse composite images on the new Cityscapes‐OP dataset and the public OPA dataset, as well as demonstrate its potential in applications such as data augmentation and multi‐object placement tasks. Our dataset and code will be released.
Hang Zhou 0007, Shida Wei, Rui Ma 0011
Comput. Graph. Forum4
2024 SAC-GAN: Structure-Aware Image Composition
abstract
We introduce an end-to-end learning framework for image-to-image composition, aiming to plausibly compose an object represented as a cropped patch from an object image into a background scene image. As our approach emphasizes more on semantic and structural coherence of the composed images, rather than their pixel-level RGB accuracies, we tailor the input and output of our network with structure-aware features and design our network losses accordingly, with ground truth established in a self-supervised setting through the object cropping. Specifically, our network takes the semantic layout features from the input scene image, features encoded from the edges and silhouette in the input object patch, as well as a latent code as inputs, and generates a 2D spatial affine transform defining the translation and scaling of the object patch. The learned parameters are further fed into a differentiable spatial transformer network to transform the object patch into the target image, where our model is trained adversarially using an affine transform discriminator and a layout discriminator. We evaluate our network, coined SAC-GAN, for various image composition scenarios in terms of quality, composability, and generalizability of the composite images. Comparisons are made to state-of-the-art alternatives, including Instance Insertion, ST-GAN, CompGAN and PlaceNet, confirming superiority of our method.
Hang Zhou 0007, Rui Ma 0011, Ling-Xiao Zhang, Lin Gao 0004, Ali Mahdavi-Amiri, Hao (Richard) Zhang
IEEE Trans. Vis. Comput. Graph.2
2023 P2M2-Net: Part-Aware Prompt-Guided Multimodal Point Cloud Completion
Linlian Jiang, Tieru Wu, Rui Ma 0011
CAD/Graphics5
2023 Shape-aware fine-grained classification of erythroid cells
Rui Ma 0011, Xiaoqing Ma, Honghua Cui, Yubin Xiao, Xuan Wu 0004, You Zhou 0008
Appl. Intell.2
2023 P3DC-shot: Prior-driven discrete data calibration for nearest-neighbor few-shot classification
Shuangmei Wang, Rui Ma 0011, Tieru Wu, Yang Cao 0010
Image Vis. Comput.2
2022 FD-CAM: Improving Faithfulness and Discriminability of Visual Explanation for CNNs
abstract
Class activation map (CAM) has been widely studied for visual explanation of the internal working mechanism of convolutional neural networks. The key of existing CAM-based methods is to compute effective weights to combine activation maps in the target convolution layer. Existing gradient and score based weighting schemes have shown superiority in ensuring either the discriminability or faithfulness of the CAM, but they normally cannot excel in both properties. In this paper, we propose a novel CAM weighting scheme, named FD-CAM, to improve both the faithfulness and discriminability of the CAM-based CNN visual explanation. First, we improve the faithfulness and discriminability of the score-based weights by performing a grouped channel switching operation. Specifically, for each channel, we compute its similarity group and switch the group of channels on or off simultaneously to compute changes in the class prediction score as the weights. Then, we combine the improved score-based weights with the conventional gradient-based weights so that the discriminability of the final CAM can be further improved. We perform extensive comparisons with the state-of-the-art CAM algorithms. The quantitative and qualitative results show our FD-CAM can produce more faithful and more discriminative visual explanations of the CNNs. We also conduct experiments to verify the effectiveness of the proposed grouped channel switching and weight combination scheme on improving the results. Our code is available at https://github.com/crishhh1998/FD-CAM.
Rui Ma 0011, Tieru Wu
ICPR3
2022 View-Aware Geometry-Structure Joint Learning for Single-View 3D Shape Reconstruction
abstract
Reconstructing a 3D shape from a single-view image using deep learning has become increasingly popular recently. Most existing methods only focus on reconstructing the 3D shape geometry based on image constraints. The lack of explicit modeling of structure relations among shape parts yields low-quality reconstruction results for structure-rich man-made shapes. In addition, conventional 2D-3D joint embedding architecture for image-based 3D shape reconstruction often omits the specific view information from the given image, which may lead to degraded geometry and structure reconstruction. We address these problems by introducing VGSNet, an encoder-decoder architecture for view-aware joint geometry and structure learning. The key idea is to jointly learn a multimodal feature representation of 2D image, 3D shape geometry and structure so that both geometry and structure details can be reconstructed from a single-view image. To this end, we explicitly represent 3D shape structures as part relations and employ image supervision to guide the geometry and structure reconstruction. Trained with pairs of view-aligned images and 3D shapes, the VGSNet implicitly encodes the view-aware shape information in the latent feature space. Qualitative and quantitative comparisons with the state-of-the-art baseline methods as well as ablation studies demonstrate the effectiveness of the VGSNet for structure-aware single-view 3D shape reconstruction.
Xuancheng Zhang, Rui Ma 0011, Changqing Zou, Xibin Zhao, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Spatial-Temporal Residual Aggregation for High Resolution Video Inpainting
Vishnu Sanjay Ramiya Srinivasan, Rui Ma 0011, Qiang Tang 0002, Zili Yi
BMVC2
2018 Language-driven synthesis of 3D scenes from scene databases
abstract
We introduce a novel framework for using natural language to generate and edit 3D indoor scenes, harnessing scene semantics and text-scene grounding knowledge learned from large annotated 3D scene databases. The advantage of natural language editing interfaces is strongest when performing semantic operations at the sub-scene level, acting on groups of objects. We learn how to manipulate these sub-scenes by analyzing existing 3D scenes. We perform edits by first parsing a natural language command from the user and transforming it into a semantic scene graph that is used to retrieve corresponding sub-scenes from the databases that match the command. We then augment this retrieved sub-scene by incorporating other objects that may be implied by the scene context. Finally, a new 3D scene is synthesized by aligning the augmented sub-scene with the user's current scene, where new objects are spliced into the environment, possibly triggering appropriate adjustments to the existing scene arrangement. A suggestive modeling interface with multiple interpretations of user commands is used to alleviate ambiguities in natural language. We conduct studies comparing our approach against both prior text-to-scene work and artist-made scenes and find that our method significantly outperforms prior work and is comparable to handmade scenes even when complex and varied natural sentences are used.
Rui Ma 0011, Akshay Gadi Patil, Matthew Fisher, Manyi Li, Sören Pirk, Binh-Son Hua, Sai-Kit Yeung, Xin Tong 0001, Leonidas J. Guibas, Hao (Richard) Zhang
ACM Trans. Graph.1
2016 Action-driven 3D indoor scene evolution
abstract
We introduce a framework for action-driven evolution of 3D indoor scenes, where the goal is to simulate how scenes are altered by human actions, and specifically, by object placements necessitated by the actions. To this end, we develop an action model with each type of action combining information about one or more human poses, one or more object categories, and spatial configurations of objects belonging to these categories which summarize the object-object and object-human relations for the action. Importantly, all these pieces of information are learned from annotated photos. Correlations between the learned actions are analyzed to guide the construction of an action graph. Starting with an initial 3D scene, we probabilistically sample a sequence of actions from the action graph to drive progressive scene evolution. Each action triggers appropriate object placements, based on object co-occurrences and spatial configurations learned for the action model. We show results of our scene evolution that lead to realistic and messy 3D scenes, as well as quantitative evaluations by user studies which compare our method to manual scene creation and state-of-the-art, data-driven methods, in terms of scene plausibility and naturalness.
Rui Ma 0011, Honghua Li, Changqing Zou, Zicheng Liao, Xin Tong 0001, Hao (Richard) Zhang
ACM Trans. Graph.1
2014 Organizing heterogeneous scene collections through contextual focal points
abstract
We introduce focal points for characterizing, comparing, and organizing collections of complex and heterogeneous data and apply the concepts and algorithms developed to collections of 3D indoor scenes. We represent each scene by a graph of its constituent objects and define focal points as representative substructures in a scene collection. To organize a heterogeneous scene collection, we cluster the scenes based on a set of extracted focal points: scenes in a cluster are closely connected when viewed from the perspective of the representative focal points of that cluster. The key concept of representativity requires that the focal points occur frequently in the cluster and that they result in a compact cluster. Hence, the problem of focal point extraction is intermixed with the problem of clustering groups of scenes based on their representative focal points. We present a co-analysis algorithm which interleaves frequent pattern mining and subspace clustering to extract a set of contextual focal points which guide the clustering of the scene collection. We demonstrate advantages of focal-centric scene comparison and organization over existing approaches, particularly in dealing with hybrid scenes, scenes consisting of elements which suggest membership in different semantic categories.
Kai Xu 0004, Rui Ma 0011, Hao (Richard) Zhang, Chenyang Zhu 0002, Ariel Shamir, Daniel Cohen-Or, Hui Huang 0004
ACM Trans. Graph.2
2014 Topology-varying 3D shape creation via structural blending
abstract
We introduce an algorithm for generating novel 3D models via topology-varying shape blending. Given a source and a target shape, our method blends them topologically and geometrically, producing continuous series of in-betweens as new shape creations. The blending operations are defined on a spatio-structural graph composed of medial curves and sheets. Such a shape abstraction is structure-oriented, part-aware, and facilitates topology manipulations. Fundamental topological operations including split and merge are realized by allowing one-to-many correspondences between the source and the target. Multiple blending paths are sampled and presented in an interactive, exploratory tool for creative 3D modeling. We show a variety of topology-varying 3D shapes generated via continuous structural blending between man-made shapes exhibiting complex topological differences, in real time.
Ibraheem Alhashim, Honghua Li, Kai Xu 0004, Junjie Cao 0001, Rui Ma 0011, Hao (Richard) Zhang
ACM Trans. Graph.5