EDBT 2026 Demo / reviewers in the wild / expert
Ning Xie 0003
dblp:55/4104-3
· DBLP profile ↗
64ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0002-1509-464XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image GenerationabstractMultimodal Large Language Models (MLLMs) with unified architectures excel across a wide range of vision-language tasks, yet aligning them with personalized image generation remains a significant challenge. Existing methods for MLLMs are frequently subject-specific, demanding a data-intensive fine-tuning process for every new subject, which limits their scalability. In this paper, we introduce MM-R1, a framework that integrates a cross-modal Chain-of-Thought (X-CoT) reasoning strategy to unlock the inherent potential of unified MLLMs for personalized image generation. Specifically, we structure personalization as an integrated visual reasoning and generation process: (1) grounding subject concepts by interpreting and understanding user-provided images and contextual cues, and (2) generating personalized images conditioned on both the extracted subject representations and user prompts. To further enhance the reasoning capability, we adopt Grouped Reward Proximal Policy Optimization(GRPO) to explicitly align the generation. Experiments demonstrate that MM-R1 unleashes the personalization capability of unified MLLMs to generate images with high subject fidelity and strong text alignment in a zero-shot manner. Yujia Wu, Kuncheng Li, Jiwei Wei, Shiyuan He, Jinyu Guo, Ning Xie 0003 |
AAAI | 7 |
| 2025 | MS-RainMamba: Learning Multi-Scale State Space Models for Single Image DerainingabstractDespite the significant advances of Convolutional neural networks (CNNs) and Transformers in image deraining, they either suffer from limited receptive fields or incur quadratic complexity, leading to an imbalance between performance and efficiency. Recently, state space models (SSMs) have demonstrated significant potential in modeling long-range dependencies while maintaining linear complexity. However, existing Mamba-based approaches lack the exploration of useful complementary information from multiple image scales, which could be beneficial for facilitating rain removal. In this paper, we propose an effective multi-scale state-space model-based framework (MS-RainMamba) to explore richer scale-space information for better image deraining. Specifically, we design a local-enhanced state space module to better aggregate rich local and global information. In contrast to existing methods that adopt fixed-scale scanning for feature extraction, we develop a multi-scale hierarchical 2D scanning technique to better help image restoration. Experimental results on six benchmarks show that the proposed method performs favorably against state-of-the-art models. Zhanshuo Liu, Tuo Zhao, Tingting Zhao 0001, Yarui Chen, Ning Xie 0003 |
ICASSP | 6 |
| 2025 | KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal Grounding
Ran Ran 0001, Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Chaoning Zhang, Ning Xie 0003, Yang Yang 0002 |
ICCV | 6 |
| 2025 | PDNet: Patch-Wise Deformation Network for Cross-Modal Point Cloud CompletionabstractPoint cloud completion aims to reconstruct complete shapes from partial data. Single-modal methods, constrained by limited prior information, see a substantial boost in completion accuracy with additional guidance provided by image modality. However, most existing methods still adopt encoder-decoder architectures, which tends to cause information loss during compression and decompression operations. To address this, we propose the Patch-wise Deformation Network for cross-modal completion. Rather than encode shapes into high-dimensional representations and subsequently decode them, this strategy directly leverages deformations on point cloud in Euclidean space to retain original information adequately. Specifically, we first split input point cloud into similar patches. Then Patch Deformation module decomposes shape attributes from cross-modal inputs into structural and semantic components as guidance for deformations, enabling high-fidelity completion results. Extensive experiments demonstrate our method’s superiority in point cloud completion task. Jingwen He, Zhenjiang Du, Ning Xie 0003 |
ICME | 3 |
| 2025 | DCCL: Discriminative Cosine Center Learning for 3D Cross-Modal Retrieval with Real-world ImageabstractCross-modal retrieval with 3D models has gained significant attention with the rapid growth of 3D assets. The core challenge lies in learning modality-invariant and discriminative features in a common space. Existing methods often rely on shared class centers in Euclidean space, overlooking directional relationships between samples and non-corresponding centers, while remaining sensitive to modality-specific scales, hindering the learning of discriminative cross-modal centers, especially for dispersed modalities like real-world images. To address these limitations, we propose the Discriminative Cosine Center Learning (DCCL) framework for 3D cross-modal retrieval. DCCL integrates the Adaptive Cosine Center Learning (ACCL) mechanism, optimizing cosine similarity on a shared hypersphere with adaptive penalties for challenging samples. Additionally, the Cross-Modal Affinity Learning (CMAL) mechanism reduces cross-modal discrepancies by pairwise matching data from different modalities. Extensive experiments on five benchmarks demonstrate that DCCL significantly outperforms baseline methods in both synthetic and real-world scenarios. Zengyu Liu, Zhitao Liu, Zhenjiang Du, Ning Xie 0003 |
ICME | 6 |
| 2025 | MSC-Net: Multi-Scale Cross-Modal Network for Point Cloud CompletionabstractPoint clouds captured by scanning devices are often sparse and incomplete. Most existing methods for point cloud completion use 3D coordinates only to infer geometric shapes, making it difficult to reconstruct accurate structures and details. We propose a novel cross-modal approach, called Multi-Scale Cross-modal Network for Point Cloud Completion (MSC-Net), which leverages the information of image modality to guide the geometric inference of missing parts. In order to obtain more abundant geometric information, we extract multi-scale features of partial point cloud. Then we design a feature fusion module, which employs multi-layer cross-attention to achieve the interaction between the image features and point cloud features at different scales. To improve the ability of local feature perception, we further devise an enhanced cross-attention block. In the decoding stage, we adopt a coarse-to-fine strategy, where the geometry-aware upsampling layer is utilized to refine the point cloud step by step. Experiments and ablation studies have demonstrated the effectiveness of our network, proving that our approach outperforms existing approaches. Zhenjiang Du, Zhitao Liu, Mingda Tang, Ning Xie 0003 |
ICME | 7 |
| 2025 | SyncGaussian: Stable 3D Gaussian-Based Talking Head Generation with Enhanced Lip Sync via Discriminative Speech FeaturesabstractGenerating high-fidelity talking heads that maintain stable head poses and achieve robust lip sync remains a significant challenge. Although methods based on 3D Gaussian Splatting (3DGS) offer a promising solution via point-based deformation, they suffer from inconsistent head dynamics and mismatched mouth movements due to unstable Gaussian initialization and incomplete speech features. To overcome these limitations, we introduce SyncGaussian, a 3DGS-based framework that ensures stable head poses, enhanced lip sync, and realistic appearances with real-time rendering. SyncGaussian employs a stable head Gaussian initialization strategy to mitigate head jitter by optimizing commonly used rough head pose parameters. To enhance lip sync, we propose a sync-enhanced encoder that leverages audio-to-text and audio-to-visual speech features. Guided by a tailored cosine similarity loss function, the encoder integrates discriminative speech features through a multi-level sync adaptation mechanism, enabling the learning of an adaptive speech feature space. Extensive experiments demonstrate that SyncGaussian outperforms state-of-the-art methods in image quality, dynamic motion, and lip sync, with the potential for real-time applications. Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Chaoning Zhang, Ning Xie 0003, Yang Yang 0002 |
IJCAI | 6 |
| 2025 | Zeitgebers-Based User Experience Analysis and Time Perception Modeling via Transformer in VRabstractVirtual Reality (VR) creates a highly realistic and controllable simulation environment that can easily manipulate users' perception of space and time. However, while the sensation of “losing track of time” is often associated with enjoyable experiences, both the relationship between time perception and user experience in VR, and the underlying mechanisms of time perception itself, remain largely unexplored. In this study, we first investigated how different zeitgebers—such as light color, music tempo, and VR task—affect time perception. We then introduced the Relative Subjective Time Change (RSTC) method to explore the link between time perception and user experience quantitatively. Furthermore, to uncover the mechanisms underlying time perception in VR, we propose a computational model based on CNN and Transformer, named the Time Perception Modeling Network (TPM-Net), which leverages multimodal physiological data to infer users' time perception states in VR. In a between-subject experiment with 56 participants, our results indicate that the VR task factor significantly influences time perception, with red light and slow-tempo music contributing to an underestimation of time. The RSTC method effectively demonstrates that a relative underestimation of time in VR is strongly associated with enhanced user experience, presence, and engagement. Moreover, the TPM-Net shows great potential in modeling time perception, enabling further inference of relative changes in both time perception and user experience. Our study comprehensively elucidates the mechanisms of time perception in VR. It provides valuable insights and promising methodologies for exploring the relationship between time perception and user experience. Modeling time perception through physiological data marks a first step toward objectively assessing users' temporal perception states, offering a promising tool for VR-based therapy and training systems that require precise temporal awareness. Zengyu Liu, Xiandi Zhu, Zhitao Liu, Yalan Ye, Ning Xie 0003 |
ISMAR | 6 |
| 2025 | SGCDiff: Sketch-Guided Cross-modal Diffusion Model for 3D shape completion
Zhenjiang Du, Zhitao Liu, Zeyu Ma 0002, Ning Xie 0003, Yang Yang 0002 |
Neurocomputing | 6 |
| 2025 | CMNet: Cross-Modal Coarse-to-Fine Network for Point Cloud Completion Based on PatchesabstractPoint clouds serve as the foundational representation of 3D objects, playing a pivotal role in both computer vision and computer graphics. Recently, the acquisition of point clouds has been effortless because of the development of hardware devices. However, the collected point clouds may be incomplete due to environmental conditions, such as occlusion. Therefore, completing partial point clouds becomes an essential task. The majority of current methods address point cloud completion via the utilization of shape priors. While these methods have demonstrated commendable performance, they often encounter challenges in preserving the global structural and geometric details of the 3D shape. In contrast to those mentioned earlier, we propose a novel cross-modal coarse-to-fine network (CMNet) for point cloud completion. Our method utilizes additional image information to provide global information, thus avoiding the loss of structure. To ensure that the generated results contain sufficient geometric details, we propose a coarse-to-fine learning approach based on multiple patches. Specifically, we encode the image and use multiple generators to generate multiple coarse patches, which are combined into a complete shape. Subsequently, based on the coarse patches generated in advance, we generate fine patches by combining partial point cloud information. Experimental results show that our method achieves state-of-the-art performance on point cloud completion. Zhenjiang Du, Zhitao Liu, Jiwei Wei, Sophyani Banaamwini Yussif, Zheng Wang 0044, Ning Xie 0003, Yang Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | CDPNet: Cross-Modal Dual Phases Network for Point Cloud CompletionabstractPoint cloud completion aims at completing shapes from their partial. Most existing methods utilized shape’s priors information for point cloud completion, such as inputting the partial and getting the complete one through an encoder-decoder deep learning structure. However, it is very often to easily cause the loss of information in the generation process because of the invisibility of missing areas. Unlike most existing methods directly inferring the missing points using shape priors, we address it as a cross-modality task. We propose a new Cross-modal Dual Phases Network (CDPNet) for shape completion. Our key idea is that the global information of the shape is obtained from the extra single-view image, and the partial point clouds provide the geometric information. After that, the multi-modal features jointly guide the specific structural information. To learn the geometric details of the shape, we chose to use patches to preserve the local geometric feature. In this way, we can generate shapes with enough geometric details. Experimental results show that our method achieves state-of-the-art performance on point cloud completion. Zhenjiang Du, Jiale Dou, Zhitao Liu, Jiwei Wei, Ning Xie 0003, Yang Yang 0002 |
AAAI | 6 |
| 2024 | PS-DeiT: A Part-Selection Based DeiT for Fine-Grained Classification
Tingting Zhao 0001, Yarui Chen, Ning Xie 0003 |
ICIC (11) | 6 |
| 2024 | PIMT: Physics-Based Interactive Motion Transition for Hybrid Character AnimationabstractMotion transitions, which serve as bridges between two sequences of character animation, play a crucial role in creating long variable animation for real-time 3D interactive applications. In this paper, we present a framework to produce hybrid character animation, which combines motion capture animation and physical simulation animation that seamlessly connects the front and back motion clips. In contrast to previous works using interpolation for transition, our physics-based approach inherently ensures physical validity, and both the transition moment of the source motion clip and the horizontal rotation of the target motion clip can be specified arbitrarily within a certain range, which achieves high responsiveness and wide latitude for user control. The control policy of character can be trained automatically using only the motion capture data that requires transition, and is enhanced by our proposed Self-Behavior Cloning (SBC), an approach to improve the unsupervised reinforcement learning of motion transition. We show that our framework can accomplish the interactive transition tasks from a fully-connected state machine constructed from nine motion clips with high accuracy and naturalness. Yanbin Deng, Ning Xie 0003, Wei Zhang 0290 |
ACM Multimedia | 3 |
| 2024 | Emotion Recognition in HMDs: A Multi-task Approach Using Physiological Signals and Occluded FacesabstractPrior research on emotion recognition in extended reality (XR) has faced challenges due to the occlusion of facial expressions by Head-Mounted Displays (HMDs). This limitation hinders accurate Facial Expression Recognition (FER), which is crucial for immersive user experiences. This study aims to overcome the occlusion challenge by integrating physiological signals with partially visible facial expressions to enhance emotion recognition in XR environments. We employed a multi-task approach, utilizing a feature-level fusion to fuse Electroencephalography (EEG) and Galvanic Skin Response (GSR) signals with occluded facial expressions. The model predicts valence and arousal simultaneously from both macro-and micro-expression. Our method demonstrated improved accuracy in emotion recognition under partial occlusion conditions. The integration of temporal physiological signals with other modalities significantly enhanced performance, particularly for half-face emotion recognition. The study presents a novel approach to emotion recognition in XR, addressing the limitations of facial occlusion by HMDs. The findings suggest that physiological signals are vital for interpreting emotions in occluded scenarios, offering potential for real-time applications and advancing social XR applications. Yunqiang Pei, Jialei Tang, Qihang Tang, Mingfeng Zha, Dongyu Xie, Guoqing Wang 0001, Zhitao Liu, Ning Xie 0003, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 8 |
| 2024 | Improving Interaction Comfort in Authoring Task in AR-HRI through Dynamic Dual-Layer Interaction AdjustmentabstractPrevious research has demonstrated the potential of Augmented Reality in enhancing psychological comfort in Human-Robot Interaction (AR-HRI) through shared robot intent, enhanced visual feedback, and increased expressiveness and creativity in interaction methods. However, the challenge of selecting interaction methods that enhance physical comfort in varying scenarios remains. This study purposes a dynamic dual-layer interaction adjustment mechanism to improve user comfort and interaction efficiency. The mechanism comprises two models: an general layer model, grounded in ergonomics principles, identifies appropriate areas for various interaction methods; a individual layer model predicts user discomfort levels using physiological signals. Interaction methods are dynamically adjusted based on discomfort level changes, enabling the system to adapt to individual differences and dynamic changes, thereby reducing misjudgments and enhancing comfort management. The mechanism's success in authoring tasks validates its effectiveness, significantly advancing AR-HRI and fostering more comfortable and enhancing efficient human-centered interactions. Yunqiang Pei, Hongrong Yang, Qihang Tang, Jialei Tang, Guoqing Wang 0001, Zhitao Liu, Ning Xie 0003, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 9 |
| 2024 | Dynamic Scene Adjustment Mechanism for Manipulating User Experience in VRabstractWith the progression of VR tech, virtual interactive environments are becoming increasingly realistic and controllable. Research has substantiated the influence of VR environmental variables on user experience and engagement. Concurrently, real-time user status monitoring advancements have unlocked dynamic adjustments to VR environments through user interaction with real-time status and feedback, increasing researchers’ focus on enhancing user experience and engagement by adjusting VR environmental variables. This paper introduces an interactive paradigm for VR environments called the Dynamic Scene Adjustment (DSA) mechanism, which seeks to modify the VR environmental variables in real-time according to the user’s status and performance to enhance user engagement and experience. We selected the perspective of the impact of visual environment variables on player status, embedding the DSA mechanism into a music VR game with brain-computer interaction for specific VR tasks. Experimental findings affirm that incorporating the DSA mechanism into the VR game enhances the user’s engagement and performance, thereby strongly validating the rationality of the proposed DSA approach. This work can assist researchers think about dynamic regulation in VR environments from a new perspective and will shed light on the design of VR healing, VR education, VR games, and other fields. Zhitao Liu, Haolan Tang, YouTeng Fan, Ning Xie 0003 |
VR | 6 |
| 2024 | Learning explainable task-relevant state representation for model-free deep reinforcement learning
Tingting Zhao 0001, Guixi Li, Tuo Zhao, Yarui Chen, Ning Xie 0003, Gang Niu 0001, Masashi Sugiyama |
Neural Networks | 5 |
| 2024 | Semantics Disentangling for Cross-Modal RetrievalabstractCross-modal retrieval (e.g., query a given image to obtain a semantically similar sentence, and vice versa) is an important but challenging task, as the heterogeneous gap and inconsistent distributions exist between different modalities. The dominant approaches struggle to bridge the heterogeneity by capturing the common representations among heterogeneous data in a constructed subspace which can reflect the semantic closeness. However, insufficient consideration is taken into the fact that learned latent representations are actually heavily entangled with those semantic-unrelated features, which obviously further compounds the challenges of cross-modal retrieval. To alleviate the difficulty, this work makes an assumption that the data are jointly characterized by two independent features: semantic-shared and semantic-unrelated representations. The former presents characteristics of consistent semantics shared by different modalities, while the latter reflects the characteristics with respect to the modality yet unrelated to semantics, such as background, illumination, and other low-level information. Therefore, this paper aims to disentangle the shared semantics from the entangled features, andthus the purer semantic representation can promote the closeness of paired data. Specifically, this paper designs a novel Semantics Disentangling approach for Cross-Modal Retrieval (termed as SDCMR) to explicitly decouple the two different features based on variational auto-encoder. Next, the reconstruction is performed by exchanging shared semantics to ensure the learning of semantic consistency. Moreover, a dual adversarial mechanism is designed to disentangle the two independent features via a pushing-and-pulling strategy. Comprehensive experiments on four widely used datasets demonstrate the effectiveness and superiority of the proposed SDCMR method by achieving a new bar on performance when compared against 15 state-of-the-art methods. Zheng Wang 0044, Xing Xu 0001, Jiwei Wei, Ning Xie 0003, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Image Process. | 4 |
| 2024 | Eye-Hand Typing: Eye Gaze Assisted Finger Typing via Bayesian Processes in ARabstractNowadays, AR HMDs are widely used in scenarios such as intelligent manufacturing and digital factories. In a factory environment, fast and accurate text input is crucial for operators' efficiency and task completion quality. However, the traditional AR keyboard may not meet this requirement, and the noisy environment is unsuitable for voice input. In this article, we introduce Eye-Hand Typing, an intelligent AR keyboard. We leverage the speed advantage of eye gaze and use a Bayesian process based on the information of gaze points to infer users' text input intentions. We improve the underlying keyboard algorithm without changing user input habits, thereby improving factory users' text input speed and accuracy. In real-time applications, when the user's gaze point is on the keyboard, the Bayesian process can predict the most likely characters, vocabulary, or commands that the user will input based on the position and duration of the gaze point and input history. The system can enlarge and highlight recommended text input options based on the predicted results, thereby improving user input efficiency. A user study showed that compared with the current HoloLens 2 system keyboard, Eye-Hand Typing could reduce input error rates by 28.31 % and improve text input speed by 14.5%. It also outperformed a gaze-only technique, being 43.05% more accurate and 39.55% faster. And it was no significant compromise in eye fatigue. Users also showed positive preferences. Yunlei Ren, Zhitao Liu, Ning Xie 0003 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | GT-Net: Variational Autoencoder Networks based on Graph Transformer for 3D Shape LearningabstractIn this paper, we introduce a novel structure-aware method to generate diverse, and realistic 3D shapes via semantic parts. Most previous works neglect the structural and context information between shape parts and only consider the geometric information. This sometimes leads to the wrong combination of parts during the generation process and brought down the generation quality. To address this issue, we learn a structure-aware latent representation for 3D shapes by training a variational autoencoder(VAE). Specially, we use a graph to express semantic parts and their structural relationship of the 3D shape. Based on that graph representation, we design a generative network based on graph transformer architecture, called Graph Transformer VAE networks(GT-Net), to encode and decode the graph-represented 3D shape. Our experimental results demonstrate that our method achieves better performance than previous methods among various shape families, especially in terms of capturing shape details information. Zhenjiang Du, Ning Xie 0003, Yang Yang 0002 |
ICME | 4 |
| 2023 | MRRA-GAN: Multi-Resolution Relation-Aware GAN for Point Cloud CompletionabstractPoint cloud completion has got increasingly attention recently. Its task is to predict a complete point cloud from a partial one, which plays a vital role in three-dimension technology. In order to better obtain the multi-level local information of the point cloud and better combine the local information with the global information for analysis, we proposed a novel Generative Adversarial Network(GAN) for point cloud completion, which called Multi-Resolution Relation-Aware GAN(MRRA-GAN). We designed a Multi-Resolution Key Points Generator(MKPG) which uses multi-resolution point cloud as input to construct key points, a Point Cloud Tree Generator(PTG) to construct Point Cloud Tree(PCT) and a penalty item called Uniformity Penalty(UP) to increase the uniformity of the output point cloud. Experiments, ablation study and robustness test demonstrate the effectiveness of our network, even chanllenging point cloud with different missing degree. Zhenjiang Du, Qifeng He, Ning Xie 0003 |
ICME | 4 |
| 2023 | Self-Relational Graph Convolution Network for Skeleton-Based Action RecognitionabstractUsing a Graph convolution network (GCN) for constructing and aggregating node features has been helpful for skeleton-based action recognition. The strength of the nodes' relation of an action sequence distinguishes it from other actions. This work proposes a novel spatial module called Multi-scale self-relational graph convolution (MS-SRGC) for dynamically modeling joint relations of action instances. Modeling the joints' relations is crucial in determining the spatial distinctiveness between skeleton sequences; hence MS-SRGC shows effectiveness for activity recognition. We also propose a Hybrid multi-scale temporal convolution network (HMS-TCN) that captures different ranges of time steps along the temporal dimension of the skeleton sequence. In addition, we propose a Spatio-temporal blackout (STB) module that randomly zeroes some continue frames for selected strategic joint groups. We sequentially stack our spatial (MS-SRGC) and temporal (HMS-TCN) modules to form a Self-relational graph convolution network (SR-GCN) block, which we use to construct our SR-GCN model. We append our STB on the SR-GCN model top for the randomized operation. With the effectiveness of ensemble networks, we perform extensive experiments on single and multiple ensembles. Our results beat the state-of-the-art methods on the NTU RGB-D, NTU RGB-D 120, and Northwestern-UCLA datasets. Sophyani Banaamwini Yussif, Ning Xie 0003, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 2 |
| 2023 | SDC: Spatial Depth Completion for Outdoor ScenesabstractDepth completion is a crucial computer vision task that aims to fill in missing or incomplete depth values in a depth map. In this paper, we propose SDC: Spatial Depth Completion for Outdoor Scenes. Our approach leverages a two-stage architecture with a spatial feature extractor (SFE) to utilize multi-scale features for accurate depth completion effectively. The proposed method incorporates attention mechanisms, including the Efficient Position Attention Module (EPAM) and Channel Attention Module (CAM), to adaptively fuse depth map features and improve the accuracy of depth completion. Additionally, the Pearson loss function is employed to further enhance the accuracy of the completed depth maps. Experimental results on the KITTI depth completion benchmark demonstrate that our method achieves comparable or better results than traditional depth completion methods while significantly reducing the number of parameters. The proposed SDC model shows great potential in practical applications of depth completion, with its ability to effectively fuse multiscale features and compact model size. Weipeng Huang, Ning Xie 0003 |
SMC | 3 |
| 2023 | Quaternion Representation Learning for cross-modal matching
Zheng Wang 0044, Xing Xu 0001, Jiwei Wei, Ning Xie 0003, Jie Shao 0001, Yang Yang 0002 |
Knowl. Based Syst. | 4 |
| 2022 | PDP-NET: Patch-Based Dual-Path Network for Point Cloud CompletionabstractPoint cloud completion has become a popular research area in 3D computer vision. It aims to recover the complete point cloud from its partial observation. However, previous methods either directly predict the whole shape, change the original distribution of points, or have limited performance in reconstructing tiny and detailed object components. In this paper, we propose a novel Patch-based Dual-Path Network (PDP-Net) for point cloud completion, which leverages the advantages of different encoder architectures, with one path providing estimation for the global structure of the missing part, and the other path filling in the details by generating several point cloud patches. We also propose an identifier to retain the original points in the partial point cloud possibly. Comprehensive experiments and robustness tests demonstrate the effectiveness of our method even against different missing scales of the point cloud. Code available at: https://github.com/QifHE/PDP-Net-Public. Qifeng He, Ning Xie 0003, Zhenjiang Du, Jiale Dou |
ICME | 2 |
| 2022 | Deep Flow Rendering: View Synthesis via Layer-aware Reflection FlowabstractAbstract Novel view synthesis (NVS) generates images from unseen viewpoints based on a set of input images. It is a challenge because of inaccurate lighting optimization and geometry inference. Although current neural rendering methods have made significant progress, they still struggle to reconstruct global illumination effects like reflections and exhibit ambiguous blurs in highly view‐dependent areas. This work addresses high‐quality view synthesis to emphasize reflection on non‐concave surfaces. We propose Deep Flow Rendering that optimizes direct and indirect lighting separately, leveraging texture mapping, appearance flow, and neural rendering. A learnable texture is used to predict view‐independent features, meanwhile enabling efficient reflection extraction. To accurately fit view‐dependent effects, we adopt a constrained neural flow to transfer image‐space features from nearby views to the target view in an edge‐preserving manner. Then we further implement a fusing renderer that utilizes the predictions of both layers to form the output image. The experiments demonstrate that our method outperforms the state‐of‐the‐art methods at synthesizing various scenes with challenging reflection effects. Pinxuan Dai, Ning Xie 0003 |
Comput. Graph. Forum | 2 |
| 2021 | Twin-Channel Gan: Repair Shape with Twin-Channel Generative Adversarial Network and Structural Constraints
Zhenjiang Du, Ning Xie 0003, Zhitao Liu, Yang Yang 0002 |
CGI | 2 |
| 2021 | Efficient Spatio-Temporal Network with Gated Fusion for Video Super-Resolution
Changyu Li, Dongyang Zhang 0001, Ning Xie 0003, Jie Shao 0001 |
ICANN (5) | 3 |
| 2021 | Learning Multi-dimensional Parallax Prior for Stereo Image Super-Resolution
Changyu Li, Dongyang Zhang 0001, Chunlin Jiang, Ning Xie 0003, Jie Shao 0001 |
ICONIP (6) | 4 |
| 2021 | PFFN: Progressive Feature Fusion Network for Lightweight Image Super-ResolutionabstractRecently, convolutional neural network (CNN) has been the core ingredient of modern models, triggering the surge of deep learning in super-resolution (SR). Despite the great success of these CNN-based methods which are prone to be deeper and heavier, it is impracticable to directly apply these methods for some low-budget devices due to the superfluous computational overhead. To alleviate this problem, a novel lightweight SR network named progressive feature fusion network (PFFN) is developed to seek for better balance between performance and running efficiency. Specifically, to fully exploit the feature maps, a novel progressive attention block (PAB) is proposed as the main building block of PFFN. The proposed PAB adopts several parallel but connected paths with pixel attention, which could significantly increase the receptive field of each layer, distill useful information and finally learn more discriminative feature representations. In PAB, a powerful dual attention module (DAM) is further incorporated to provide the channel and spatial attention mechanism in fairly lightweight manner. Besides, we construct a pretty concise and effective upsampling module with the help of multi-scale pixel attention, named MPAU. All of the above modules ensure the network can benefit from attention mechanism while still being lightweight enough. Furthermore, a novel training strategy following the cosine annealing learning scheme is proposed to maximize the representation ability of the model. Comprehensive experiments show that our PFFN achieves the best performance against all existing lightweight state-of-the-art SR methods with less number of parameters and even performs comparably to computationally expensive networks. Dongyang Zhang 0001, Changyu Li, Ning Xie 0003, Guoqing Wang 0001, Jie Shao 0001 |
ACM Multimedia | 3 |
| 2021 | A model-based reinforcement learning method based on conditional generative adversarial networks
Tingting Zhao 0001, Guixi Li, Le Kong, Yarui Chen, Yuan Wang 0021, Ning Xie 0003, Jucheng Yang 0001 |
Pattern Recognit. Lett. | 7 |
| 2021 | Denoising Monte Carlo renderings via a multi-scale featured dual-residual GAN
Siyuan Fu, Xiao Hua Zhang, Ning Xie 0003 |
Vis. Comput. | 4 |
| 2020 | Data-Driven Spatio-Temporal Analysis via Multi-Modal Zeitgebers and Cognitive Load in VRabstractVirtual Reality (VR) produces a highly realistic simulation environment to engage users with Immersive Virtual Environments (IVEs). To interact effectively with users, VR builds intensive media through the multi-modal sense functions in the lower level, such as visual, auditory, tactile, and olfactory senses. However, the higher-level perceptions, e.g., the temporal duration, the sense of presence, and the cognitive load are less explored. These higher-level perceptions are part of the critical evaluation criteria for VR design. In this paper, we divide the external zeitgebers into visual and auditory zeitgebers. We then combine these zeitgebers with the attention-oriented cognitive load to investigate their effects on temporal estimation and presence, particularly in IVEs. We propose a data-driven method to build a multi-modal predictive equation for time estimation and presence, in an effort to figure out the essential elements of users' spatial and temporal perception in VR. We also design a complicated application and validate the predictive equation. Our feature-based model is able to guide the VR application design in terms of the subjective time length judgment and presence of users as well as achieve a better VR user experience. Haodong Liao, Ning Xie 0003, Jianping Su, Weipeng Huang, Heng Tao Shen |
VR | 2 |
| 2019 | A Siamese-Detection Network for Real-Time Object TrackingabstractSiamese networks recently have caused great attention in video tracking community. It successfully achieves not only real-time tracking, but also high precision. However, the typical Siamese networks are relatively mean performance to the target of scale variation, rapid deformation and etc due to the operation on the central point of the object only, rather than the entire spatial information. In this paper, we propose a spatial-aware video tracking method called SiamRFC to predict the central point of the tracked object and estimate its spatial configuration simultaneously. Specifically, our SiamRFC firstly predicts the central point of tracked object via Siamese network by judging the similarity. According to this central, we further estimate its optimal anchor configuration of the tracked object among multiple candidates, and then based on the optimal anchor to regress target position. Extensive evaluations on VOT2015, VOT2016 and VOT2017 benchmarks demonstrate that the proposed SiamRFC tracker is competitive with state-ofthe-art trackers. Ning Xie 0003, Yang Yang 0002 |
ICTAI | 2 |
| 2019 | Attention Transfer (ANT) Network for View-invariant Action RecognitionabstractWith wide applications in surveillance and human-robot interaction, view-invariant human action recognition is critical, however, challenging, due to the action occlusion and information loss caused by view change. Current methods mainly seek for a common feature space for different views. However, such solutions become invalid when there exist few common features, e.g. large view change. To tackle the problem, we propose an AttentioN Transfer (ANT) Network for view-invariant action recognition. Other than transferring features, ANT transfers attention from the reference view to arbitrary views, which correctly emphasize crucial body joints and their relations for view-invariant representation. In addition, the attention calculation method taking into account both recognition contribution and reliability of skeleton joints generates effective attention. Experiments showed its effectiveness for correctly locating crucial body joints in action sequences. We exhaustively evaluate our approach on the UESTC and the NTU dataset with three types of view-invariant evaluations, i.e. X-view, X-sub, and Arbitrary-view evaluation. Experiment results demonstrate its superiority in view-invariant representation and recognition. Yanli Ji, Feixiang Xu, Yang Yang 0002, Ning Xie 0003, Heng Tao Shen, Tatsuya Harada |
ACM Multimedia | 4 |
| 2019 | Word-to-region attention network for visual question answering
Yang Yang 0002, Yi Bin, Ning Xie 0003, Fumin Shen, Yanli Ji, Xing Xu 0001 |
Multim. Tools Appl. | 4 |
| 2019 | Web-based SBLR method of multimedia tools for computer-aided drawing
Ning Xie 0003, Tingting Zhao 0001, Yang Yang 0002, Heng Tao Shen |
Multim. Tools Appl. | 1 |
| 2019 | Describing Video With Attention-Based Bidirectional LSTMabstractVideo captioning has been attracting broad research attention in the multimedia community. However, most existing approaches heavily rely on static visual information or partially capture the local temporal knowledge (e.g., within 16 frames), thus hardly describing motions accurately from a global view. In this paper, we propose a novel video captioning framework, which integrates bidirectional long-short term memory (BiLSTM) and a soft attention mechanism to generate better global representations for videos as well as enhance the recognition of lasting motions in videos. To generate video captions, we exploit another long-short term memory as a decoder to fully explore global contextual information. The benefits of our proposed method are two fold: 1) the BiLSTM structure comprehensively preserves global temporal and visual information and 2) the soft attention mechanism enables a language decoder to recognize and focus on principle targets from the complex content. We verify the effectiveness of our proposed video captioning framework on two widely used benchmarks, that is, microsoft video description corpus and MSR-video to text, and the experimental results demonstrate the superiority of the proposed approach compared to several state-of-the-art methods. Yi Bin, Yang Yang 0002, Fumin Shen, Ning Xie 0003, Heng Tao Shen, Xuelong Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2019 | Collective Reconstructive Embeddings for Cross-Modal HashingabstractIn this paper, we study the problem of cross-modal retrieval by hashing-based approximate nearest neighbor (ANN) search techniques. Most existing cross-modal hashing work mainly addresses the issue of multi-modal integration complexity using the same mapping and similarity calculation for data from different media types. Nonetheless, this may cause information loss during the mapping process due to overlooking the specifics of each individual modality. In this work, we propose a simple yet effective cross-modal hashing approach, termed Collective Reconstructive Embeddings (CRE), which can simultaneously solve the heterogeneity and integration complexity of multi-modal data. To address the heterogeneity challenge, we propose to process heterogeneous types of data using different modalityspecific models. Specifically, we model textual data with cosine similarity based reconstructive embedding to alleviate the data sparsity to the greatest extent, while for image data we utilize the Euclidean distance to characterize the relationships of the projected hash codes. Meanwhile, we unify the projections of text and image to the Hamming space into a common reconstructive embedding through rigid mathematical reformulation, which not only reduces the optimization complexity significantly but also facilitates the inter-modal similarity preservation among different modalities. We further incorporate the code balance and uncorrelation criteria into the problem, and devise an efficient iterative algorithm for optimization. Comprehensive experiments on four widely-used multimodal benchmarks show that the proposed CRE can achieve superior performance compared to the state-of-the-arts on several challenging cross-modal tasks. Mengqiu Hu, Yang Yang 0002, Fumin Shen, Ning Xie 0003, Richang Hong, Heng Tao Shen |
IEEE Trans. Image Process. | 4 |
| 2019 | Hierarchical Multi-Clue Modelling for POI Popularity Prediction with Heterogeneous Tourist InformationabstractPredicting the popularity of Point of Interest (POI) has become increasingly crucial for location-based services, such as POI recommendation. Most of the existing methods can seldom achieve satisfactory performance due to the scarcity of POI's information, which tendentiously confines the recommendation to popular scene spots, and ignores the unpopular attractions with potentially precious values. In this paper, we propose a novel approach, termed Hierarchical Multi-Clue Fusion (HMCF), for predicting the popularity of POIs. Specifically, in order to cope with the problem of data sparsity, we propose to comprehensively describe POI using various types of user generated content (UGC) (e.g., text and image) from multiple sources. Then, we devise an effective POI modelling method in a hierarchical manner, which simultaneously injects semantic knowledge as well as multi-clue representative power into POIs. For evaluation, we construct a multi-source POI dataset by collecting all the textual and visual content of several specific provinces in China from four main-stream tourism platforms during 2006 to 2017. Extensive experimental results show that the proposed method can significantly improve the performance of predicting the attractions' popularity as compared to several baseline methods. Yang Yang 0002, Yaqian Duan, Xinze Wang, Zi Huang, Ning Xie 0003, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Index and Retrieve Multimedia Data: Cross-Modal Hashing by Learning Subspace Relation
Luchen Liu, Yang Yang 0002, Mengqiu Hu, Xing Xu 0001, Fumin Shen, Ning Xie 0003, Zi Huang |
DASFAA (2) | 6 |
| 2018 | Statistical Modeling of the 3D Geometry and Topology of Botanical TreesabstractAbstract We propose a framework for statistical modeling of the 3D geometry and topology of botanical trees. We treat botanical trees as points in a tree‐shape space equipped with a proper metric that captures the geometric and the topological differences between trees. Geodesics in the tree‐shape space correspond to the optimal sequence of deformations, i.e. bending, stretching, and topological changes, which align one tree onto another. In this way, the 3D tree modeling and synthesis problem becomes a problem of exploring the tree‐shape space either in a controlled fashion, using statistical regression, or randomly by sampling from probability distributions fitted to populations in the tree‐shape space. We show how to use this framework for (1) computing statistical summaries, e.g. the mean and modes of variations, of a population of botanical trees, (2) synthesizing random instances of botanical trees from probability distributions fitted to a population of botanical trees, and (3) modeling, interactively, 3D botanical trees using a simple sketching interface. The approach is fast and only requires as input 3D botanical tree models with a known upright orientation. Hamid Laga, Jinyuan Jia 0002, Ning Xie 0003, Hedi Tabia |
Comput. Graph. Forum | 4 |
| 2018 | Stroke-based stylization by learning sequential drawing examples
Ning Xie 0003, Yang Yang 0002, Heng Tao Shen, Tingting Zhao 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Semantic binary coding for visual recognition via joint concept-attribute modelling
Xing Xu 0001, Haiping Wu, Yang Yang 0002, Fumin Shen, Ning Xie 0003, Yanli Ji |
Multim. Tools Appl. | 5 |
| 2018 | Zero-shot learning via discriminative representation extraction
Teng Long 0002, Xing Xu 0001, Fumin Shen, Li Liu 0004, Ning Xie 0003, Yang Yang 0002 |
Pattern Recognit. Lett. | 5 |
| 2018 | Recurrent attention network using spatial-temporal relations for action recognition
Yang Yang 0002, Yanli Ji, Ning Xie 0003, Fumin Shen |
Signal Process. | 4 |
| 2018 | Hashing with Angular Reconstructive EmbeddingsabstractLarge-scale search methods are increasingly critical for many content-based visual analysis applications, among which hashing-based approximate nearest neighbor search techniques have attracted broad interests due to their high efficiency in storage and retrieval. However, existing hashing works are commonly designed for measuring data similarity by the Euclidean distances. In this paper, we focus on the problem of learning compact binary codes using the cosine similarity. Specifically, we proposed novel angular reconstructive embeddings (ARE) method, which aims at learning binary codes by minimizing the reconstruction error between the cosine similarities computed by original features and the resulting binary embeddings. Furthermore, we devise two efficient algorithms for optimizing our ARE in continuous and discrete manners, respectively. We extensively evaluate the proposed ARE on several large-scale image benchmarks. The results demonstrate that ARE outperforms several state-of-the-art methods. Mengqiu Hu, Yang Yang 0002, Fumin Shen, Ning Xie 0003, Heng Tao Shen |
IEEE Trans. Image Process. | 4 |
| 2018 | The Shape Space of 3D Botanical Tree ModelsabstractWe propose an algorithm for generating novel 3D tree model variations from existing ones via geometric and structural blending. Our approach is to treat botanical trees as elements of a tree-shape space equipped with a proper metric that quantifies geometric and structural deformations. Geodesics, or shortest paths under the metric, between two points in the tree-shape space correspond to optimal deformations that align one tree onto another, including the possibility of expanding, adding, or removing branches and parts. Central to our approach is a mechanism for computing correspondences between trees that have different structures and a different number of branches. The ability to compute geodesics and their lengths enables us to compute continuous blending between botanical trees, which, in turn, facilitates statistical analysis, such as the computation of averages of tree structures. We show a variety of 3D tree models generated with our approach from 3D trees exhibiting complex geometric and structural differences. We also demonstrate the application of the framework in reflection symmetry analysis and symmetrization of botanical trees. Hamid Laga, Ning Xie 0003, Jinyuan Jia 0002, Hedi Tabia |
ACM Trans. Graph. | 3 |
| 2017 | Efficient Binary Coding for Subspace-based Query-by-Image Video RetrievalabstractSubspace representations have been widely applied for videos in many tasks. In particular, the subspace-based query-by-image video retrieval (QBIVR), facing high challenges on similarity-preserving measurements and efficient retrieval schemes, urgently needs considerable research attention. In this paper, we propose a novel subspace-based QBIVR framework to enable efficient video search. We first define a new geometry-preserving distance metric to measure the image-to-video distance, which transforms the QBIVR task to be the Maximum Inner Product Search (MIPS) problem. The merit of this distance metric lies in that it helps to preserve the genuine geometric relationship between query images and database videos to the greatest extent. To boost the efficiency of solving the MIPS problem, we introduce two asymmetric hashing schemes which can bridge the domain gap of images and videos properly. The first approach, termed Inner-product Binary Coding (IBC), achieves high-quality binary codes by learning the binary codes and coding functions simultaneously without continuous relaxations. The other one, Bilinear Binary Coding (BBC) approach, employs compact bilinear projections instead of a single large projection matrix to further improve the retrieval efficiency. Extensive experiments on four real-world video datasets verify the effectiveness of our proposed approaches, as compared to the state-of-the-art methods. Ruicong Xu, Yang Yang 0002, Fumin Shen, Ning Xie 0003, Heng Tao Shen |
ACM Multimedia | 4 |
| 2017 | POI Popularity Prediction via Hierarchical Fusion of Multiple Social CluesabstractPredicting the popularity of Point of Interest (POI) has become increasingly crucial for location-based services, such as POI recommendation. Most of the existing methods can seldom achieve satisfactory performance due to the scarcity of POI's information, which tendentiously confines the recommendation to popular scenic spots, and ignores the unpopular attractions with potentially precious values. In this paper, we propose a novel approach, termed Hierarchical Multi-Clue Fusion (HMCF), for predicting the popularity of POIs. Specifically, we devise an effective hierarchy to comprehensively describe POI by integrating various types of media information (e.g., image and text) from multiple social sources. For each individual POI, we simultaneously inject semantic knowledge as well as multi-clue representative power. We collect a multi-source POI dataset from four widely-used tourism platforms. Extensive experimental results show that the proposed method can significantly improve the performance of predicting the attractions' popularity as compared to several baselines. Yaqian Duan, Xinze Wang, Yang Yang 0002, Zi Huang, Ning Xie 0003, Heng Tao Shen |
SIGIR | 5 |
| 2016 | Research in Web3D Virtual Technology based Online Education Platform for Historical BattleabstractThis paper explores how the reconstruction of special history scenario will be applied in online education. After investigating various virtual reality techniques including design of virtual educational system, reconstruction of virtual scene, management of scene, AI and light shadow rendering, we build an online education platform for touring a web3D virtual battlefield scenario called Huangyangjie in China. We firstly present the solution and scheme for rebuilding the web 3D battlefield Scenario using lightweight 3D models. Secondly, we present voxel of interesting (VOI) scene management strategy. Thirdly, we optimize A* algorithm in AI management process. Finally, we design an experiment for comparing virtual reality based technique teaching mode with traditional teaching mode. Chang Liu 0037, Jinyuan Jia 0002, Ning Xie 0003 |
ICCE | 3 |
| 2016 | Client-Driven Strategy of Large-Scale Scene Streaming
Laixiang Wen, Ning Xie 0003, Jinyuan Jia 0002 |
MMM (2) | 2 |
| 2016 | Lightweighting for Web3D visualization of large-scale BIM scenes in real-time
Ning Xie 0003, Kai Tang 0001, Jinyuan Jia 0002 |
Graph. Model. | 2 |
| 2016 | Interest-driven avatar neighbor-organizing for P2P transmission in distributed virtual worldsabstractAbstract The neighbor table/distributed hash table (DHT) is used to choose the data supplier for data‐dispatching services in distributed virtual environments based on peer‐to‐peer networks. It is essential that a stable and efficient neighbor table/DHT be maintained. Because the avatar has much freedom to roam, the spatial distribution of nodes is not uniform, and the logical topology may change dramatically. Therefore, traditional construction mechanisms, such as the neighbor‐discovery mechanism based on spatial distance or DHT, may involve fierce churn in the neighbor table and frequent message exchanges. In this paper, we proposed a dynamic node‐organizing mechanism that aims to solve these challenging problems by applying the avatar's behavioral characteristics to the neighbor maintenance mechanism and scene data transmission. First, we have summarized the common social behaviors of avatars and extracted their characteristics. We then propose an interest‐similarity measuring algorithm to divide the node into diverse clusters. Next, we measure the cluster stability in terms of interest entropy while constructing a stable neighbor mesh for each node in a cluster. We have conducted extensive simulation experiments that simulate avatar behaviors in a popular massively multiplayer online game. The results show that our proposed mechanism achieved a substantial alleviation of neighbor churn and reduced information exchange, which improves the transmission efficiency in distributed virtual environments. Copyright © 2015 John Wiley & Sons, Ltd. Mingfei Wang, Jinyuan Jia 0002, Ning Xie 0003 |
Comput. Animat. Virtual Worlds | 3 |
| 2016 | Fast accessing Web3D contents using lightweight progressive meshesabstractAbstract Accessing Web3D contents is relatively slow through Internet under limited bandwidth. Preprocessing of 3D models can certainly alleviate the problem, such as 3D compression and progressive meshes (PM). But none of them considers the similarity between components of a 3D model, so that we could take advantage of this to further improve the efficiency. This paper proposes a similarity‐aware data reduction method together with PM, called lightweight progressive meshes (LPM). LPM aims to excavate similar components in a 3D model, generates PM representation of each component left after removing redundant components, and organizes all the processed data using a structure called lightweight scene graph. The proposed LPM possesses four significant advantages. First, it can minimize the file size of 3D model dramatically without almost any precision loss. Because of this, minimal data is delivered. Second, PM enables the delivery to be progressive, so called streaming. Third, when rendering at client side, due to lightweight scene graph, decompression is not necessary and instanced rendering is fully exerted. Fourth, it is extremely efficient and effective under very limited bandwidth, especially when delivering large 3D scenes. Performance on real data justifies the effectiveness of our LPM, which improves the state‐of‐the‐art in accessing Web3D contents. Copyright © 2015 John Wiley & Sons, Ltd. Laixiang Wen, Ning Xie 0003, Jinyuan Jia 0002 |
Comput. Animat. Virtual Worlds | 2 |
| 2016 | 3D tree skeletonization from multiple images based on PyrLK optical flow
Dejia Zhang, Ning Xie 0003, Shuang Liang 0001, Jinyuan Jia 0002 |
Pattern Recognit. Lett. | 2 |
| 2015 | Regularized Policy Gradients: Direct Variance Reduction in Policy Gradient Estimation
Tingting Zhao 0001, Gang Niu 0001, Ning Xie 0003, Jucheng Yang 0001, Masashi Sugiyama |
ACML | 3 |
| 2015 | Web3D-Based Online Walkthrough of Large-Scale Underground ScenesabstractLarge-scale scenes' processing has become the major trend today. We mainly address the online walkthrough of Large-scale underground (UG) scenes in this paper. Taking into account the characteristics of UG scene, we first propose a lightweight preprocessing to optimize the raw UG scene and unify the raw data with scene, sub-scene and simple model. Then we generate a three-layered grid structure for organizing the scene to facilitate the visibility culling and data accessing. Finally, we design two scene management strategies, named SOI-Exterior Shell and Portal-Interior Shell, and integrate our methods in an experimental prototype. The experimental result shows that our method can remove a large amount of redundancies from the raw data, reduce resource consumption greatly and make it possible to walkthrough in large-scale UG scenes online without any web browsers plugins. Ning Xie 0003, Jinyuan Jia 0002 |
DS-RT | 2 |
| 2015 | Stroke-Based Stylization Learning and Rendering with Inverse Reinforcement Learning
Ning Xie 0003, Tingting Zhao 0001, Masashi Sugiyama |
IJCAI | 1 |
| 2015 | Conditional Density Estimation with Dimensionality Reduction via Squared-Loss Conditional Entropy MinimizationabstractRegression aims at estimating the conditional mean of output given input. However, regression is not informative enough if the conditional density is multimodal, heteroskedastic, and asymmetric. In such a case, estimating the conditional density itself is preferable, but conditional density estimation (CDE) is challenging in high-dimensional space. A naive approach to coping with high dimensionality is to first perform dimensionality reduction (DR) and then execute CDE. However, a two-step process does not perform well in practice because the error incurred in the first DR step can be magnified in the second CDE step. In this letter, we propose a novel single-shot procedure that performs CDE and DR simultaneously in an integrated way. Our key idea is to formulate DR as the problem of minimizing a squared-loss variant of conditional entropy, and this is solved using CDE. Thus, an additional CDE step is not needed after DR. We demonstrate the usefulness of the proposed method through extensive experiments on various data sets, including humanoid robot transition and computer art. Voot Tangkaratt, Ning Xie 0003, Masashi Sugiyama |
Neural Comput. | 2 |
| 2012 | Artist Agent: A Reinforcement Learning Approach to Automatic Stroke Generation in Oriental Ink Painting
Ning Xie 0003, Hirotaka Hachiya, Masashi Sugiyama |
ICML | 1 |
| 2011 | Anisotropic Band-Pass Trilateral Filter for Creating Artistic Painting from Real ImageabstractIn this work, we propose a new non-photorealistic rendering approach to the creation of artistic painting from a color image. Our algorithm consists chiefly of two steps: a trilateral filter is firstly applied to the original image for creating drawings vertical to the edges and then a DoG-like band-pass filter is adapted for generating streams along the eigenvectors and therefore smoothes image along stream lines. The proposed trilateral filter in fact is an extension of bilateral filter by considering additional gradient space. On the other hand, DoG-like band-pass filter is computed from eigenvectors and eigenvalues of a tensor matrix calculated at each pixel. Our approach effectively preserves image structures while anisotropically blurring image regions. In regions with low contrast, stream-like structures are also well produced due to gradient relaxation. The experiments demonstrate that the two-steps algorithm can produces good and pleasant visual results. Yuto Yoshida, Ning Xie 0003 |
TrustCom | 3 |
| 2011 | Contour-driven Sumi-e rendering of real photos
Ning Xie 0003, Hamid Laga, Suguru Saito, Masayuki Nakajima 0001 |
Comput. Graph. | 1 |
| 2009 | Contour-driven brush stroke synthesisabstractWe propose in this paper an interactive sketch-based system for simulating oriental brush strokes on complex shapes. We introduce a contour-driven approach where the user inputs contours to represent complex shapes, the system estimates automatically the optimal trajectory of the brush, and then renders them into oriental ink painting. Unlike previous work where the brush trajectory is explicitly specified as input, we automatically estimate this trajectory given the outline of the shape to paint. Existing methods can be classified into: (1) methods that explicitly model a virtual 3D brush and mimic its effect on a paper [Wang and Wang 2007], and (2) methods that simulate the rendering effect on a 2D canvas without an explicit 3D brush model [Okabe et al. 2007]. Our approach falls into the second category. Figure 1 shows four results generated by our algorithm. Ning Xie 0003, Hamid Laga, Suguru Saito, Masayuki Nakajima 0001 |
SIGGRAPH ASIA Sketches | 1 |