VLDB 2026 Research / reviewers in the wild / expert
Naye Ji
dblp:159/9836
· DBLP profile ↗
13ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-6986-3766ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ConstFS: Controlled and Stable Face Stylization with High Identity-PreservedabstractFacial style transfer often faces challenges, including content loss, style degradation, and difficulty balancing style and content consistency. Existing methods, including those based on StyleGAN and Stable Diffusion, encounter issues such as facial feature distortion, style degradation, and color leakage. We propose ConstFS (Controlled and Stable Face Stylization) to overcome these challenges, which enhances traditional diffusion models by incorporating advanced style extraction, identity preservation, and dynamic control mechanisms. The proposed ConstFS achieves superior image control, enabling it to adaptively match the style image. The experimental results demonstrate that our method significantly enhances style integrity and content consistency, outperforming existing techniques in both qualitative and quantitative experiments. Feichi Chen, Xiaokang Chen, Xuanhao Lou, Naye Ji |
VINCI | 4 |
| 2025 | Learning Multi-Grained Interpretable Latent Representation for 3D Face ManipulationabstractRepresenting 3D faces using generative models has been investigated for several years for its numerous applications in computer vision and graphics. However, general 3D face manipulation is often limited by the lack of multi-level interpretability of the latent space in 3D generative models. To address this problem, we propose a novel generative approach dubbed hierarchically semantic regularized variational auto-encoders (HSR-VAE), which explicitly endows latent variables with multi-grained semantics of the synthesized 3D face shapes. Specifically, to accommodate the hierarchical structure of the human face, we decompose the latent space to represent variations in facial features at different scales, from local facial segments to fine-grained attributes. Moreover, part-aware and attribute-aware semantic regularizers are introduced to establish a linkage between hierarchically organized latent variables and multi-grained facial semantics, allowing more interpretable and meaningful representations of the 3D face. Extensive quantitative and qualitative experiments show the effectiveness of HSR-VAE and demonstrate that it can provide a more interpretable, manipulable, and generalizable latent representation than current approaches, facilitating a wide range of 3D face shape manipulation tasks. Wen Gao 0001, Naye Ji, Dingguo Yu |
Comput. Vis. Media | 2 |
| 2025 | Relationship-Incremental Scene Graph Generation by a Divide-and-Conquer Pipeline With Feature AdapterabstractAs a challenging computer vision task, Scene Graph Generation (SGG) finds the latent semantic relationships among objects from a given image, which may be limited by the datasets and real-world scenarios. In this paper, we consider a novel incremental learning task called Relationship-Incremental Scene Graph Generation (RISGG) that learns the semantic relationships among objects in an incremental way. Compared with classic Class-Incremental Learning (CIL) problem, RISGG suffers from its special issues: 1) Old class shift - the relationship-labeled object pair may have different labels during different learning sessions; 2) Background shift - the relationship-unlabeled object pair may not be a real unlabeled one. In this work, we address the above issues from the following aspects. First, we present a Divide-and-Conquer (DaC) pipeline to deal with the old class shift via decoupling the recognition of relationship classes and recognizing relationships individually. In this way, label confusion and interaction among different relationships are eliminated during training. Second, we propose a Feature Adapter (FA) to bridge the feature space gap between the current session and the previous one and use our extra supervision to mine old relationship information in the current session. Our proposed network combined DaC and FA, abbreviated DaCFA-Net, for RISGG. Experimental results on the benchmark dataset demonstrate the significant performance gain of DaCFA-Net in RISGG. It gains about 20% improvement against the SGG baselines on the popular VG dataset. Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Naye Ji, Xi Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | GAMA-Pose: Graph-Aware Multi-Representation Aggregation for 3D Human Pose EstimationabstractMonocular 3D human pose estimation presents a considerable challenge owing to the intrinsic depth ambiguity associated with single-camera observations. Existing methods primarily rely on mean per joint position error (MPJPE) loss to train models for the conversion from 2D to 3D coordinates. However, empirical analysis reveals that models trained solely with point-based supervision may produce biomechanically implausible poses or exhibit significant depth ambiguity, even when achieving low MPJPE. This limitation arises from the fact that point-based loss only considers individual joint locations without accounting for inter-joint relationships. Fortunately, edges of human pose encode critical prior knowledge, including skeleton connectivity and biomechanical distributions. Explicitly modeling edge representations enables the model to overcome the constraints associated with point-only approaches, reducing the uncertainty in the optimization process of the 2D-3D inverse mapping and directly constraining depth ambiguity. Therefore, we propose the Graph-Aware Multi-Representation Aggregation (GAMA-Pose) framework that jointly predicts points and edges, with their fusion serving as the final output. To ensure the accuracy of edge predictions and mitigate depth ambiguity, Anti-Depth-Ambiguity Loss (ADA-Loss) is introduced to supervise the properties of edges and give direct supervision on depth ambiguity. Correspondingly, edge-based metrics are proposed to quantify the error of predicted edges. Experiments conducted on Human3.6M and MPI-INF-3DHP datasets demonstrate that GAMA-Pose effectively addresses the limitations of models relying solely on point constraints, mitigates depth ambiguity, enhances the accuracy of both point and edge predictions, and achieves state-of-the-art (SOTA) performance on both datasets. Songran Zhou, Xuewei Li 0003, Xiubo Liang, Naye Ji, Xi Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Speech-Driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature InferenceabstractSpeech-driven gesture generation is an emerging field within virtual human creation. However, a significant challenge lies in accurately determining and processing the multitude of input features (such as acoustic, semantic, emotional, personality, and even subtle unknown features). Traditional approaches, reliant on various explicit feature inputs and complex multimodal processing, constrain the expressiveness of resulting gestures and limit their applicability. To address these challenges, we present Persona-Gestor, a novel end-to-end generative model designed to generate highly personalized 3D full-body gestures solely relying on raw speech audio. The model combines a fuzzy feature extractor and a non-autoregressive Adaptive Layer Normalization (AdaLN) transformer diffusion architecture (DiTs-based). The fuzzy feature extractor harnesses a fuzzy inference strategy that automatically infers implicit, continuous fuzzy features. These fuzzy features, represented as a unified latent feature, are fed into the AdaLN transformer. The AdaLN transformer introduces a conditional mechanism that applies a uniform function across all tokens, thereby effectively modeling the correlation between the fuzzy features and the gesture sequence. This module ensures a high level of gesture-speech synchronization while preserving naturalness. Finally, we employ the diffusion model to train and infer various gestures. Extensive subjective and objective evaluations on the Trinity, ZEGGS, and BEAT datasets confirm our model's superior performance to the current state-of-the-art approaches. Persona-Gestor improves the system's usability and generalization capabilities, setting a new benchmark in speech-driven gesture synthesis and broadening the horizon for virtual human technology. Fan Zhang 0105, Zhaohan Wang, Xin Lyu 0004, Mengjian Li, Weidong Geng, Naye Ji, Fuxing Gao, Hao Wu 0141, Shunman Li |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | DiffMotion: Speech-Driven Gesture Synthesis Using Denoising Diffusion Model
Fan Zhang 0105, Naye Ji, Fuxing Gao, Yongping Li |
MMM (1) | 2 |
| 2023 | Adaptive cooperative exploration for reinforcement learning from imperfect demonstrations
Fuxian Huang, Naye Ji, Huajian Ni, Shijian Li, Xi Li 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Memory-efficient distribution-guided experience sampling for policy consolidation
Fuxian Huang, Weichao Li 0003, Yining Lin, Naye Ji, Shijian Li, Xi Li 0001 |
Pattern Recognit. Lett. | 4 |
| 2021 | NBA Basketball Video Summarization for News Report via Hierarchical-Grained Deep Reinforcement Learning
Naye Ji, Dingguo Yu, Youbing Zhao |
ICIG (3) | 1 |
| 2018 | Design and Evaluation of Tactile Number Reading Methods on SmartphonesabstractIn this paper, we propose two active tactile texture interaction methods enabling eye-free tactile number (0-9) reading on a variable-friction surface for a smartphone equipped with a TPad device. We employ four discriminative stripe-based tactile textures to represent different numbers, namely, 0, 1, 2, and 5, and combine them with the other numbers. For the first interaction method, we adopt simple left/right slider motion gestures (LRSM), and for the second, we utilize up to down slider motion gestures (UToDSM). We examined the interactive efficiency, recognition accuracy, and user-subjective satisfaction of both techniques. We also compared the two methods with an approach that converts a variable-friction perception into a typical vibrotactile perception. Our results indicate that UToDSM leads to the highest interactive efficiency and achieves a high recognition accuracy. Fan Zhang 0105, Shaowei Chu, Naye Ji, Ruifang Pan |
ICIS | 3 |
| 2017 | Pan-and-tilt self-portrait system using gesture interfaceabstractDigital cameras are widely used in desktop and notebook PCs. Taking self-portraits is one of the important function of such cameras, which allows users to capture memories, create art, and improve photography techniques. A desktop environment with a large display and a pan-and-tilt camera provides users with a good area for exploring more angles and postures while taking self-portraits. However, most of the existing camera interfaces of this type are limited to device-based systems (i.e., mouse and keyboard) that prevent users from efficiently controlling the camera while taking self-portraits. This study proposes a vision-based system equipped with a gesture interface that control a pan-and-tilt camera for taking self-portraits. This interface uses gestures, particularly slight hand movements (i.e., sweeps, circles, and waves), to control the pan, tilt, and shutter functions of the camera. The gesture-recognition achieved good efficiency in performance (less than 2ms) and the recognition rate (0.9 on average in lighting conditions range 100 - 200). Experimental results indicate that the proposed system effectively controls the options in a self-portrait camera, this approach provides significantly higher satisfaction, particularly in terms of the intuitive motion gestures, freedom, and enjoyment, than when using a hand-held remote control or a conventional mouse-based interface. The proposed system is a promising technique for taking self-portraits in a desktop environment. Shaowei Chu, Fan Zhang 0105, Naye Ji, Zhefan Jin, Ruifang Pan |
ICIS | 3 |
| 2017 | Double hand-gesture interaction for walk-through in VR environmentabstractIn this paper, we present a double hand-gesture interaction (DHGI) method for walk-through in VR environment with an Oculus Rift headset and Leap Motion function. The user can control the avatar (first-person view) to move (walk/run) forward or backward by turning the user's left palm upward or downward, and by turning the avatar to the left or right with the right thumb pointing toward either direction. Compared with the results of the joystick input device and portal method using Oculus Rift Touches, the objective and subjective findings of this study indicate that DHGI is intuitive, easy to learn, easy to use, and causes low fatigue. Moreover, the user feedback shows that DHGI significantly improves immersion and reduces the sense of motion sickness in VR. Fan Zhang 0105, Shaowei Chu, Ruifang Pan, Naye Ji, Lian Xi |
ICIS | 4 |
| 2011 | Local Regression Model for Automatic Face Sketch GenerationabstractAs one of the important artistic styles of portrait, sketch portrait has wide applications for both digital entertainment and law enforcement. In this paper, an automatic face sketch generation approach is presented by learning from photo-sketch pair examples. Specifically, the relationship between a face photo and its corresponding face sketch is learned on image patch level. By applying this relationship to the input face photo patch, we can infer the output face sketch patch by exploiting some regression techniques such as kNN, the Lasso and so on. Via our local regression model, we can synthesize an appealing sketch portrait from a given face photo in a few minutes. Experiments conducted on CUHK database have shown that our results are more compelling than previous methods especially in two respects: (1) our synthesized sketches preserve more identity information of the original face photo, (2) our synthesized sketches presents more pencil sketch texture. Naye Ji, Xiujuan Chai, Shiguang Shan, Xilin Chen 0001 |
ICIG | 1 |