Bin Ji 0004

dblp:119/1943-4 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-8981-5251ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RetailSQL: A Chinese Text-to-SQL Dataset for Retail Industry
Bin Ji 0004, Jianhua Li 0009
ICIC (24)2
2025 POMP: Physics-constrainable Motion Generative Model through Phase Manifolds
abstract
Numerous researches on real-time motion generation primarily focus on kinematic aspects, often resulting in physically implausible outcomes. In this paper, we present POMP ("Physics-cOnstrainable Motion Generative Model through Phase Manifolds"), a kinematics-based framework that synthesizes physically realistic motions by leveraging phase manifolds to align motion priors with physics constraints. POMP operates as a frame-by-frame autoregressive model with three core components: a diffusion-based kinematic module, a simulation-based dynamic module, and a phase encoding module. At each timestep, the kinematic module first generates an initial pose, which is subsequently revised by the dynamic module through a simulation step to incorporate physical constraints. While individual simulation steps induce negligible kinematic distortion, accumulated discrepancies can drive the result beyond the motion prior learned by the kinematic module, leading to failure in subsequent motion generation. To address this, the phase encoding module applies semantic alignment in the phase manifold, projecting the simulated result back to the motion prior. Moreover, we present a pipeline in Unity for generating terrain maps and capturing full-body motion impulses from existing motion capture dataset. The collected terrain topology and motion impulse data facilitate the training of POMP, enabling it to robustly respond to underlying contact forces and applied dynamics. Extensive evaluations demonstrate the efficacy of POMP across various tasks.
Bin Ji 0004, Zhimeng Liu, Shuai Tan 0002, Xiaogang Jin 0001, Xiaokang Yang 0001
CVPR1
2025 FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
abstract
Talking head generation is gaining significant importance across various domains, with a growing demand for high-quality rendering. However, existing methods often suffer from identity leakage (IL) and rendering artifacts (RA), particularly in extreme cases. Through an in-depth analysis of previous approaches, we identify two key insights: (1) IL arises from identity information embedded within motion features, and (2) this identity information can be leveraged to address RA. Building on these findings, this paper introduces FixTalk, a novel framework designed to simultaneously resolve both issues for high-quality talking head generation. Firstly, we propose an Enhanced Motion Indicator (EMI) to effectively decouple identity information from motion features, mitigating the impact of IL on generated talking heads. To address RA, we introduce an Enhanced Detail Indicator (EDI), which utilizes the leaked identity information to supplement missing details, thus fixing the artifacts. Extensive experiments demonstrate that FixTalk effectively mitigates IL and RA, achieving superior performance compared to state-of-the-art methods.
Shuai Tan 0002, Bill Gong, Bin Ji 0004
ICCV3
2025 MaskTalker: Audio-Driven Talking Head Generation from Masked Face Using StyleGAN
Shuai Tan 0002, Bin Ji 0004, Chuhang Ma
ICXR2
2025 SPORT: From Zero-Shot Prompts to Real-Time Motion Generation
abstract
Real-time motion generation has garnered significant attention within the fields of computer animation and gaming. Existing methods typically realize motion control via isolated style or content labels, resulting in short, simply motion clips. In this paper, we propose a motion generation framework, called SPORT ("from zero-Shot Prompt tO Real-Time motion generation"), for generating real-time and ever-changing motions using zero-shot prompts. SPORT consists of three primary components: (1) a body-part phase autoencoder that ensures smooth transitions between diverse motions; (2) a body-part content encoder that mitigates semantic gap between texts and motions; (3) a diffusion-based decoder that accelerates the denoising process while enhancing the diversity and realism of motions. Moreover, we develop a prototype for real-time application in Unity, demonstrating that our approach effectively considering the semantic gap caused by abstract style texts and rapidly changing terrains. Through qualitative and quantitative comparisons, we show that SPORT outperforms other approaches in terms of motion quality, style diversity and inference speed.
Bin Ji 0004, Zhimeng Liu, Shuai Tan 0002, Xiaokang Yang 0001
IEEE Trans. Vis. Comput. Graph.1
2024 Say Anything with Any Style
abstract
Generating stylized talking head with diverse head motions is crucial for achieving natural-looking videos but still remains challenging. Previous works either adopt a regressive method to capture the speaking style, resulting in a coarse style that is averaged across all training data, or employ a universal network to synthesize videos with different styles which causes suboptimal performance. To address these, we propose a novel dynamic-weight method, namely Say Anything with Any Style (SAAS), which queries the discrete style representation via a generative model with a learned style codebook. Specifically, we develop a multi-task VQ-VAE that incorporates three closely related tasks to learn a style codebook as a prior for style extraction. This discrete prior, along with the generative model, enhances the precision and robustness when extracting the speaking styles of the given style clips. By utilizing the extracted style, a residual architecture comprising a canonical branch and style-specific branch is employed to predict the mouth shapes conditioned on any driving audio while transferring the speaking style from the source to any desired one. To adapt to different speaking styles, we steer clear of employing a universal network by exploring an elaborate HyperStyle to produce the style-specific weights offset for the style branch. Furthermore, we construct a pose generator and a pose codebook to store the quantized pose representation, allowing us to sample diverse head motions aligned with the audio and the extracted style. Experiments demonstrate that our approach surpasses state-of-the-art methods in terms of both lip-synchronization and stylized expression. Besides, we extend our SAAS to video-driven style editing field and achieve satisfactory performance as well.
Shuai Tan 0002, Bin Ji 0004, Yu Ding 0001
AAAI2
2024 Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art Style
abstract
Although automatically animating audio-driven talking heads has recently received growing interest, previous efforts have mainly concentrated on achieving lip synchronization with the audio, neglecting two crucial elements for generating expressive videos: emotion style and art style. In this paper, we present an innovative audio-driven talking face generation method called Style2Talker. It involves two stylized stages, namely Style-E and Style-A, which integrate text-controlled emotion style and picture-controlled art style into the final output. In order to prepare the scarce emotional text descriptions corresponding to the videos, we propose a labor-free paradigm that employs large-scale pretrained models to automatically annotate emotional text labels for existing audio-visual datasets. Incorporating the synthetic emotion texts, the Style-E stage utilizes a large-scale CLIP model to extract emotion representations, which are combined with the audio, serving as the condition for an efficient latent diffusion model designed to produce emotional motion coefficients of a 3DMM model. Moving on to the Style-A stage, we develop a coefficient-driven motion generator and an art-specific style path embedded in the well-known StyleGAN. This allows us to synthesize high-resolution artistically stylized talking head videos using the generated emotional motion coefficients and an art style source picture. Moreover, to better preserve image details and avoid artifacts, we provide StyleGAN with the multi-scale content features extracted from the identity image and refine its intermediate feature maps by the designed content encoder and refinement network, respectively. Extensive experimental results demonstrate our method outperforms existing state-of-the-art methods in terms of audio-lip synchronization and performance of both emotion style and art style.
Shuai Tan 0002, Bin Ji 0004
AAAI2
2024 FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and Quantization
abstract
Generating emotional talking faces is a practical yet challenging endeavor. To create a lifelike avatar, we draw upon two critical insights from a human perspective: 1) The connection between audio and the non-deterministic facial dynamics, encompassing expressions, blinks, poses, should exhibit synchronous and one-to-many mapping. 2) Vibrant expressions are often accompanied by emotion-aware high-definition (HD) textures and finely detailed teeth. However, both aspects are frequently overlooked by existing methods. To this end, this paper proposes using normalizing Flow and Vector-Quantization modeling to produce emotional talking faces that satisfy both insights concurrently (FlowVQTalker). Specifically, we develop a flow-based coefficient generator that encodes the dynamics of facial emotion into a multi-emotion-class latent space represented as a mixture distribution. The generation process commences with random sampling from the modeled distribution, guided by the accompanying audio, enabling both lip-synchronization and the uncertain nonverbal facial cues generation. Furthermore, our designed vector-quantization image generator treats the creation of expressive facial images as a code query task, utilizing a learned codebook to provide rich, high-quality textures that enhance the emotional perception of the results. Extensive experiments are conducted to showcase the effectiveness of our approach.
Shuai Tan 0002, Bin Ji 0004
CVPR2
2024 EDTalk: Efficient Disentanglement for Emotional Talking Head Synthesis
Shuai Tan 0002, Bin Ji 0004, Mengxiao Bi
ECCV (6)2
2024 StyleVR: Stylizing Character Animations With Normalizing Flows
abstract
The significance of artistry in creating animated virtual characters is widely acknowledged, and motion style is a crucial element in this process. There has been a long-standing interest in stylizing character animations with style transfer methods. However, this kind of models can only deal with short-term motions and yield deterministic outputs. To address this issue, we propose a generative model based on normalizing flows for stylizing long and aperiodic animations in the VR scene. Our approach breaks down this task into two sub-problems: motion style transfer and stylized motion generation, both formulated as the instances of conditional normalizing flows with multi-class latent space. Specifically, we encode high-frequency style features into the latent space for varied results and control the generation process with style-content labels for disentangled edits of style and content. We have developed a prototype, StyleVR, in Unity, which allows casual users to apply our method in VR. Through qualitative and quantitative comparisons, we demonstrate that our system outperforms other methods in terms of style transfer as well as stochastic stylized motion generation.
Bin Ji 0004, Yichao Yan, Ruizhao Chen, Xiaokang Yang 0001
IEEE Trans. Vis. Comput. Graph.1
2023 EMMN: Emotional Motion Memory Network for Audio-driven Emotional Talking Face Generation
abstract
Synthesizing expression is essential to create realistic talking faces. Previous works consider expressions and mouth shapes as a whole and predict them solely from audio inputs. However, the limited information contained in audio, such as phonemes and coarse emotion embedding, may not be suitable as the source of elaborate expressions. Besides, since expressions are tightly coupled to lip motions, generating expression from other sources is tricky and always neglects expression performed on mouth region, leading to inconsistency between them. To tackle the issues, this paper proposes Emotional Motion Memory Net (EMMN) that synthesizes expression overall on the talking face via emotion embedding and lip motion instead of the sole audio. Specifically, we extract emotion embedding from audio and design Motion Reconstruction module to decompose ground truth videos into mouth features and expression features before training, where the latter encode all facial factors about expression. During training, the emotion embedding and mouth features are used as keys, and the corresponding expression features are used as values to create key-value pairs stored in the proposed Motion Memory Net. Hence, once the audio-relevant mouth features and emotion embedding are individually predicted from audio at inference time, we treat them as a query to retrieve the best-matching expression features, performing expression overall on the face and thus avoiding inconsistent results. Extensive experiments demonstrate that our method can generate high-quality talking face videos with accurate lip movements and vivid expressions on unseen subjects.
Shuai Tan 0002, Bin Ji 0004
ICCV2
2022 Poxture: Human Posture Imitation Using Neural Texture
abstract
Human pose imitation, which aims to generate an image with a source character’s appearance, the source character’s shape, and a target character’s posture, has many potential applications in virtual reality, augmented reality, games, movies, etc. It is incredibly challenging due to non-rigid human body motions, significant variations in clothing textures, and self-occluded human bodies in 2D images. In this paper, we propose Poxture, a novel human posture imitation method with neural texture, to address the challenges mentioned above. Concretely, first, we build a dense mapping between a source SMPL human body model (shape and posture) and its corresponding texture (appearance). Then, we apply a neural texture generator to recover the complete texture of the source character. At last, we wrap the source neural texture to the source SMLP model with a target pose to generate the desired image by a GAN model. Poxture does not require any annotations, and our framework can fully disentangle the source character’s appearance, shape, and pose, which enjoys several advantages: 1) It can synthesize high-resolution images with detailed textures, thanks to the learned neural textures containing both visible and invisible parts and high-frequency information; 2) It can imitate complex actions with various appearances and body figures since the complete texture of the source character is acquired. We compare our method with previous methods, showing state-of-the-art results on two challenging benchmarks. Extensive experiments demonstrate that, given any character, our method can manipulate this avatar imitating arbitrary posture.
Chen Yang 0023, Zanwei Zhou, Bin Ji 0004, Guangtao Zhai, Wei Shen 0002
IEEE Trans. Circuits Syst. Video Technol.4
2022 Early Screening of Autism in Toddlers via Response-To-Instructions Protocol
abstract
Early screening of autism spectrum disorder (ASD) is crucial since early intervention evidently confirms significant improvement of functional social behavior in toddlers. This article attempts to bootstrap the response-to-instructions (RTIs) protocol with vision-based solutions in order to assist professional clinicians with an automatic autism diagnosis. The correlation between detected objects and toddler's emotional features, such as gaze, is constructed to analyze their autistic symptoms. Twenty toddlers between 16-32 months of age, 15 of whom diagnosed with ASD, participated in this study. The RTI method is validated against human codings, and group differences between ASD and typically developing (TD) toddlers are analyzed. The results suggest that the agreement between clinical diagnosis and the RTI method achieves 95% for all 20 subjects, which indicates vision-based solutions are highly feasible for automatic autistic diagnosis.
Zhiyong Wang 0009, Bin Ji 0004, Jingxin Deng, Xiu Xu, Honghai Liu 0001
IEEE Trans. Cybern.4
2021 HPOF: 3D Human Pose Recovery from Monocular Video with Optical Flow
abstract
This paper introduces HPOF, a novel deep neural network to reconstruct the 3D human motion from a monocular video. Recently, model-based methods have been proposed to simplify the reconstruction task by estimating several parameters that control a deformable surface model to fit the person in the image. However, learning the parameters from a single image is a highly ill-posed problem, and the process is ultimately data-hungry. Existing 3D datasets are not sufficient, and the usage of 2D in-the-wild datasets is often susceptible to the inadequate precision of manual annotations. To address the above issues, our method yields substantial improvements in two domains. First, we leverage optical flow to supervise the 2D rendered images of predicted SMPL models to learn short-term temporal features. Besides, taking long-term temporal consistency into account, we define a novel temporal encoder based on a dilated convolutional network. The encoder decomposes the learning process of human shape and pose, first guarantees the invariance of the body shape, and then simulates a more reasonable forward kinematics process on this basis to achieve more accurate pose estimation. In addition, an adversarial learning framework is applied to supervise the reconstruction progress in a coarse-grained way. We show that HPOF not only improves the accuracy of 3D poses but ensures the realistic body structure throughout the video. We perform extensive experimentation to demonstrate the superiority of our method and analyze the effectiveness of our model, surpassing other state-of-the-arts.
Bin Ji 0004, Chen Yang 0023
ICMR1
2020 Free-Head Pose Estimation under Low-Resolution Scenarios
abstract
Head pose offers vital cues to infer one's social attention in wide applications. Most existing head pose estimation algorithms have demonstrated competitive results taking high resolution images of frontal view as input. However, these approaches still work poorly if they are fed with low-resolution images. In a more common realistic scene such as computer vision assisted autism screening, images of unconstrained patients that are taken from distant cameras often have low-resolution and non-frontal faces. To this end, we present a multi-view scheme based CNN for free-head pose classification. Residual Networks are taken as the backbone model to generate effective feature representations of the low-resolution images. A novel multi-view feature fusion layer is proposed to address facial appearance variation over multi perspectives owing to free movement. Also, a multi loss function by combining binned classification and regression losses of different pose angles is employed to obtain a more precise pose estimation. The proposed method is evaluated on two scenarios: (1) a publicly available dataset and (2) a practical application to explore the social attention of a group with social deficits: autistic children. Experimental results suggest that our method significantly outperforms state-of-the-art multi-view pose classification methods and achieves comparable pose estimation results. Moreover, the proposed method can be extended to applications of quantitative analysis of social deficits.
Zhiyong Wang 0009, Haibo Qin, Bin Ji 0004, Honghai Liu 0001
SMC5
2020 An Auxiliary Screening System for Autism Spectrum Disorder Based on Emotion and Attention Analysis
abstract
The screening and diagnosis of Autism Spectrum Disorder(ASD) suffer from great challenges due to insufficient professional clinicians and complex procedures. It is urgent to introduce an effective auxiliary system in the diagnosis and treatment process to assist in the completion of pathological information collection tasks, consequently simplifying the screening method and improving the accuracy of screening. We propose a computer vision-based early screening system for ASD to characterize the facial expressions and eye gaze attention which are considered to remarkable indicators for early screening of autism. The system provides the subjects with three different virtual interaction modes: video, picture, and virtual interactive game. During the interaction between the subject and the computer, the system extracts and analyzes the quantitative information of the subject's performance. Then, through computer vision-based emotion analysis and attention analysis methods, the subject's emotions and attention features in the three interaction modes are automatically calculated to assist in the early screening of autism. Finally, the accuracy and feasibility of the system are verified through experiments on both the publicly available dataset and the data collected from 10 ASD children.
Bin Ji 0004, Zhiyong Wang 0009, Honghai Liu 0001
SMC2