Tianxiang Ma

dblp:50/10800 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 10 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
abstract
Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, images, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of modality-complete data and the difficulty of jointly modeling triplet conditions without performance degradation. In this work, we present HuMo, a unified HCVG framework for collaborative multimodal control. For the first challenge, we construct an incomplete-yet-complementary dataset for improved data utilization efficiency and training scalability. For the second challenge, we propose a two-stage progressive multimodal training paradigm with task-specific strategies at each stage. In the first stage, to balance the text-following and subject-preservation abilities, we adopt the minimal-invasive image injection strategy. In the second stage, to enhance audio-visual sync, we propose a focus-by-predicting strategy that implicitly guides the model to associate audio with facial regions. For joint learning of controllabilities across multi-modal inputs, we progressively incorporate the audio-visual sync task, building on previously acquired capabilities. During inference, for flexible and fine-grained multimodal control, we design a stage-adaptive Classifier-Free Guidance strategy that dynamically adjusts guidance weights across denoising steps. Extensive experimental results demonstrate that HuMo surpasses specialized state-of-the-art methods in sub-tasks, establishing a unified framework for collaborative multimodal-conditioned HCVG.
Liyang Chen, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Lijie Liu, Zhiyong Wu 0001
AAAI2
2025 I2VControl: Disentangled and Unified Video Motion Synthesis Control
abstract
Motion controllability is crucial in video synthesis. However, most previous methods are limited to single control types, and combining them often results in logical conflicts. In this paper, we propose a disentangled and unified framework, namely I2VControl, to overcome the logical conflicts. We rethink camera control, object dragging, and motion brush, reformulating all tasks into a consistent representation based on point trajectories, each managed by a dedicated formulation. Accordingly, we propose a spatial partitioning strategy, where each unit is assigned to a concomitant control category, enabling diverse control types to be dynamically orchestrated within a single synthesis pipeline without conflicts. Furthermore, we design an adapter structure that functions as a plug-in for pre-trained models and is agnostic to specific model architectures. We conduct extensive experiments, achieving excellent performance on various control tasks, and our method further facilitates user-driven creative combinations, enhancing innovation and creativity. Project page: https://wanquanf.github.io/I2VControl .
Wanquan Feng, Tianhao Qi, Jiawei Liu 0001, Mingzhen Sun, Pengqi Tu, Tianxiang Ma, Songtao Zhao, SiYu Zhou 0002
ICCV6
2025 Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
abstract
The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts subject elements from reference images and generates subject-consistent videos following textual instructions. We believe that the essence of subject-to-video lies in balancing the dual-modal prompts of text and image, thereby deeply and simultaneously aligning both text and visual content. To this end, we propose Phantom, a unified video generation framework for both single- and multi-subject references. Building on existing text-to-video and image-to-video architectures, we redesign the joint text-image injection model and drive it to learn cross-modal alignment via text-image-video triplet data. The proposed method achieves high-fidelity subject-consistent video generation while addressing issues of image content leakage and multi-subject confusion. Evaluation results indicate that our method outperforms other state-of-the-art closed-source commercial solutions. In particular, we emphasize subject consistency in human generation, covering existing ID-preserving video generation while offering enhanced advantages.
Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen
ICCV2
2025 OptiACL: Optimized Anchor Contrastive Learning Framework for Multimodal Conversational Emotion Recognition
Yujie Guan, Weijie Feng, Tianxiang Ma
ICIC (24)4
2025 I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
abstract
Video generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However, existing camera control methods still suffer from several limitations, including control precision and the neglect of the control for subject motion dynamics. In this work, we propose I2VControl-Camera, a novel camera control method that significantly enhances controllability while providing adjustability over the strength of subject motion. To improve control precision, we employ point trajectory in the camera coordinate system instead of only extrinsic matrix information as our control signal. To accurately control and adjust the strength of subject motion, we explicitly model the higher-order components of the video trajectory expansion, not merely the linear terms, and design an operator that effectively represents the motion strength. We use an adapter architecture that is independent of the base model structure. Experiments on static and dynamic scenes show that our framework outperformances previous methods both quantitatively and qualitatively. Project page: https://wanquanf.github.io/I2VControlCamera.
Wanquan Feng, Jiawei Liu 0001, Pengqi Tu, Tianhao Qi, Mingzhen Sun, Tianxiang Ma, Songtao Zhao, SiYu Zhou 0002
ICLR6
2025 TuplePick: A High Stability Packet Classification based on Neural Network
abstract
Packet classification is one of the crucial components of networking. With the advent of Software Defined Network (SDN), packet classification has become more challenging. So far, the proposed schemes have shown good performance. However, packet classification has different application scenarios, such as access control and firewalls. The distribution characteristics of rulesets vary in different application scenarios, which affects packet classification throughput. To this end, a tuple selection model named Picking Model (PM) is designed in this paper to perform packet matching via a neural network. Moreover, based on PM, a packet classification scheme called TuplePick (TP) is proposed, which enables to pick a possible good tuple rather than an exhaustive search in the tuple space. The experimental results indicate that its throughput variances of different rulesets are less than state-of-the-art schemes, which means it outperforms current schemes on stability of throughput in different application scenarios.
Zhuo Li 0009, Jindian Liu, Yu Zhang 0036, Tianxiang Ma
WoWMoM5
2024 Freestyle 3D-Aware Portrait Synthesis Based on Compositional Generative Priors
Tianxiang Ma, Jianxin Sun 0003, Yingya Zhang, Jing Dong 0003
ICPR (6)1
2023 ReGANIE: Rectifying GAN Inversion Errors for Accurate Real Image Editing
abstract
The StyleGAN family succeed in high-fidelity image generation and allow for flexible and plausible editing of generated images by manipulating the semantic-rich latent style space. However, projecting a real image into its latent space encounters an inherent trade-off between inversion quality and editability. Existing encoder-based or optimization-based StyleGAN inversion methods attempt to mitigate the trade-off but see limited performance. To fundamentally resolve this problem, we propose a novel two-phase framework by designating two separate networks to tackle editing and reconstruction respectively, instead of balancing the two. Specifically, in Phase I, a W-space-oriented StyleGAN inversion network is trained and used to perform image inversion and edit- ing, which assures the editability but sacrifices reconstruction quality. In Phase II, a carefully designed rectifying network is utilized to rectify the inversion errors and perform ideal reconstruction. Experimental results show that our approach yields near-perfect reconstructions without sacrificing the editability, thus allowing accurate manipulation of real images. Further, we evaluate the performance of our rectifying net- work, and see great generalizability towards unseen manipulation types and out-of-domain images.
Bingchuan Li, Tianxiang Ma, Miao Hua, Zili Yi
AAAI2
2023 Semantic 3D-Aware Portrait Synthesis and Manipulation Based on Compositional Neural Radiance Field
abstract
Recently 3D-aware GAN methods with neural radiance field have developed rapidly. However, current methods model the whole image as an overall neural radiance field, which limits the partial semantic editability of synthetic results. Since NeRF renders an image pixel by pixel, it is possible to split NeRF in the spatial dimension. We propose a Compositional Neural Radiance Field (CNeRF) for semantic 3D-aware portrait synthesis and manipulation. CNeRF divides the image by semantic regions and learns an independent neural radiance field for each region, and finally fuses them and renders the complete image. Thus we can manipulate the synthesized semantic regions independently, while fixing the other parts unchanged. Furthermore, CNeRF is also designed to decouple shape and texture within each semantic region. Compared to state-of-the-art 3D-aware GAN methods, our approach enables fine-grained semantic region manipulation, while maintaining high-quality 3D-consistent synthesis. The ablation studies show the effectiveness of the structure and loss function used by our method. In addition real image inversion and cartoon portrait 3D editing experiments demonstrate the application potential of our method.
Tianxiang Ma, Bingchuan Li, Jing Dong 0003, Tieniu Tan
AAAI1
2023 CFFT-GAN: Cross-Domain Feature Fusion Transformer for Exemplar-Based Image Translation
abstract
Exemplar-based image translation refers to the task of generating images with the desired style, while conditioning on certain input image. Most of the current methods learn the correspondence between two input domains and lack the mining of information within the domain. In this paper, we propose a more general learning approach by considering two domain features as a whole and learning both inter-domain correspondence and intra-domain potential information interactions. Specifically, we propose a Cross-domain Feature Fusion Transformer (CFFT) to learn inter- and intra-domain feature fusion. Based on CFFT, the proposed CFFT-GAN works well on exemplar-based image translation. Moreover, CFFT-GAN is able to decouple and fuse features from multiple domains by cascading CFFT modules. We conduct rich quantitative and qualitative experiments on several image translation tasks, and the results demonstrate the superiority of our approach compared to state-of-the-art methods. Ablation studies show the importance of our proposed CFFT. Application experimental results reflect the potential of our method.
Tianxiang Ma, Bingchuan Li, Wei Liu 0035, Miao Hua, Jing Dong 0003, Tieniu Tan
AAAI1
2022 AdaDeId: Adjust Your Identity Attribute Freely
abstract
Face de-identification has drawn increasing attention in recent years. It is important to protect people’s identity information meanwhile keeping the utility of the face data in many computer vision tasks. We propose a Adaptive De-identification (AdaDeId) method, a novel approach that can freely manipulate the identity attributes of given faces. We introduce an identity decoupling representation learning method, which is based on the autoencoder decoupling model as well as our proposed Identity Decoupling Representation (IDR) loss and Content Retention (CR) loss. Our method encodes the identity information of a face into a unit spherical space, where we can continuously manipulate the identity representation vector. Various de-identified faces derived from an original face can be generated through our method and maintain high similarity to the original image contents. Quantitative and qualitative experiments demonstrate our method achieves state-of-the-art on visual quality and de-identification validity.
Tianxiang Ma, Wei Wang 0025, Jing Dong 0003
ICPR1
2022 DesignerGAN: Sketch Your Own Photo
abstract
Person image generation is a challenging problem due to the complexity of human body structure and the richness of clothing texture. Recent works have made great progress on pose transfer by using keypoints, but cannot characterize the personalized shape attributes. Hence, they have limited person image editing ability, especially in respect of shape editing. In this paper, we propose to use sketches as the expression of the target image, which can not only represent the pose and shape simultaneously but is also flexible to manipulate at the semantic level. We propose DesignerGAN, a novel two-stage model for pose transfer and shape-related attributes editing. The first stage predicts the target semantic parsing using the target sketch and obtains parsing feature maps. In the second stage, with the parsing feature maps and the scaled target sketch, we devise a domain-matching spatially-adaptive normalization method to guide target image generation in multi-level. Qualitative and quantitative comparison results demonstrate our method’s superiority over state-of-the-arts on pose transfer. Besides, we achieve flexible person image editing through simple hand-drawings on sketches.
Binghao Zhao, Tianxiang Ma, Bo Peng 0002, Jing Dong 0003
ICPR2
2021 MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image Generation
abstract
Pose-guided person image generation usually involves using paired source-target images to supervise the training, which significantly increases the data preparation effort and limits the application of the models. To deal with this problem, we propose a novel multi-level statistics transfer model, which disentangles and transfers multi-level appearance features from person images and merges them with pose features to reconstruct the source person images themselves. So that the source images can be used as supervision for self-driven person image generation. Specifically, our model extracts multi-level features from the appearance encoder and learns the optimal appearance representation through attention mechanism and attributes statistics. Then we transfer them to a pose-guided generator for re-fusion of appearance and pose. Our approach allows for flexible manipulation of person appearance and pose properties to perform pose transfer and clothes style transfer tasks. Experimental results on the DeepFashion dataset demonstrate our method’s superiority compared with state-of-the-art supervised and unsupervised methods. In addition, our approach also performs well in the wild.
Tianxiang Ma, Bo Peng 0002, Wei Wang 0025, Jing Dong 0003
CVPR1
2011 Interference Alignment-Like Behaviors of MMSE Designs for General Multiuser MIMO Systems
abstract
Interference Alignment (IA) transceiver designs are currently of great interest to the community. They are, however, limited to certain configurations. This paper seeks to show that for general multiuser MIMO systems, a) there is a relationship between MMSE designs and IA, and b) MMSE designs should be used instead of IA ones. The former is done by proving that MMSE designs naturally have IA-like behaviors. The latter is done using several arguments. One of them is based on the MIMO X network numerical results involving the generalized iterative approach (GIA) and the previously proposed MMSE-IA. In the simulation, the hybrid IA approach is clearly outperformed by the purely MMSE based GIA. The GIA is also run for the MIMO M network.
Enoch Lu, Tianxiang Ma, I-Tai Lu
GLOBECOM2