EDBT 2026 Demo / reviewers in the wild / expert
Qixuan Zhang
dblp:229/0579
· DBLP profile ↗
29ranked-venue papers
7as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 19 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Artificial intelligence for virtual reality: a review
Lili Wang 0006, Yebin Liu, Miao Wang 0004, Xubo Yang, Lan Xu 0003, Zhangyao Tan, Runze Fan, Hongwen Zhang 0001, Yijian Wen, Haozhong Yang, Jian Wu 0033, Jiahui Fan, Hui Wang 0045, Qixuan Zhang, Yongtian Wang, Qinping Zhao |
Sci. China Inf. Sci. | 18 |
| 2026 | A whole-life fatigue crack growth rate prediction method based on active learning and physics-informed loss
Qixuan Zhang, Wei Zhang 0021, Rui Huang 0001, Xinghui Chen, Changyu Zhou |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Learning a Delighting Prior for Facial Appearance Capture in the WildabstractHigh-quality facial appearance capture has traditionally required costly studio recording. Recent works consider an in-the-wild smartphone-based setup; however, their model-based inverse rendering paradigm struggles with the complex disentanglement of reflectance from unknown illumination. To bridge this gap, we propose to shift the paradigm into training a powerful delighting network as a prior to constrain the optimization. We leverage the OLAT dataset and the rendered Light Stage scans for training, and propose Dataset Latent Modulation (DLM) to seamlessly integrate these heterogeneous data sources. Specifically, by conditioning the core network on learnable source-aware tokens, we decouple dataset-specific styles from physical delighting principles, enabling the emergence of a delighting prior that outperforms existing proprietary models. This powerful delighting prior enables a simple and automatic appearance capture pipeline that achieves high-quality reflectance estimation from casual video inputs, outperforming prior arts by a large margin. Furthermore, we leverage our appearance capture method to transform the multi-view NeRSemble dataset into NeRSemble-Scan, a large-scale collection of 4K-resolution relightable scans. By open-sourcing our model and the NeRSemble-Scan dataset, we democratize high-end facial capture and provide a new foundation for the research community to build photorealistic digital humans. Xin Ming, Zhuofan Shen, Qixuan Zhang, Lan Xu 0003, Feng Xu 0005 |
ACM Trans. Graph. | 5 |
| 2026 | ACT: A Unified Framework for Rigging and Animating Characters with Arbitrary TopologiesabstractRecent advances in generative models have democratized the creation of high-quality static 3D assets, yet animating these meshes remains a labor-intensive bottleneck. Traditional pipelines fracture this process into sequential stages—rigging, skinning, and motion synthesis—ignoring the inherent coupling between morphological structure and motor function. To bridge this gap, we introduce ACT, a unified generative framework that reformulates rigging and animation not as independent tasks, but as complementary views of a single hyper-kinematic process. Our key insight is to model the joint distribution of skeletal topology and temporal motion within a shared latent space. ACT utilizes a Vision Language Model (VLM) to extract semantic topological priors from arbitrary meshes, which then condition a Diffusion Transformer (DiT) backbone. By treating static rest poses and dynamic trajectories as a unified sequence, our model employs a task-aware masking strategy to flexibly perform zero-shot rigging, text-guided motion generation, and motion completion within a single end-to-end architecture. Furthermore, a geometry-guided decoder ensures that surface deformations are tightly coupled with the generated kinematics. Extensive experiments demonstrate that ACT generalizes robustly to diverse, non-humanoid characters without retraining. By replacing brittle cascaded pipelines with a holistic prior, our method enables novel applications such as semantic-driven topology editing and generative in-betweening, offering a versatile and efficient solution for automating 3D character animation. Pengyu Long, Weirui Wang, Qingcheng Zhao, Qixuan Zhang, Jiaqing Zhou, Tianlei Hu, Wei Yang 0034, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 6 |
| 2026 | GenPIE: A Time-Resolved Plenoptic ImagerabstractCapturing the full plenoptic light transport across spatial, angular, and temporal dimensions has long been a pursuit in computational imaging, yet it remains fundamentally constrained by the high dimensionality of the sampling space and the physical inaccessibility of scene regions due to self-occlusions. While time-resolved imaging records the temporal axis, existing methods are bottlenecked by the combinatorial complexity of the plenoptic function. This high dimensionality makes dense omni-dimensional sampling physically prohibitive. Simultaneously, tight coupling between illumination and viewpoint in current systems also precludes the full acquisition of plenoptic light transport. In this work, we present GenPIE, a Generative Plenoptic Imager designed to bridge the gap between sparse physical observations and high-dimensional light transport. We introduce a decoupled laser-detector hardware setup that enables independent control over illumination and detection, allowing for active probing of indirect light paths. To overcome the ill-posedness of sparse sampling and physical blind spots, we propose a generative inverse transient rendering framework. Our approach leverages 3D foundation models to provide strong semantic and 3D geometric priors for initialization, which are subsequently refined through a differentiable transient path tracer to ensure physically grounded adherence to the Transient Rendering Equation. We demonstrate that GenPIE supports a range of applications that are challenging for steady-state or purely neural methods, including disentangling multi-bounce light transport directly from captured transient videos, time unwarping, and time-resolved relighting. The project page is at https://wangzh1.github.io/GenPIE. Huanyu Xu, Kaichun Qiao, Longwen Zhang, Qixuan Zhang, Qilin Sun 0001, Jingyi Yu 0001 |
ACM Trans. Graph. | 6 |
| 2026 | Strips as Tokens: Artist Mesh Generation with Native UV SegmentationabstractRecent advancements in autoregressive transformers have demonstrated remarkable potential for generating artist-quality meshes. However, the token ordering strategies employed by existing methods typically fail to meet professional artist standards, where coordinate-based sorting yields inefficiently long sequences, and patch-based heuristics disrupt the continuous edge flow and structural regularity essential for high-quality modeling. To address these limitations, we propose Strips as Tokens ( SATO ), a novel framework with a token ordering strategy inspired by triangle strips. By constructing the sequence as a connected chain of faces that explicitly encodes UV boundaries, our method naturally preserves the organized edge flow and semantic layout characteristic of artist-created meshes. A key advantage of this formulation is its unified representation, enabling the same token sequence to be decoded into either a triangle or quadrilateral mesh. This flexibility facilitates joint training on both data types: large-scale triangle data provides fundamental structural priors, while high-quality quad data enhances the geometric regularity of the outputs. Extensive experiments demonstrate that SATO consistently outperforms prior methods in terms of geometric quality, structural coherence, and UV segmentation. Rui Xu 0016, Dafei Qin, Kaichun Qiao, Qiujie Dong, Huaijin Pi, Qixuan Zhang, Longwen Zhang, Lan Xu 0003, Jingyi Yu 0001, Wenping Wang 0001, Taku Komura |
ACM Trans. Graph. | 6 |
| 2025 | JMTF: A Joint Model for Chinese Measurable Quantitative Information Extraction based on Table FillingabstractRecently, measurable quantitative information extraction from unstructured texts has attracted increasing attention in various fields of industry. However, due to the issues of error propagation and insufficient deep interactions between entities and relations, the accurate extraction of measurable quantitative information remains as a challenging task. To address these issues, this paper proposes a joint model based on table filling (JMTF) for measurable quantitative information extraction task. The core of this model introduces a co-attention mechanism to achieve bidirectional interactions between recognition and association subtasks, thereby avoiding problems such as feature confusion or insufficient interaction. Additionally, the model employs an optimized BERT-based encoder (TEncoder) for text encoding. TEncoder improves ability of the model to capture long-range contextual information of text by incorporating direction-awareness, distance-awareness, and unscaled attention. To further improve performance of the model in Chinese text, TEncoder also integrates the inherent features of Chinese characters, such as pinyin and glyphs, which help handle the ambiguity of polysemous words and homophones. The experiments evaluate the JMTF model on a standardized quantitative information dataset of 3106 Chinese text sentences. The results show that our JMTF achieves F1 values of 87.01% and 85.95% for MQI recognition and association, respectively, outperforming the best baseline at 86.42% and 84.65%, demonstrating its advantages in measurable quantitative information extraction. Qixuan Zhang, Fu Lee Wang, Tianyong Hao |
IJCNN | 1 |
| 2025 | Interference Mitigation in 4.9 GHz ISAC Networks: A Multi-domain Resource Management ApproachabstractIntegrated Sensing and Communication (ISAC) is a key enabler for sixth-generation (6G) networks and the low-altitude economy (LAE). However, its deployment faces significant challenges due to limited resources in multi-base station (BS) scenarios. Existing approaches struggle with scalability in multi-BS networks and lack rigorous practical validation. To address these challenges, this paper introduces a novel Time-Frequency-Space Interference Mitigation (TFS-IM) approach that jointly allocates non-overlapping resources, controlling interference lifting below 10 dB in a large-scale 4.9 GHz field trial with 8 mono-static BSs. Compared with the baseline of full-spectrum reuse, the TFS-IM approach achieves a 5 dB interference mitigation gain and improves the position accuracy from 57 m to 15 m. This work provides the first empirical validation and a practical paradigm for large-scale ISAC networks deployment. Junyu Dong, Songtao Gao, Qixuan Zhang, Du Pan |
VTC2025-Fall | 5 |
| 2025 | PFCC-Net: A Polarized Fusion Cross-Correlation Network for efficient and high-quality embedding in few-shot classification
Qixuan Zhang |
Neurocomputing | 1 |
| 2025 | Visual and Textual Prompts in VLLMs for Enhancing Emotion RecognitionabstractVision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. Traditional approaches, which prioritize isolated facial features, often neglect critical non-verbal cues such as body language, environmental context, and social interactions, leading to reduced robustness in real-world scenarios. To address this gap, we propose Set-of-Vision-Text Prompting (SoVTP), a novel framework that enhances zero-shot emotion recognition by integrating spatial annotations (e.g., bounding boxes, facial landmarks), physiological signals (facial action units), and contextual cues (body posture, scene dynamics, others’ emotions) into a unified prompting strategy. SoVTP preserves holistic scene information while enabling fine-grained analysis of facial muscle movements and interpersonal dynamics. Extensive experiments show that SoVTP achieves substantial improvements over existing visual prompting methods, demonstrating its effectiveness in enhancing VLLMs’ video emotion recognition capabilities. Zhifeng Wang 0004, Qixuan Zhang, Peter Zhang, Wenjia Niu, Kaihao Zhang, Ramesh S. Sankaranarayana, Sabrina B. Caldwell, Tom Gedeon |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Facial Appearance Capture at Home with Patch-Level Reflectance PriorabstractExisting facial appearance capture methods can reconstruct plausible facial reflectance from smartphone-recorded videos. However, the reconstruction quality is still far behind the ones based on studio recordings. This paper fills the gap by developing a novel daily-used solution with a co-located smartphone and flashlight video capture setting in a dim room. To enhance the quality, our key observation is to solve facial reflectance maps within the data distribution of studio-scanned ones. Specifically, we first learn a diffusion prior over the Light Stage scans and then steer it to produce the reflectance map that best matches the captured images. We propose to train the diffusion prior at the patch level to improve generalization ability and training stability, as current Light Stage datasets are in ultra-high resolution but limited in data size. Tailored to this prior, we propose a patch-level posterior sampling technique to sample seamless full-resolution reflectance maps from this patch-level diffusion model. Experiments demonstrate our method closes the quality gap between low-cost and studio recordings by a large margin, opening the door for everyday users to clone themselves to the digital world. Junfeng Lyu, Kuan Sheng, Minghao Que, Qixuan Zhang, Lan Xu 0003, Feng Xu 0005 |
ACM Trans. Graph. | 5 |
| 2025 | CAST: Component-Aligned 3D Scene Reconstruction from an RGB ImageabstractRecovering high-quality 3D scenes from a single RGB image is a challenging task in computer graphics. Current methods often struggle with domain-specific limitations or low-quality object generation. To address these, we propose CAST (Component-Aligned 3D Scene Reconstruction from a Single RGB Image), a novel method for 3D scene reconstruction. CAST starts by extracting object-level 2D segmentation and relative depth information from the input image, followed by using a GPT-based model to analyze inter-object spatial relations. This enables understanding of how objects relate to each other within the scene, ensuring more coherent reconstruction. CAST then employs an occlusion-aware large-scale 3D generation model to independently generate each object's full geometry, using Masked Auto Encoder (MAE) and point cloud conditioning to mitigate the effects of occlusions and partial object information, ensuring accurate alignment with the source image's geometry and texture. To align each object with the scene, the alignment generation model computes the necessary transformations, allowing the generated meshes to be accurately placed and integrated into the scene's point cloud. Finally, CAST applies a physics-aware correction mechanism, which leverages a fine-grained relation graph to generate a constraint graph. This graph guides the optimization of object poses, ensuring physical consistency and spatial coherence. By utilizing Signed Distance Fields (SDF), the model effectively addresses issues such as occlusions, object penetration, and floating objects, ensuring that the generated scene accurately reflects real-world physical interactions. Experimental results demonstrate that CAST significantly improves the quality of single-image 3D scene reconstruction, offering enhanced realism and accuracy in scene understanding and reconstruction tasks. CAST has practical applications in virtual content creation, such as immersive game environments and film production, where real-world setups can be seamlessly integrated into virtual landscapes. Additionally, CAST can be leveraged in robotics, enabling efficient real-to-simulation workflows and providing realistic, scalable simulation environments for robotic systems. Kaixin Yao, Longwen Zhang, Xinhao Yan, Qixuan Zhang, Lan Xu 0003, Wei Yang 0034, Jiayuan Gu, Jingyi Yu 0001 |
ACM Trans. Graph. | 5 |
| 2025 | BANG: Dividing 3D Assets via Generative Exploded Dynamicsabstract3D creation has always been a unique human strength, driven by our ability to deconstruct and reassemble objects using our eyes, mind and hand. However, current 3D design tools struggle to replicate this natural process, requiring considerable artistic expertise and manual labor. This paper introduces BANG, a novel generative approach that bridges 3D generation and reasoning, allowing for intuitive and flexible part-level decomposition of 3D objects. At the heart of BANG is "Generative Exploded Dynamics", which creates a smooth sequence of exploded states for an input geometry, progressively separating parts while preserving their geometric and semantic coherence. BANG utilizes a pre-trained large-scale latent diffusion model, fine-tuned for exploded dynamics with a lightweight exploded view adapter, allowing precise control over the decomposition process. It also incorporates a temporal attention module to ensure smooth transitions and consistency across time. BANG enhances control with spatial prompts, such as bounding boxes and surface regions, enabling users to specify which parts to decompose and how. This interaction can be extended with multimodal models like GPT-4, enabling 2D-to-3D manipulations for more intuitive and creative workflows. The capabilities of BANG extend to generating detailed part-level geometry, associating parts with functional descriptions, and facilitating component-aware 3D creation and manufacturing workflows. Additionally, BANG offers applications in 3D printing, where separable parts are generated for easy printing and reassembly. In essence, BANG enables seamless transformation from imaginative concepts to detailed 3D assets, offering a new perspective on creation that resonates with human intuition. Longwen Zhang, Qixuan Zhang, Yinuo Bai 0001, Wei Yang 0034, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 2 |
| 2024 | Visual Prompting in LLMs for Enhancing Emotion RecognitionabstractVision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing.Nonetheless, the potential of using visual prompts for emotion recognition in these models remains largely unexplored and untapped.Traditional methods in VLLMs struggle with spatial localization and often discard valuable global context.To address this problem, we propose a Set-of-Vision prompting (SoV) approach that enhances zero-shot emotion recognition by using spatial information, such as bounding boxes and facial landmarks, to mark targets precisely.SoV improves accuracy in face count and emotion categorization while preserving the enriched image context.Through a battery of experimentation and analysis of recent commercial or open-source VLLMs, we evaluate the SoV model's ability to comprehend facial expressions in natural environments.Our findings demonstrate the effectiveness of integrating spatial visual prompts into VLLMs for improving emotion recognition performance. Question:How many visible faces are there in the image?Analyze each face's emotion and group the face's emotions. Answer:In the image, there are 22 visible faces.Emotions can be grouped into three main categories. Neutral Emotion: Approximately 16Faces.This neutral expression is typical in settings requiring attention to a speaker or event, reflecting an engaged yet reserved emotional state. Mildly Positive Emotion: Approximately 5Faces.This could be a response to something interesting or mildly amusing happening at the event, but their reactions are subdued, in keeping with the formal setting.Happy: 1 Face.This expression stands out as more joyful or enthusiastic about the proceedings. Qixuan Zhang, Zhifeng Wang 0004, Dylan Zhang, Wenjia Niu, Sabrina B. Caldwell, Tom Gedeon, Yang Liu 0003, Zhenyue Qin |
EMNLP | 1 |
| 2024 | Distributed Adaptive Coordinated Control for High-Speed Trains with Input Saturation Based on RBFNN and Sliding Mode ControlabstractThis paper addresses the distributed adaptive coordinated control for high-speed train (HST) fleet with uncertain parameters. The motion of the train in the fleet is constrained by its adjacent trains, necessitating dynamic adjustment mechanism facilitated through inter-train communication. For the uncertainty, radial basic function neural network (RBFNN) is introduced into the distributed adaptive coordinated control algorithm, which ensures behavioral consistency and short inter-train intervals for each train in the fleet. This paper compares the proposed method with distributed adaptive sliding mode control (DASMC). The simulation demonstrates better performance and benefits of this new algorithm. We show that the algorithm substantially reduces inter-train distance and ensures heightened level of behavioral consistency among all individual trains within the train fleet. Pengfei Sun 0002, Qixuan Zhang, Youxing Guo, Qingyuan Wang 0001, Xiaoyun Feng |
IV | 2 |
| 2024 | DressCode: Autoregressively Sewing and Generating Garments from Text GuidanceabstractApparel's significant role in human appearance underscores the importance of garment digitalization for digital human creation. Recent advances in 3D content creation are pivotal for digital human creation. Nonetheless, garment generation from text guidance is still nascent. We introduce a text-driven 3D garment generation framework, DressCode, which aims to democratize design for novices and offer immense potential in fashion design, virtual try-on, and digital human creation. We first introduce SewingGPT, a GPT-based architecture integrating cross-attention with text-conditioned embedding to generate sewing patterns with text guidance. We then tailor a pre-trained Stable Diffusion to generate tile-based Physically-based Rendering (PBR) textures for the garments. By leveraging a large language model, our framework generates CG-friendly garments through natural language interaction. It also facilitates pattern completion and texture editing, streamlining the design process through user-friendly interaction. This framework fosters innovation by allowing creators to freely experiment with designs and incorporate unique elements into their work. With comprehensive evaluations and comparisons with other state-of-the-art methods, our method showcases superior quality and alignment with input prompts. User studies further validate our high-quality rendering results, highlighting its practical utility and potential in production settings. Our project page is https://IHe-KaiI.github.io/DressCode/. Kaixin Yao, Qixuan Zhang, Jingyi Yu 0001, Lingjie Liu, Lan Xu 0003 |
ACM Trans. Graph. | 3 |
| 2024 | Implicit Swept Volume SDF: Enabling Continuous Collision-Free Trajectory Generation for Arbitrary ShapesabstractIn the field of trajectory generation for objects, ensuring continuous collision-free motion remains a huge challenge, especially for non-convex geometries and complex environments. Previous methods either oversimplify object shapes, which results in a sacrifice of feasible space or rely on discrete sampling, which suffers from the "tunnel effect". To address these limitations, we propose a novel hierarchical trajectory generation pipeline, which utilizes the Swept Volume Signed Distance Field (SVSDF) to guide trajectory optimization for Continuous Collision Avoidance (CCA). Our interdisciplinary approach, blending techniques from graphics and robotics, exhibits outstanding effectiveness in solving this problem. We formulate the computation of the SVSDF as a Generalized Semi-Infinite Programming model, and we solve for the numerical solutions at query points implicitly, thereby eliminating the need for explicit reconstruction of the surface. Our algorithm has been validated in a variety of complex scenarios and applies to robots of various dynamics, including both rigid and deformable shapes. It demonstrates exceptional universality and superior CCA performance compared to typical algorithms. The code will be released at https://github.com/ZJU-FAST-Lab/Implicit-SVSDF-Planner for the benefit of the community. Qixuan Zhang, Chuxiao Zeng, Jingyi Yu 0001, Chao Xu 0001, Lan Xu 0003, Fei Gao 0011 |
ACM Trans. Graph. | 3 |
| 2024 | CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D AssetsabstractIn the realm of digital creativity, our potential to craft intricate 3D worlds from imagination is often hampered by the limitations of existing digital tools, which demand extensive expertise and efforts. To narrow this disparity, we introduce CLAY, a 3D geometry and material generator designed to effortlessly transform human imagination into intricate 3D digital structures. CLAY supports classic text or image inputs as well as 3D-aware controls from diverse primitives (multi-view images, voxels, bounding boxes, point clouds, implicit representations, etc). At its core is a large-scale generative model composed of a multi-resolution Variational Autoencoder (VAE) and a minimalistic latent Diffusion Transformer (DiT), to extract rich 3D priors directly from a diverse range of 3D geometries. Specifically, it adopts neural fields to represent continuous and complete surfaces and uses a geometry generative module with pure transformer blocks in latent space. We present a progressive training scheme to train CLAY on an ultra large 3D model dataset obtained through a carefully designed processing pipeline, resulting in a 3D native geometry generator with 1.5 billion parameters. For appearance generation, CLAY sets out to produce physically-based rendering (PBR) textures by employing a multi-view material diffusion model that can generate 2K resolution textures with diffuse, roughness, and metallic modalities. We demonstrate using CLAY for a range of controllable 3D asset creations, from sketchy conceptual designs to production ready assets with intricate details. Even first time users can easily use CLAY to bring their vivid 3D imaginations to life, unleashing unlimited creativity. Longwen Zhang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Wei Yang 0034, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 3 |
| 2023 | Relightable Neural Human Assets from Multi-view Gradient IlluminationsabstractHuman modeling and relighting are two fundamental problems in computer vision and graphics, where high-quality datasets can largely facilitate related research. However, most existing human datasets only provide multi-view human images captured under the same illumination. Although valuable for modeling tasks, they are not read-ily used in relighting problems. To promote research in both fields, in this paper, we present UltraStage, a new 3D human dataset that contains more than 2, 000 high-quality human assets captured under both multi-view and multi-illumination settings. Specifically, for each example, we provide 32 surrounding views illuminated with one white light and two gradient illuminations. In addition to regular multi-view images, gradient illuminations help recover de-tailed surface normal and spatially-varying material maps, enabling various relighting applications. Inspired by recent advances in neural representation, we further interpret each example into a neural human asset which allows novel view synthesis under arbitrary lighting conditions. We show our neural human assets can achieve extremely high capture performance and are capable of representing fine details such as facial wrinkles and cloth folds. We also validate UltraStage in single image relighting tasks, training neural networks with virtual relighted data from neural assets and demonstrating realistic rendering improvements over prior arts. UltraStage will be publicly available to the community to stimulate significant future developments in various human modeling and rendering tasks. The dataset is available at https://miaoing.github.io/RNHA. Taotao Zhou 0006, Teng Xu 0008, Qixuan Zhang, Kuixiang Shao, Wenzheng Chen, Lan Xu 0003, Jingyi Yu 0001 |
CVPR | 5 |
| 2023 | DreamFace: Progressive Generation of Animatable 3D Faces under Text GuidanceabstractEmerging Metaverse applications demand accessible, accurate and easy-to-use tools for 3D digital human creations in order to depict different cultures and societies as if in the physical world. Recent large-scale vision-language advances pave the way for novices to conveniently customize 3D content. However, the generated CG-friendly assets still cannot represent the desired facial traits for human characteristics. In this paper, we present Dream-Face, a progressive scheme to generate personalized 3D faces under text guidance. It enables layman users to naturally customize 3D facial assets that are compatible with CG pipelines, with desired shapes, textures and fine-grained animation capabilities. From a text input to describe the facial traits, we first introduce a coarse-to-fine scheme to generate the neutral facial geometry with a unified topology. We employ a selection strategy in the CLIP embedding space to generate coarse geometry, and subsequently optimize both the detailed displacements and normals using Score Distillation Sampling (SDS) from the generic Latent Diffusion Model (LDM). Then, for neutral appearance generation, we introduce a dual-path mechanism, which combines the generic LDM with a novel texture LDM to ensure both the diversity and textural specification in the UV space. We also employ a two-stage optimization to perform SDS in both the latent and image spaces to significantly provide compact priors for fine-grained synthesis. It also enables learning the mapping from the compact latent space into physically-based textures (diffuse albedo, specular intensity, normal maps, etc.). Our generated neutral assets naturally support blendshapes-based facial animations, thanks to the unified geometric topology. We further improve the animation ability with personalized deformation characteristics. To this end, we learn the universal expression prior in a latent space with neutral asset conditioning using the cross-identity hypernetwork, we subsequently train a neural facial tracker from video input space into the pre-trained expression space for personalized fine-grained animation. Extensive qualitative and quantitative experiments validate the effectiveness and generalizability of DreamFace. Notably, DreamFace can generate realistic 3D facial assets with physically-based rendering quality and rich animation ability from video footage, even for fashion icons or exotic characters in cartoons and fiction movies. Longwen Zhang, Qiwei Qiu, Hongyang Lin, Qixuan Zhang, Cheng Shi 0001, Wei Yang 0034, Ye Shi 0001, Sibei Yang, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 4 |
| 2023 | HACK: Learning a Parametric Head and Neck Model for High-fidelity AnimationabstractSignificant advancements have been made in developing parametric models for digital humans, with various approaches concentrating on parts such as the human body, hand, or face. Nevertheless, connectors such as the neck have been overlooked in these models, with rich anatomical priors often unutilized. In this paper, we introduce HACK (Head-And-neCK), a novel parametric model for constructing the head and cervical region of digital humans. Our model seeks to disentangle the full spectrum of neck and larynx motions, facial expressions, and appearance variations, providing personalized and anatomically consistent controls, particularly for the neck regions. To build our HACK model, we acquire a comprehensive multi-modal dataset of the head and neck under various facial expressions. We employ a 3D ultrasound imaging scheme to extract the inner biomechanical structures, namely the precise 3D rotation information of the seven vertebrae of the cervical spine. We then adopt a multi-view photometric approach to capture the geometry and physically-based textures of diverse subjects, who exhibit a diverse range of static expressions as well as sequential head-and-neck movements. Using the multi-modal dataset, we train the parametric HACK model by separating the 3D head and neck depiction into various shape, pose, expression, and larynx blendshapes from the neutral expression and the rest skeletal pose. We adopt an anatomically-consistent skeletal design for the cervical region, and the expression is linked to facial action units for artist-friendly controls. We also propose to optimize the mapping from the identical shape space to the PCA spaces of personalized blendshapes to augment the pose and expression blendshapes, providing personalized properties within the framework of the generic model. Furthermore, we use larynx blendshapes to accurately control the larynx deformation and force the larynx slicing motions along the vertical direction in the UV-space for precise modeling of the larynx beneath the neck skin. HACK addresses the head and neck as a unified entity, offering more accurate and expressive controls, with a new level of realism, particularly for the neck regions. This approach has significant benefits for numerous applications, including geometric fitting and animation, and enables inter-correlation analysis between head and neck for fine-grained motion synthesis and transfer. Longwen Zhang, Zijun Zhao, Xinzhou Cong, Qixuan Zhang, Shuqi Gu, Yuchong Gao, Wei Yang 0034, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 4 |
| 2022 | ARL: An adaptive reinforcement learning framework for complex question answering over knowledge base
Qixuan Zhang, Xinyi Weng, Guangyou Zhou, Yi Zhang 0118, Jimmy Huang 0001 |
Inf. Process. Manag. | 1 |
| 2022 | SCULPTOR: Skeleton-Consistent Face Creation Using a Learned Parametric GeneratorabstractRecent years have seen growing interest in 3D human face modeling due to its wide applications in digital human, character generation and animation. Existing approaches overwhelmingly emphasized on modeling the exterior shapes, textures and skin properties of faces, ignoring the inherent correlation between inner skeletal structures and appearance. In this paper, we present SCULPTOR, 3D face creations with Skeleton Consistency Using a Learned Parametric facial generaTOR , aiming to facilitate the easy creation of both anatomically correct and visually convincing face models via a hybrid parametric-physical representation. At the core of SCULPTOR is LUCY, the first large-scale shape-skeleton face dataset in collaboration with plastic surgeons. Named after the fossils of one of the oldest known human ancestors, our LUCY dataset contains high-quality Computed Tomography (CT) scans of the complete human head before and after orthognathic surgeries, which are critical for evaluating surgery results. LUCY consists of 144 scans of 72 subjects (31 male and 41 female), where each subject has two CT scans taken pre- and post-orthognathic operations. Based on our LUCY dataset, we learned a novel skeleton consistent parametric facial generator, SCULPTOR, which can create unique and nuanced facial features that help define a character and at the same time maintain physiological soundness. Our SCULPTOR jointly models the skull, face geometry and face appearance under a unified data-driven framework by separating the depiction of a 3D face into shape blend shape, pose blend shape and facial expression blend shape. SCULPTOR preserves both anatomic correctness and visual realism in facial generation tasks compared with existing methods. Finally, we showcase the robustness and effectiveness of SCULPTOR in various fancy applications unseen before, like archaeological skeletal facial completion, bone-aware character fusion, skull inference from images, face generation with lipo-Level change and facial animations, etc. Zesong Qiu, Dongming He, Qixuan Zhang, Longwen Zhang, Jingya Wang 0001, Lan Xu 0003, Yuyao Zhang 0005, Jingyi Yu 0001 |
ACM Trans. Graph. | 4 |
| 2022 | Video-Driven Neural Physically-Based Facial Asset for ProductionabstractProduction-level workflows for producing convincing 3D dynamic human faces have long relied on an assortment of labor-intensive tools for geometry and texture generation, motion capture and rigging, and expression synthesis. Recent neural approaches automate individual components but the corresponding latent representations cannot provide artists with explicit controls as in conventional tools. In this paper, we present a new learning-based, video-driven approach for generating dynamic facial geometries with high-quality physically-based assets. For data collection, we construct a hybrid multiview-photometric capture stage, coupling with ultra-fast video cameras to obtain raw 3D facial assets. We then set out to model the facial expression, geometry and physically-based textures using separate VAEs where we impose a global MLP based expression mapping across the latent spaces of respective networks, to preserve characteristics across respective attributes. We also model the delta information as wrinkle maps for the physically-based textures, achieving high-quality 4K dynamic textures. We demonstrate our approach in high-fidelity performer-specific facial capture and cross-identity facial motion retargeting. In addition, our multi-VAE-based neural asset, along with the fast adaptation schemes, can also be deployed to handle in-the-wild videos. Besides, we motivate the utility of our explicit facial disentangling strategy by providing various promising physically-based editing results with high realism. Comprehensive experiments show that our technique provides higher accuracy and visual fidelity than previous video-driven facial reconstruction and animation methods. Longwen Zhang, Chuxiao Zeng, Qixuan Zhang, Hongyang Lin, Ruixiang Cao, Wei Yang 0034, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 3 |
| 2021 | Convolutional Neural Opacity Radiance FieldsabstractPhoto-realistic modeling and rendering of fuzzy objects with complex opacity are critical for numerous immersive VR/AR applications, but it suffers from strong view-dependent brightness, color. In this paper, we propose a novel scheme to generate opacity radiance fields with a convolutional neural renderer for fuzzy objects, which is the first to combine both explicit opacity supervision and convolutional mechanism into the neural radiance field framework so as to enable high-quality appearance and global consistent alpha mattes generation in arbitrary novel views. More specifically, we propose an efficient sampling strategy along with both the camera rays and image plane, which enables efficient radiance field sampling and learning in a patch-wise manner, as well as a novel volumetric feature integration scheme that generates per-patch hybrid feature embeddings to reconstruct the view-consistent fine-detailed appearance and opacity output. We further adopt a patch-wise adversarial training scheme to preserve both high-frequency appearance and opacity details in a self-supervised framework. We also introduce an effective multi-view image capture system to capture high-quality color and alpha maps for challenging fuzzy objects. Extensive experiments on existing and our new challenging fuzzy object dataset demonstrate that our method achieves photo-realistic, globally consistent, and fined detailed appearance and opacity free-viewpoint rendering for various fuzzy objects. Haimin Luo, Anpei Chen, Qixuan Zhang, Bai Pang, Minye Wu, Lan Xu 0003, Jingyi Yu 0001 |
ICCP | 3 |
| 2021 | Neural Video Portrait Relighting in Real-time via Consistency ModelingabstractVideo portraits relighting is critical in user-facing human photography, especially for immersive VR/AR experience. Recent advances still fail to recover consistent relit result under dynamic illuminations from monocular RGB stream, suffering from the lack of video consistency supervision. In this paper, we propose a neural approach for real-time, high-quality and coherent video portrait relighting, which jointly models the semantic, temporal and lighting consistency using a new dynamic OLAT dataset. We propose a hybrid structure and lighting disentanglement in an encoder-decoder architecture, which combines a multi-task and adversarial training strategy for semantic-aware consistency modeling. We adopt a temporal modeling scheme via flow-based supervision to encode the conjugated temporal consistency in a cross manner. We also propose a lighting sampling strategy to model the illumination consistency and mutation for natural portrait light manipulation in real-world. Extensive experiments demonstrate the effectiveness of our approach for consistent video portrait light-editing and relighting, even using mobile computing. Longwen Zhang, Qixuan Zhang, Minye Wu, Jingyi Yu 0001, Lan Xu 0003 |
ICCV | 2 |
| 2021 | Towards Controllable and Photorealistic Region-wise Image ManipulationabstractAdaptive and flexible image editing is a desirable function of modern generative models. In this work, we present a generative model with auto-encoder architecture for per-region style manipulation. We apply a code consistency loss to enforce an explicit disentanglement between content and style latent representations, making the content and style of generated samples consistent with their corresponding content and style references. The model is also constrained by a content alignment loss to ensure the foreground editing will not interfere background contents. As a result, given interested region masks provided by users, our model supports foreground region-wise style transfer. Specially, our model receives no extra annotations such as semantic labels except for self-supervision. Extensive experiments show the effectiveness of the proposed method and exhibit the flexibility of the proposed model for various applications, including region-wise style editing, latent space interpolation, cross-domain style transfer. Ansheng You, Chenglin Zhou, Qixuan Zhang, Lan Xu 0003 |
ACM Multimedia | 3 |
| 2019 | Variational Bayesian Channel Estimation for Wideband Multiuser mmWave SystemsabstractIn this paper, a frequency-distributed variational Bayesian (F-DVB) channel estimation algorithm is proposed for wideband multiuser millimeter wave (mmWave) multiple-input multiple-output (MIMO) systems, where hybrid precoding architectures are adopted and frequency selective fading channels are assumed. First, a distributed compressed sensing-based method is employed by leveraging the joint sparsity of different subcarriers in the frequency domain, reducing the required pilot overhead significantly. Next, a hierarchical channel model, which adopts an identify-and-reject strategy to deal with hardware impairments, is designed to enhance the robustness of the proposed algorithm. Finally, the channel information is estimated by a modified variational Bayesian method, which improves the channel estimation accuracy dramatically. Simulation results verify that the proposed algorithm outperforms the state-of-the-art channel estimation strategies at low SNR and pilot overhead. Qixuan Zhang, Tiejun Lv, Zhipeng Lin 0001 |
ICC | 1 |
| 2018 | Fast Sparse Bayesian Channel Estimation for Wideband mmWave SystemsabstractIn this paper, we propose a fast channel estimation algorithm for wideband millimeter wave (mmWave) massive MIMO systems, where the hybrid precoding architectures are adopted. Stimulated by the joint sparsity of different subcarriers, a distributed compressed sensing-based strategy is presented to reduce the required pilot overhead. Based on a hierarchical channel model, a fast sparse Bayesian learning method, which can drastically release the relaxed evidence low bound, is designed to accelerate the convergence rate. Simulation results verify that the proposed algorithm is capable of achieving substantially higher estimation accuracy and convergence rate as compared to other existing Bayesian channel estimation strategies. Qixuan Zhang, Zhipeng Lin 0001, Tiejun Lv |
PIMRC | 1 |