Jingxi Xu 0001

dblp:33/10762 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0002-6262-4199ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-authorComputer networks · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 41% Deep learning architectures and training · 32% Generative modeling · 27%
Computer networks
2 papers
Vehicular, aerial and satellite networks · 58% Transport protocols and congestion control · 36% Network optimization and economics · 6%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 61% Geometric modeling and processing · 30% Image and video coding · 9%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Transport protocols and congestion control
real-time communication
1.422025
SpaceRTC: Unleashing the Low-Latency Potential of Mega-Constellations for Wide-Area Real-Time Communications · IEEE Trans. Mob. Comput. 2025
SpaceRTC: Unleashing the Low-latency Potential of Mega-constellations for Real-Time Communications · INFOCOM 2022
Vehicular, aerial and satellite networks
satellite networks
1.422025
SpaceRTC: Unleashing the Low-Latency Potential of Mega-Constellations for Wide-Area Real-Time Communications · IEEE Trans. Mob. Comput. 2025
SpaceRTC: Unleashing the Low-latency Potential of Mega-constellations for Real-Time Communications · INFOCOM 2022
Machine learning › Deep learning architectures and training › transformer › efficient transformer
sparse transformer
0.912025
High-quality Text-to-3D Character Generation with SparseCubes and Sparse Transformers · ICLR 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
High-quality Text-to-3D Character Generation with SparseCubes and Sparse Transformers · ICLR 2025
Visual content generation and editing
3d content generation
0.912025
High-quality Text-to-3D Character Generation with SparseCubes and Sparse Transformers · ICLR 2025
Geometric modeling and processing › shape representation
mesh representation
0.912025
High-quality Text-to-3D Character Generation with SparseCubes and Sparse Transformers · ICLR 2025
Visual content generation and editing › 3d content generation
text-to-3d generation
0.912025
High-quality Text-to-3D Character Generation with SparseCubes and Sparse Transformers · ICLR 2025
Vehicular, aerial and satellite networks › satellite networks
LEO mega-constellation
0.912025
SpaceRTC: Unleashing the Low-Latency Potential of Mega-Constellations for Wide-Area Real-Time Communications · IEEE Trans. Mob. Comput. 2025
Computer vision › 3D vision
3d generation
0.812024
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer · NeurIPS 2024
Computer vision › 3D vision
3d shape representation
0.812024
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer · NeurIPS 2024
Computer vision › 3D vision › 3d generation
image-to-3d generation
0.812024
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer · NeurIPS 2024
Machine learning › Generative modeling › diffusion model › diffusion transformer
latent diffusion transformer
0.812024
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer · NeurIPS 2024
Image and video coding › image quality assessment
just noticeable difference
0.212016
Optimality of Greedy Algorithm for Generating Just-Noticeable Difference Surfaces · IEEE Trans. Multim. 2016
Mathematical optimization › combinatorial optimization
greedy algorithm
0.212016
Optimality of Greedy Algorithm for Generating Just-Noticeable Difference Surfaces · IEEE Trans. Multim. 2016

Methods — techniques the papers use, named apart from their topics

sparse transformer · 1.7overlay routing · 1.7measurement study · 1.7differentiable mesh representation · 1.7bitrate adaptation · 0.9bit rate adaptation · 0.9variational autoencoder · 0.8semi-continuous surface sampling · 0.8diffusion transformer · 0.8subjective testing · 0.8greedy algorithm · 0.8satellite-cloud cooperative relay selection · 0.6overlay network construction · 0.6flow allocation · 0.6
YearPublicationVenuePosition
2025 High-quality Text-to-3D Character Generation with SparseCubes and Sparse Transformers
abstract
Current state-of-the-art text-to-3D generation methods struggle to produce 3D models with fine details and delicate structures due to limitations in differentiable mesh representation techniques. This limitation is particularly pronounced in anime character generation, where intricate features such as fingers, hair, and facial details are crucial for capturing the essence of the characters. In this paper, we introduce a novel, efficient, sparse differentiable mesh representation method, termed SparseCubes, alongside a sparse transformer network designed to generate high-quality 3D models. Our method significantly reduces computational requirements by over 95% and storage memory by 50%, enabling the creation of higher resolution meshes with enhanced details and delicate structures. We validate the effectiveness of our approach through its application to text-to-3D anime character generation, demonstrating its capability to accurately render subtle details and thin structures (e.g. individual fingers) in both meshes and textures.
Jiachen Qian, Hongye Yang, Jingxi Xu 0001, Feihu Zhang
ICLR4
2025 SpaceRTC: Unleashing the Low-Latency Potential of Mega-Constellations for Wide-Area Real-Time Communications
abstract
User-perceived latency is important for the quality of experience (QoE) of wide-area real-time communications (RTC). With the rapid development of low Earth orbit (LEO) mega-constellations, this paper explores a futuristic yet important problem facing the RTC community:can we exploit emerging mega-constellations to facilitate low-latency RTC globally?We carry out our quest in three steps. First, through a measurement study associated with a large number of geo-distributed RTC users, we quantitatively expose that themeandering routesin theclient-to-cloudandinter-cloud-sitesegment of existing cloud-based RTC architecture are critical culprits for the high latency issue suffered by wide-area RTC sessions. Second, we proposeSpaceRTC, a satellite-cloud cooperative framework that dynamically selectsrelay serversupon satellites and cloud sites to build an overlay network which enables diverse close-to-optimal paths.SpaceRTCjudiciously allocates RTC flows of different sessions upon the network to facilitate low-latency interactions and adaptively selects bitrates to offer high user-perceived QoE in energy-limited space circumstance. Finally, we implement a testbed based on public constellation information and real-world RTC traces. Extensive experiments demonstrate thatSpaceRTCcan deliver near-optimal interactive latency, with up to 53.3% average latency reduction and 103.6% average bitrate improvement as compared to other state-of-the-art cloud-based solutions.
Zeqi Lai, Weisen Liu, Qian Wu 0001, Hewu Li, Jingxi Xu 0001, Yuanjie Li, Jun Liu 0063
IEEE Trans. Mob. Comput.5
2024 Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer
abstract
Generating high-quality 3D assets from text and images has long been challenging, primarily due to the absence of scalable 3D representations capable of capturing intricate geometry distributions. In this work, we introduce Direct3D, a native 3D generative model scalable to in-the-wild input images, without requiring a multi-view diffusion model or SDS optimization. Our approach comprises two primary components: a Direct 3D Variational Auto-Encoder (D3D-VAE) and a Direct 3D Diffusion Transformer (D3D-DiT). D3D-VAE efficiently encodes high-resolution 3D shapes into a compact and continuous latent triplane space. Notably, our method directly supervises the decoded geometry using a semi-continuous surface sampling strategy, diverging from previous methods relying on rendered images as supervision signals. D3D-DiT models the distribution of encoded 3D latents and is specifically designed to fuse positional information from the three feature maps of the triplane latent, enabling a native 3D generative model scalable to large-scale 3D datasets. Additionally, we introduce an innovative image-to-3D generation pipeline incorporating semantic and pixel-level image conditions, allowing the model to produce 3D shapes consistent with the provided conditional image input. Extensive experiments demonstrate the superiority of our large-scale pre-trained Direct3D over previous image-to-3D approaches, achieving significantly better generation quality and generalization ability, thus establishing a new state-of-the-art for 3D content creation. Project page: https://www.neural4d.com/research/direct3d.
Youtian Lin, Yifei Zeng, Feihu Zhang, Jingxi Xu 0001, Philip Torr 0001, Xun Cao, Yao Yao 0008
NeurIPS5
2022 SpaceRTC: Unleashing the Low-latency Potential of Mega-constellations for Real-Time Communications
abstract
User-perceived latency is important for the quality of experience (QoE) of wide-area real-time communications (RTC). This paper explores a futuristic yet important problem facing the RTC community: can we exploit emerging mega-constellations to facilitate low-latency RTC globally? We carry out our quest in three steps. First, through a measurement study associated with a large number of geo-distributed RTC users, we quantitatively expose that the meandering routes in the client-cloud and inter-cloud-site segment of existing cloud-based RTC architecture are critical culprits for the high latency issue suffered by wide-area RTC sessions. Second, we propose SPACERTC, a satellite-cloud cooperative framework that adaptively selects relay servers upon satellites and cloud sites to build an overlay network which enables diverse close-to-optimal paths, and then judiciously allocates RTC flows upon the network to facilitate low-latency interactions. Finally, we implement our SPACERTC prototype on an experimental environment based on public constellation information and RTC trace, and extensive experiments demonstrate that SPACERTC can deliver near-optimal interactive latency, with up to 64.9% latency reduction as compared to other state-of-the-art cloud-based solutions under representative videoconferencing traffic.
Zeqi Lai, Weisen Liu, Qian Wu 0001, Hewu Li, Jingxi Xu 0001
INFOCOM5
2016 Optimality of Greedy Algorithm for Generating Just-Noticeable Difference Surfaces
abstract
Tuning multimedia applications at run time to achieve high perceptual quality entails the search of nonlinear mappings that determine how control inputs should be set in order to lead to high user-perceived quality. Offline subjective tests are often used for this purpose but they are expensive to conduct because each can only evaluate one mapping at a time and there can be infinitely many such mappings to be evaluated. In this paper, we present a greedy algorithm that uses a small number of subjective test results to accurately approximate this space of mappings. Based on an axiom on monotonicity and the property of just-noticeable differences, we prove its optimality in minimizing the average absolute error between the approximate and the original mappings. We further demonstrate the results using numerical simulations and the application of the mappings found to tune the control of the multimedia game BZFlag.
Jingxi Xu 0001, Benjamin W. Wah
IEEE Trans. Multim.1
2016 Consistent Synchronization of Action Order with Least Noticeable Delays in Fast-Paced Multiplayer Online Games
abstract
When running multiplayer online games on IP networks with losses and delays, the order of actions may be changed when compared to the order run on an ideal network with no delays and losses. To maintain a proper ordering of events, traditional approaches either use rollbacks to undo certain actions or local lags to introduce additional delays. Both may be perceived by players because their changes are beyond the just-noticeable-difference (JND) threshold. In this article, we propose a novel method for ensuring a strongly consistent completion order of actions, where strong consistency refers to the same completion order as well as the same interval between any completion time and the corresponding ideal reference completion time under no network delay. We find that small adjustments within the JND on the duration of an action would not be perceivable, as long as the duration is comparable to the network round-trip time. We utilize this property to control the vector of durations of actions and formulate the search of the vector as a multidimensional optimization problem. By using the property that players are generally more sensitive to the most prominent delay effect (with the highest probability of noticeability P notice or the probability of correctly noticing a change when compared to the reference), we prove that the optimal solution occurs when P notice of the individual adjustments are equal. As this search can be done efficiently in polynomial time ( ∼ 5ms) with a small amount of space ( ∼ 160KB), the search can be done at runtime to determine the optimal control. Last, we evaluate our approach on the popular open-source online shooting game BZFlag.
Jingxi Xu 0001, Benjamin W. Wah
ACM Trans. Multim. Comput. Commun. Appl.1
2013 Concealing network delays in delay-sensitive online interactive games based on just-noticeable differences
abstract
Online delay-sensitive games with fast interactive actions, like fighting games (FTG) and sports games, require the synchronization of coupled multi-player actions. With network impairments, action commands from other players can be delayed or lost, leading to compromised real-time perception of these games. In contrast to slower strategy games, online fighting and sports games studied in this paper require realtime judgment and instant feedbacks. Traditional methods for optimizing the delay effects of these games are focused on quantifying round-trip delays, without examining the perceptual effects of players. In this paper, we develop a new criterion using just-noticeable differences (JND) for optimizing the duration of actions and responses. Our approach aims to reduce the probability of players perceiving the delay effects, when compared to a reference game with zero network delay. Using statistics collected in offline subjective tests, the timing of actions is modified at run time. Experimental results show significant reduction of players' awareness of network delays using our approach when compared to existing delay-concealment schemes.
Jingxi Xu 0001, Benjamin W. Wah
ICME1
2013 Exploiting just-noticeable difference of delays for improving quality of experience in video conferencing
abstract
This paper proposes a novel approach for improving the quality of experience (QoE) of real-time video conferencing systems. In these systems, QoE is affected by signal quality as well as interactivity, both depending on the packet loss rate, delay jitters, and mouth-to-ear delay (MED) that measures the sender-receiver delay on audio signals (and will be the same as that of video signals when video and audio is synchronized). We notice in the current Internet that increasing MED as well as reducing packet rate can help reduce the delay-aware loss rate in congested connections. Between the two methods, the former plays a more important role and applies well to a variety of network conditions for improving audiovisual signal quality, although overly increasing the MED will degrade interactivity. Based on a psychophysical concept called just-noticeable difference (JND), we find the extent to which MED can be increased, without humans perceiving the difference from the original conversation. The approach can be applied to improve existing video conferencing systems. Starting from the operating point of an existing system, we increase its MED to within JND in order to have more room for smoothing network delay spikes as well as recovering lost packets, without incurring noticeable degradation in interactivity. We demonstrate the idea on Skype and Windows Live Messenger by designing a traffic interceptor to extend their buffering time and to perform packet scheduling/recovery. Our experimental results show significant improvements in QoE, with much better signal quality while maintaining similar interactivity.
Jingxi Xu 0001, Benjamin W. Wah
MMSys1
2011 Delay-Aware Loss-Concealment Strategies for Real-Time Video Conferencing
abstract
One-way audiovisual quality and mouth-to-ear delay (MED) are two important quality metrics in the design of real-time video-conferencing systems, and their trade-offs have significant impact on the user-perceived quality. In this paper, we address one aspect of this larger problem by developing efficient loss-concealment schemes that optimize the one-way quality under given MED and network conditions. Our experimental results show that our approach can attain significant improvements over the LARDo reference scheme that does not consider MED in its optimization.
Jingxi Xu 0001, Benjamin W. Wah
ISM1