EDBT 2026 Demo / reviewers in the wild / expert
Xianghao Kong
dblp:212/0360
· DBLP profile ↗
13ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 44% Vision and language · 19% Autonomous driving · 10% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% | |
| Theoretical computer science
2 papers |
Information theory · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.3 | 3 | 2025 | Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget · CVPR 2025 Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training · NeurIPS 2024 Information-Theoretic Diffusion · ICLR 2023 |
Visual content generation and editing › video generation
diffusion-based video generation |
1.0 | 1 | 2026 | UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors · ACM Trans. Graph. 2026 |
Visual content generation and editing
video generation |
1.0 | 1 | 2026 | UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors · ACM Trans. Graph. 2026 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget · CVPR 2025 |
Computer vision › Vision and language › compositionality
compositional understanding |
0.8 | 1 | 2024 | Interpretable Diffusion via Information Decomposition · ICLR 2024 |
Natural language and speech › Language models and text generation
controllable text generation |
0.8 | 1 | 2024 | Controllable Navigation Instruction Generation with Chain of Thought Prompting · ECCV (29) 2024 |
Machine learning › Generative modeling › diffusion model
denoising |
0.8 | 1 | 2024 | Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
diffusion model interpretability |
0.8 | 1 | 2024 | Interpretable Diffusion via Information Decomposition · ICLR 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Interpretable Diffusion via Information Decomposition · ICLR 2024 |
Computer vision › Vision and language › vision-and-language navigation
navigation instruction generation |
0.8 | 1 | 2024 | Controllable Navigation Instruction Generation with Chain of Thought Prompting · ECCV (29) 2024 |
Machine learning › Generative modeling › diffusion model
parallel sampling |
0.8 | 1 | 2024 | Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training · NeurIPS 2024 |
Computer vision › Vision and language
vision-and-language navigation |
0.8 | 1 | 2024 | Controllable Navigation Instruction Generation with Chain of Thought Prompting · ECCV (29) 2024 |
Information theory › information measures
information decomposition |
0.8 | 1 | 2024 | Interpretable Diffusion via Information Decomposition · ICLR 2024 |
Robotics › Autonomous driving
collaborative perception |
0.7 | 1 | 2023 | DUSA: Decoupled Unsupervised Sim2Real Adaptation for Vehicle-to-Everything Collaborative Perception · ACM Multimedia 2023 |
Robotics › Autonomous driving › collaborative perception
vehicle-to-everything perception |
0.7 | 1 | 2023 | DUSA: Decoupled Unsupervised Sim2Real Adaptation for Vehicle-to-Everything Collaborative Perception · ACM Multimedia 2023 |
Information theory › information measures
mutual information |
0.7 | 1 | 2023 | Information-Theoretic Diffusion · ICLR 2023 |
Computer vision › 3D vision › 3d scene understanding
3d visual grounding |
0.6 | 1 | 2022 | 3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection · CVPR 2022 |
Machine learning › Generative modeling › diffusion model
diffusion transformer |
0.3 | 1 | 2025 | Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget · CVPR 2025 |
Computer vision › 3D vision
3d object detection |
0.2 | 1 | 2023 | DUSA: Decoupled Unsupervised Sim2Real Adaptation for Vehicle-to-Everything Collaborative Perception · ACM Multimedia 2023 |
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
synthetic-to-real domain adaptation |
0.2 | 1 | 2023 | DUSA: Decoupled Unsupervised Sim2Real Adaptation for Vehicle-to-Everything Collaborative Perception · ACM Multimedia 2023 |
Computer vision › Vision and language
cross-modal matching |
0.2 | 1 | 2022 | 3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection · CVPR 2022 |
Methods — techniques the papers use, named apart from their topics
pointwise information estimation · 1.5mutual information · 1.5information theory · 1.3stochastic condition masking · 1.0diffusion prior · 1.0decoupled gated LoRA · 1.0cross-modal self-attention · 1.0synthetic data training · 0.9patch masking · 0.9mixture of experts · 0.9self-supervised learning · 0.8log-likelihood ratio · 0.8contrastive training · 0.8chain-of-thought prompting · 0.8variational bounds · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion PriorsabstractRecent progress has shown that video diffusion models (VDMs) can be repurposed to solve various multimodal graphics tasks. However, existing approaches predominantly train separate models for each specific problem setting. This practice locks models into fixed input-output mappings, and typically ignores the joint correlations across modalities. In this paper, we present UniVidX , a unified multimodal framework designed to leverage VDM priors to enable versatile video generation. Our goal is to (i) master diverse pixel-aligned tasks by formulating them as conditional generation problems within multimodal space, (ii) adapt to modality-specific distributions without compromising the backbone's native priors, and (iii) ensure cross-modal consistency during synthesis. Concretely, we propose three key designs: 1) Stochastic Condition Masking (SCM): by randomly partitioning modalities into clean conditions and noisy targets during training, we enable the model to learn omni-directional conditional generation rather than fixed mappings. 2) Decoupled Gated LoRA (DGL): we attach per-modality LoRAs and activate them when a modality serves as a generation target, thereby preserving the VDM's strong priors. 3) Cross-Modal Self-Attention (CMSA): we explicitly share keys/values across modalities while maintaining modality-specific queries, facilitating information exchange and inter-modal alignment. We validate our framework by instantiating it in two domains: 1) UniVid-Intrinsic for RGB videos and their intrinsic maps (albedo, irradiance, normal), and 2) UniVid-Alpha for blended RGB videos and their constituent RGBA layers. Experimental results demonstrate that both models achieve performance competitive with state-of-the-art methods across distinct tasks. Notably, they exhibit robust generalization capabilities in in-the-wild scenarios, even when trained on limited datasets of fewer than 1k videos. Houyuan Chen, Hong Li 0016, Xianghao Kong, Tianrui Zhu, Shaocong Xu, Weiqing Xiao, Yuwei Guo 0002, Chongjie Ye, Lvmin Zhang, Hao Zhao 0002, Anyi Rao |
ACM Trans. Graph. | 3 |
| 2025 | Stretching Each Dollar: Diffusion Training from Scratch on a Micro-BudgetabstractAs scaling laws in generative AI push performance, they simultaneously concentrate the development of these models among actors with large computational resources. With a focus on text-to-image (T2I) generative models, we aim to unlock this bottleneck by demonstrating very low-cost training of large-scale T2I diffusion transformer models. As the computational cost of transformers increases with the number of patches in each image, we propose randomly masking up to 75% of the image patches during training. We propose a deferred masking strategy that preprocesses all patches using a patch-mixer before masking, thus significantly reducing the performance degradation with masking, making it superior to model downscaling in reducing computational cost. We also incorporate the latest improvements in transformer architecture, such as the use of mixture-of-experts layers, to improve performance and further identify the critical benefit of using synthetic images in micro-budget training. Finally, using only 37M publicly available real and synthetic images, we train a 1.16 billion parameter sparse transformer with only $1,890 economical cost and achieve a 12.7 FID in zero-shot generation on the COCO dataset. Notably, our model achieves competitive performance across both automated and human-centric evaluations, as well as high-quality generations, while incurring 118× lower costs than Stable Diffusion models and 14× lower costs than the current state-of-the-art approach, which costs $28,400. We also further investigate the influence of synthetic images on performance and demonstrate that micro-budget training on only synthetic images is sufficient for achieving high-quality data generation. Our end-to-end training pipeline and model checkpoints are available at https://github.com/SonyResearch/micro_diffusion. Vikash Sehwag, Xianghao Kong, Michael Spranger, Lingjuan Lyu |
CVPR | 2 |
| 2025 | Denoising diffusion wavelet models for zero-shot medical image translation
Xianghao Kong, Greg Ver Steeg |
Knowl. Based Syst. | 2 |
| 2024 | Controllable Navigation Instruction Generation with Chain of Thought Prompting
Xianghao Kong, Wenguan Wang, Hang Su 0006, Xiaolin Hu 0001, Yi Yang 0001, Si Liu 0001 |
ECCV (29) | 1 |
| 2024 | Interpretable Diffusion via Information DecompositionabstractDenoising diffusion models enable conditional generation and density modeling of complex relationships like images and text.
However, the nature of the learned relationships is opaque making it difficult to understand precisely what relationships between words and parts of an image are captured, or to predict the effect of an intervention. We illuminate the fine-grained relationships learned by diffusion models by noticing a precise relationship between diffusion and information decomposition. Exact expressions for mutual information and conditional mutual information can be written in terms of the denoising model. Furthermore, ${pointwise}$ estimates can be easily estimated as well, allowing us to ask questions about the relationships between specific images and captions. Decomposing information even further to understand which variables in a high-dimensional space carry information is a long-standing problem. For diffusion models, we show that a natural non-negative decomposition of mutual information emerges, allowing us to quantify informative relationships between words and pixels in an image. We exploit these new relations to measure the compositional understanding of diffusion models, to do unsupervised localization of objects in images, and to measure effects when selectively editing images through prompt interventions. Xianghao Kong, Ollie Liu, Dani Yogatama, Greg Ver Steeg |
ICLR | 1 |
| 2024 | Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive TrainingabstractDiffusion models learn to denoise data and the trained denoiser is then used to generate new samples from the data distribution.
In this paper, we revisit the diffusion sampling process and identify a fundamental cause of sample quality degradation: the denoiser is poorly estimated in regions that are far Outside Of the training Distribution (OOD), and the sampling process inevitably evaluates in these OOD regions.
This can become problematic for all sampling methods, especially when we move to parallel sampling which requires us to initialize and update the entire sample trajectory of dynamics in parallel, leading to many OOD evaluations.
To address this problem, we introduce a new self-supervised training objective that differentiates the levels of noise added to a sample, leading to improved OOD denoising performance. The approach is based on our observation that diffusion models implicitly define a log-likelihood ratio that distinguishes distributions with different amounts of noise, and this expression depends on denoiser performance outside the standard training distribution.
We show by diverse experiments that the proposed contrastive diffusion training is effective for both sequential and parallel settings, and it improves the performance and speed of parallel samplers significantly. Code for our paper can be found at https://github.com/yunshuwu/ContrastiveDiffusionLoss Yunshu Wu, Yingtao Luo, Xianghao Kong, Evangelos E. Papalexakis, Greg Ver Steeg |
NeurIPS | 3 |
| 2024 | Anti-Jamming Hybrid Beamforming Design for Millimeter-Wave Massive MIMO SystemsabstractIn this paper, we investigate an anti-jamming hybrid beamforming (HBF) design in millimeter-wave (mmWave) massive MIMO systems for reliable wireless communications. Different from the conventional schemes designed by assuming perfect channel state information (CSI) of both communication channels and jamming channels. We aim to design the HBF scheme by minimizing the dispersion of the output signals in statistics without the knowledge of the jamming channel information. Two partially-connected HBF architectures are considered. Specifically, we propose a two-stage robust HBF design by solving the formulated non-convex problem. First, we introduce an auxiliary vector that represents the product of an analog beamformer (ABF) and a digital beamformer (DBF), and propose a pseudo-Newton method assisted beamforming algorithm by relaxing the non-convex constraint on the ABF. In the second stage, we separately optimize the ABF and DBF for two different partially-connected HBF architectures by proposing a matrix decomposition-based alternating minimization method. Finally, simulation results are provided for demonstrating the superiority of our proposed robust HBF schemes over other benchmark schemes by considering the security threat of potential jamming attacks, especially when the number of snapshots is small. Xiaolei Qi, Mugen Peng, Hongming Zhang 0001, Xianghao Kong |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Information-Theoretic Diffusion
Xianghao Kong, Rob Brekelmans, Greg Ver Steeg |
ICLR | 1 |
| 2023 | DUSA: Decoupled Unsupervised Sim2Real Adaptation for Vehicle-to-Everything Collaborative PerceptionabstractVehicle-to-Everything (V2X) collaborative perception is crucial for the advancement of autonomous driving. However, achieving high-precision V2X perception requires a significant amount of annotated real-world data, which can always be expensive and hard to acquire. Simulated data have raised much attention since they can be massively produced at an extremely low cost. Nevertheless, the significant domain gap between simulated and real-world data, including differences in sensor type, reflectance patterns, and road surroundings, often leads to poor performance of models trained on simulated data when evaluated on real-world data. In addition, there remains a domain gap between real-world collaborative agents, e.g. different types of sensors may be installed on autonomous vehicles and roadside infrastructures with different extrinsics, further increasing the difficulty of sim2real generalization. To take full advantage of simulated data, we present a new unsupervised sim2real domain adaptation method for V2X collaborative detection named Decoupled Unsupervised Sim2Real Adaptation (DUSA). Our new method decouples the V2X collaborative sim2real domain adaptation problem into two sub-problems: sim2real adaptation and inter-agent adaptation. For sim2real adaptation, we design a Location-adaptive Sim2Real Adapter (LSA) module to adaptively aggregate features from critical locations of the feature map and align the features between simulated data and real-world data via a sim/real discriminator on the aggregated global feature. For inter-agent adaptation, we further devise a Confidence-aware Inter-agent Adapter (CIA) module to align the fine-grained features from heterogeneous agents under the guidance of agent-wise confidence maps. Experiments demonstrate the effectiveness of the proposed DUSA approach on unsupervised sim2real adaptation from the simulated V2XSet dataset to the real-world DAIR-V2X-C dataset. Xianghao Kong, Jinrang Jia, Yifeng Shi, Runsheng Xu, Si Liu 0001 |
ACM Multimedia | 1 |
| 2022 | 3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selectionabstract3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage paradigm, i.e., language-irrelevant detection and cross-modal matching, which is limited by the isolated architecture. In such a paradigm, the detector needs to sample keypoints from raw point clouds due to the inherent properties of 3D point clouds (irregular and large-scale), to generate the corresponding object proposal for each keypoint. However, sparse proposals may leave out the target in detection, while dense proposals may confuse the matching model. Moreover, the language-irrelevant detection stage can only sample a small proportion of keypoints on the target, deteriorating the target prediction. In this paper, we propose a 3D Single-Stage Referred Point Progressive Selection (3D-SPS) method, which progressively selects keypoints with the guidance of language and directly locates the target. Specifically, we propose a Description-aware Keypoint Sampling (DKS) module to coarsely focus on the points of language-relevant objects, which are significant clues for grounding. Besides, we devise a Target-oriented Progressive Mining (TPM) module to finely concentrate on the points of the target, which is enabled by progressive intra-modal relation modeling and inter-modal target mining. 3D-SPS bridges the gap between detection and matching in the 3D visual grounding task, localizing the target at a single stage. Experiments demonstrate that 3D-SPS achieves state-of-the-art performance on both ScanRe-fer and Nr3D/Sr3D datasets. Junyu Luo 0002, Jiahui Fu 0003, Xianghao Kong, Chen Gao 0005, Haibing Ren, Huaxia Xia, Si Liu 0001 |
CVPR | 3 |
| 2018 | Task-Independent EEG Identification via Low-Rank Matrix Decomposition
Xianghao Kong, Wanzeng Kong, Qiaonan Fan, Qibin Zhao, Andrzej Cichocki |
BIBM | 1 |
| 2018 | Super-Resolution for GaoFen-4 Remote Sensing ImagesabstractIn this letter, the application of super-resolution (SR) techniques to GaoFen(GF)-4, which is the most advanced geostationary-orbit earth observing satellite in China, remote sensing images is investigated and tested. One of the shortcomings of the geostationary-orbit-based earth observing satellite is the limitation of spatial resolution. However, human beings never stop pursuing higher resolution in images. This is the first experiment of applying SR to a sequence of low-resolution (LR) images captured by GF-4 within a short time period. One of the barriers for applying SR to remote sensing images is the large time gaps between those LR image acquisition, because the reflection characteristic of the ground may change within the time period when those LR images were captured. However, GF-4 has the unique advantage of capturing a sequence of LR images of the same region in minutes, i.e., working as a staring camera from the point view of SR. The reconstructed high-resolution images of some regions in Beijing and Hainan are shown and evaluated in this letter. This letter demonstrates that the application of SR to geostationary-orbit-based earth observation data is feasible and valuable, and it has the potential to be applied to the images acquired by all other geostationary-orbit-based earth observing systems. Feng Li 0003, Lei Xin, Yi Guo 0001, Dongsheng Gao, Xianghao Kong, Xiuping Jia |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2014 | Zero-sequence current suppression for parallel-connected open-winding permanent magnet synchronous generation systemabstractIn order to overcome the difficulties of poor voltage regulation and narrow speed range for the traditional permanent magnet synchronous machine used in the generation system, a parallel-connected open-winding topology is adopted with the inverter in series to one side of the windings and the rectifier to another side. Due to the parallel-connected dc sides of the rectifier and inverter, the zero-sequence current would flow in the windings under the conventional space vector pulse width modulation method for the voltage regulation, and it was result in the low efficiency. Therefore, the operation principle of the parallel-connected topology was given, and the causes and components of zero-sequence current for conventional space vector pulse with modulation algorithm was analyzed. Then, the hysteresis modulation method is adopted to suppress the zero-sequence current, and the suppression efficiency and feasibility of hysteresis modulation algorithm is verified by Simulation and experimental results. Qingqing Zheng, Jiadan Wei, Bo Zhou 0014, Xianghao Kong |
IECON | 5 |