VLDB 2026 Research / reviewers in the wild / expert
Xinyu Hou
dblp:226/5294
· DBLP profile ↗
13ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
Xinyu Hou, Xiaoming Li 0002, Chen Change Loy |
Int. J. Comput. Vis. | 1 |
| 2025 | mmDiffusion: mmWave Diffusion for Sequential 3D Human Dense Point Cloud GenerationabstractMillimeter-wave (mmWave) point-cloud radar shows great promise in enabling responsive human-machine interfaces (e.g., through pose and gesture tracking and for emerging augmented reality approaches). However, generating dense and temporally consistent 3D human point clouds from sequential mmWave signals is challenging due to point-cloud sparsity, jitter, and noise. Existing approaches have made progress in single-frame densification, but are inaccurate over multiple frames. This work redefines the problem as a 3D point cloud denoising task, leveraging reverse diffusion processes to transform sparse mmWave data into detailed and accurate whole-body representations. Our proposed method, mmDiffusion, effectively exploits diffusion models and temporal context within mmWave sequences to learn the denoising process, resulting in denser and temporally coherent human point clouds. For the first time, we also introduce an evaluation metric tailored to measure temporal consistency for sequential 3D human point clouds. Experimental results demonstrate that mmDiffusion significantly outperforms existing methods. Qian Xie 0001, Xinyu Hou, Qianyi Deng, Amir Patel, Agathoniki Trigoni, Andrew Markham |
3DV | 2 |
| 2025 | Ransomware 3.0: Enhancing Risk Management and Mitigation Options with Proof-of-Decryptability and Smart Contracts
Xinyu Hou, Yang Lu 0010, Rabimba Karanjai, Lei Xu 0012, Larry Shi |
ICBC | 1 |
| 2025 | Omegance: A Single Parameter for Various Granularities in Diffusion-Based SynthesisabstractIn this work, we introduce a single parameter $ω$, to effectively control granularity in diffusion-based synthesis. This parameter is incorporated during the denoising steps of the diffusion model's reverse process. Our approach does not require model retraining, architectural modifications, or additional computational overhead during inference, yet enables precise control over the level of details in the generated outputs. Moreover, spatial masks or denoising schedules with varying $ω$ values can be applied to achieve region-specific or timestep-specific granularity control. Prior knowledge of image composition from control signals or reference images further facilitates the creation of precise $ω$ masks for granularity control on specific objects. To highlight the parameter's role in controlling subtle detail variations, the technique is named Omegance, combining "omega" and "nuance". Our method demonstrates impressive performance across various image and video synthesis tasks and is adaptable to advanced diffusion models. The code is available at https://github.com/itsmag11/Omegance. Xinyu Hou, Zongsheng Yue, Xiaoming Li 0002, Chen Change Loy |
ICCV | 1 |
| 2025 | DiffRefine: Diffusion-Based Proposal Specific Point Cloud Densification for Cross-Domain Object Detection
Sang-Yun Shin, Xinyu Hou, Samuel Hodgson, Andrew Markham, Agathoniki Trigoni |
ICCV | 3 |
| 2025 | ROCIP: robust continuous inertial position tracking for complex actions emerging from the interaction of human actors and environmentabstractInertial navigation is advancing rapidly due to improvements in sensor technology and tracking algorithms, with consumer-grade inertial measurement units (IMUs) becoming increasingly compact and affordable. Despite progress in pedestrian dead reckoning (PDR), IMU-based positional tracking still faces significant noise and bias issues. While traditional model-based methods and recent machine learning approaches have been employed to reduce signal drift, error accumulation remains a barrier to long-term system performance. Inertial tracking's self-contained nature offers broad applicability but limits integration with a global reference frame. To solve this problem, a system that could "introspect its error" and "learn from the past" is proposed. It consists of a neural statistical motion model that regresses both poses and uncertainties with DenseNet, which are then fed into Rao-Blackwellised particle filter (RBPF) for calibration with a probabilistic transition map. An inertial tracking dataset with head-mounted IMUs was collected, including walking and running with different speeds while allowing participants to rotate their heads in a self-selected manner. The dataset consisted of 19 volunteers that generated 151 sequences in 4 scenarios with a total time of 929.8 min. It was shown that our proposed method (ROCIP) outperformed the leading methods in the field, with a relative trajectory error (RTE) of 4.94m and absolute trajectory error (ATE) of 4.36m. ROCIP could also solve the problem of error accumulation in dead reckoning and maintain a small and consistent error during long-term tracking. Xinyu Hou, Jeroen H. M. Bergmann |
Appl. Intell. | 1 |
| 2024 | When StyleGAN Meets Stable Diffusion: a $\mathcal{W}_{+}$ Adapter for Personalized Image GenerationabstractText-to-image diffusion models have remarkably excelled in producing diverse, high-quality, and photo-realistic images. This advancement has spurred a growing interest in incorporating specific identities into generated content. Most current methods employ an inversion approach to em-bed a target visual concept into the text embedding space using a single reference image. However, the newly synthe-sized faces either closely resemble the reference image in terms of facial attributes, such as expression, or exhibit a reduced capacity for identity preservation. Text descriptions intended to guide the facial attributes of the synthesized face may fall short, owing to the intricate entanglement of identity information with identity-irrelevant facial attributes derived from the reference image. To address these issues, we present the novel use of the extended StyleGAN embed-ding space$\mathcal{W}+$, to achieve enhanced identity preservation and disentanglement for diffusion models. By aligning this semantically meaningful human face latent space with text-to-image diffusion models, we succeed in maintaining high fidelity in identity preservation, coupled with the capacity for semantic editing. Additionally, we propose new training objectives to balance the influences of both prompt and identity conditions, ensuring that the identity-irrelevant back-ground remains negligibly affected during facial attribute modifications. Extensive experiments reveal that our method adeptly generates personalized text-to-image outputs that are not only compatible with prompt descriptions but also amenable to common StyleGAN editing directions in diverse settings. Our code and model are available at https://github.com/csxmli2016/w-plus-adapter. Xiaoming Li 0002, Xinyu Hou, Chen Change Loy |
CVPR | 2 |
| 2024 | A self-adaptive arithmetic optimization algorithm with hybrid search modes for 0-1 knapsack problem
Mengdie Lu, Haiyan Lu, Xinyu Hou |
Neural Comput. Appl. | 3 |
| 2023 | Video Infilling with Rich Motion Prior
Xinyu Hou, Liming Jiang 0001, Chen Change Loy |
BMVC | 1 |
| 2023 | HINNet: Inertial navigation with head-mounted sensors using a neural networkabstractHuman inertial navigation systems have been developing rapidly in recent years, and it has shown great potential for applications within healthcare, smart homes, sports, and emergency services. Placing inertial measurement units on the head for localisation is relatively new. However, it provides a very interesting option, as there are several everyday head-worn items that could easily be equipped with sensors. Yet, there remains a lack of research in this area and currently no localisation solutions have been offered that allow for free head-rotations during long periods of walking. To solve this problem, we present HINNet, the first deep neural network (DNN) pedestrian inertial navigation system allowing free head movements with head-mounted inertial measurement units (IMUs), which deploys a 2-layer bi-directional LSTM. A new ’peak ratio’ feature is introduced and utilised as part of the input to the neural network. This information can be leveraged to solve the issue of differentiating between changes in movements related to the head and those that are associated with the walking pattern. A dataset with 8 subjects totalling 528 min has been collected on three different tracks for training and verification. The HINNet could effectively distinguish head rotations and changes in walking direction with a distance percentage error of 0.46%, a relative trajectory error of 3.88 m, and a absolute trajectory error of 5.98 m, which outperforms the current best head-mounted Pedestrian Dead Reckoning (PDR) method. Xinyu Hou, Jeroen H. M. Bergmann |
Eng. Appl. Artif. Intell. | 1 |
| 2021 | Joint Link Rate Selection and Channel State Change Detection in Block-Fading ChannelsabstractIn this work, we consider the problem of transmission rate selection for a discrete time point-to-point block fading wire-less communication link. The wireless channel remains constant within the channel coherence time but can change rapidly across blocks. The goal is to design a link rate selection strategy that can identify the best transmission rate quickly and adaptively in quasi-static channels. This problem can be cast into the stochastic bandit framework, and the unawareness of time-stamps where channel changes necessitates running change-point detection simultaneously with stochastic bandit algorithms to improve adaptivity. We present a joint channel change-point detection and link rate selection algorithm based on Thompson Sampling (CD-TS) and show it can achieve a sublinear regret with respect to the number of time steps$T$when the channel coherence time is larger than a threshold. We then improve the CD-TS algorithm by considering the fact that higher transmission rate has higher packet-loss probability. Finally, we validate the performance of the proposed algorithms through numerical simulations. Haoyue Tang, Xinyu Hou, Jintao Wang 0001, Jian Song 0004 |
GLOBECOM | 2 |
| 2021 | A Self-Training Approach for Point-Supervised Object Detection and Counting in CrowdsabstractIn this article, we propose a novel self-training approach named Crowd-SDNet that enables a typical object detector trained only with point-level annotations (i.e., objects are labeled with points) to estimate both the center points and sizes of crowded objects. Specifically, during training, we utilize the available point annotations to supervise the estimation of the center points of objects directly. Based on a locally-uniform distribution assumption, we initialize pseudo object sizes from the point-level supervisory information, which are then leveraged to guide the regression of object sizes via a crowdedness-aware loss. Meanwhile, we propose a confidence and order-aware refinement scheme to continuously refine the initial pseudo object sizes such that the ability of the detector is increasingly boosted to detect and count objects in crowds simultaneously. Moreover, to address extremely crowded scenes, we propose an effective decoding method to improve the detector's representation ability. Experimental results on the WiderFace benchmark show that our approach significantly outperforms state-of-the-art point-supervised methods under both detection and counting tasks, i.e., our method improves the average precision by more than 10% and reduces the counting error by 31.2%. Besides, our method obtains the best results on the crowd counting and localization datasets (i.e., ShanghaiTech and NWPU-Crowd) and vehicle counting datasets (i.e., CARPK and PUCPR+) compared with state-of-the-art counting-by-detection methods. The code will be publicly available at https://github.com/WangyiNTU/Point-supervised-crowd-detection. Yi Wang 0068, Junhui Hou, Xinyu Hou, Lap-Pui Chau |
IEEE Trans. Image Process. | 3 |
| 2019 | Vehicle Tracking Using Deep SORT with Low Confidence Track FilteringabstractMulti-object tracking (MOT) becomes an attractive topic due to its wide range of usability in video surveillance and traffic monitoring. Recent improvements on MOT has focused on tracking-by-detection manner. However, as a relatively complicated and integrated computer vision mission, state-of-the-art tracking-by-detection techniques are still suffering from issues such as a large number of false-positive tracks. To reduce the effect of unreliable detections on vehicle tracking, in this paper, we propose to incorporate a low confidence track filtering into the Simple Online and Realtime Tracking with a Deep association metric (Deep SORT) algorithm. We present a self-generated UA-DETRAC vehicle re-identification dataset which can be used to train the convolutional neural network of Deep SORT for data association. We evaluate our proposed tracker on UA-DETRAC test dataset. Experimental results show that the proposed method can improve the original Deep SORT algorithm with a significant margin. Our tracker outperforms the state-of-the-art online trackers and is comparable with batch-mode trackers. Xinyu Hou, Yi Wang 0068, Lap-Pui Chau |
AVSS | 1 |